A buddy of mine works at a place where they did the "please!
-
A buddy of mine works at a place where they did the "please! use as many llm tokes as you can!" to "oh shit this is way more expensive than the people and services we thought it was going to replace" speedrun in just a couple of months.
Their suggestion for reducing token use? "Ask Claude how you can use fewer tokens."
Corporate America is out of ideas, folks
@ricci Ask your dealer how to use less drugs.
-
A buddy of mine works at a place where they did the "please! use as many llm tokes as you can!" to "oh shit this is way more expensive than the people and services we thought it was going to replace" speedrun in just a couple of months.
Their suggestion for reducing token use? "Ask Claude how you can use fewer tokens."
Corporate America is out of ideas, folks
@ricci This is funny, but its right up there with the Nigerian prince saying "We hit a snag, but if you forward another $10,000 I'm sure we can get the money moving your way!"
-
A buddy of mine works at a place where they did the "please! use as many llm tokes as you can!" to "oh shit this is way more expensive than the people and services we thought it was going to replace" speedrun in just a couple of months.
Their suggestion for reducing token use? "Ask Claude how you can use fewer tokens."
Corporate America is out of ideas, folks
@ricci oh. They probably haven’t heard of code mode, or model routing.
-
A buddy of mine works at a place where they did the "please! use as many llm tokes as you can!" to "oh shit this is way more expensive than the people and services we thought it was going to replace" speedrun in just a couple of months.
Their suggestion for reducing token use? "Ask Claude how you can use fewer tokens."
Corporate America is out of ideas, folks
@ricci It's cooked af
-
A buddy of mine works at a place where they did the "please! use as many llm tokes as you can!" to "oh shit this is way more expensive than the people and services we thought it was going to replace" speedrun in just a couple of months.
Their suggestion for reducing token use? "Ask Claude how you can use fewer tokens."
Corporate America is out of ideas, folks
@ricci omg my friend too. along with statements like "we lost $10k in tokens due to a bug in the llm"
-
@ricci omg my friend too. along with statements like "we lost $10k in tokens due to a bug in the llm"
@ricci this and previous statements are from the utterly deranged
-
A buddy of mine works at a place where they did the "please! use as many llm tokes as you can!" to "oh shit this is way more expensive than the people and services we thought it was going to replace" speedrun in just a couple of months.
Their suggestion for reducing token use? "Ask Claude how you can use fewer tokens."
Corporate America is out of ideas, folks
@ricci make sure that shit ends up in someone's lap, but they are also doing it to squeeze headcount, while navigating a society that is becoming more dangerous because of desperate need and lack of engagement
-
@ricci my corporate training at <garbage consulting firm> back in April included the suggestion to determine if using AI for a certain task is a good idea by... asking the AI
-
@ricci it’s called session analysis and is part of tokenomics. Model choice, agents, routing, wrapping tools etc all can all substantially reduce costs. Out of the box most of these systems are designed to maximize cost with more token creation and consumption. Simple things like a cheaper capable model can save 50%.
The spend as much as you can is literally ridiculous, but sadly not surprising to still see.@drewdaniels @ricci I observe some sort of AI-enabled narcism(?) among my AI-native collegues. Their every task requires Fable, and now Astra, because ofc their problems can't be solved by small, older or local LLMs. Where, well, 99% of the tasks is some basic script like : sort according to another column.
This is the same mindset as buying polar-expedition grade gear to walk 100m from a metro station to the office. Tech industry's culture is sometimes plain cringe.
-
@ricci I can see the response.
"I'd be happy to help with that! Enable agent and planning modes, select {{latest model}}, and make sure to upload your entire hard drive with every request"
-
A buddy of mine works at a place where they did the "please! use as many llm tokes as you can!" to "oh shit this is way more expensive than the people and services we thought it was going to replace" speedrun in just a couple of months.
Their suggestion for reducing token use? "Ask Claude how you can use fewer tokens."
Corporate America is out of ideas, folks
Drawing on personal experience I can provide useful advice on saving tokens.
Hire some freaking humans, bozos! That way you get REAL intelligence!!! Not all that make believe crap!!
-
A buddy of mine works at a place where they did the "please! use as many llm tokes as you can!" to "oh shit this is way more expensive than the people and services we thought it was going to replace" speedrun in just a couple of months.
Their suggestion for reducing token use? "Ask Claude how you can use fewer tokens."
Corporate America is out of ideas, folks
@ricci What about the military-industrial complex in the United States?
-
A buddy of mine works at a place where they did the "please! use as many llm tokes as you can!" to "oh shit this is way more expensive than the people and services we thought it was going to replace" speedrun in just a couple of months.
Their suggestion for reducing token use? "Ask Claude how you can use fewer tokens."
Corporate America is out of ideas, folks
@ricci I have never watched the movie Idiocracy (I picked up enough of the synopsis I was quite sure it would make me itch) but I feel like someone has taken it to be an instructional video.
-
@drewdaniels @ricci I observe some sort of AI-enabled narcism(?) among my AI-native collegues. Their every task requires Fable, and now Astra, because ofc their problems can't be solved by small, older or local LLMs. Where, well, 99% of the tasks is some basic script like : sort according to another column.
This is the same mindset as buying polar-expedition grade gear to walk 100m from a metro station to the office. Tech industry's culture is sometimes plain cringe.
@hiiamfrompoland @ricci it can be cringe. There are a great many people using the technology that choose cheaper defaults and consult benchmarks. Benchmarks are a hot topic for LLM’s. I’ve heard many people and companies default to cheaper models like Sonnet or even run their own models.
Open weight models are hot for a reason too. Qwen3-coder-next (a local MoE model) benchmarks almost at (many people’s old default) Sonnet 4.6 and can be quantized to run in 30gb. -
@hiiamfrompoland @ricci it can be cringe. There are a great many people using the technology that choose cheaper defaults and consult benchmarks. Benchmarks are a hot topic for LLM’s. I’ve heard many people and companies default to cheaper models like Sonnet or even run their own models.
Open weight models are hot for a reason too. Qwen3-coder-next (a local MoE model) benchmarks almost at (many people’s old default) Sonnet 4.6 and can be quantized to run in 30gb.@drewdaniels or, and bear with us here: STOP USING THE FUCKING SLOP BOT
-
@drewdaniels or, and bear with us here: STOP USING THE FUCKING SLOP BOT
@atax1a that makes a lot of sense.
-
A buddy of mine works at a place where they did the "please! use as many llm tokes as you can!" to "oh shit this is way more expensive than the people and services we thought it was going to replace" speedrun in just a couple of months.
Their suggestion for reducing token use? "Ask Claude how you can use fewer tokens."
Corporate America is out of ideas, folks
@ricci
Working at a place where the leadership is falling deeper and deeper into full-on AI psychosis is fucking rough. -
@stib it was, but I left it on purpose after I spotted it
-
A buddy of mine works at a place where they did the "please! use as many llm tokes as you can!" to "oh shit this is way more expensive than the people and services we thought it was going to replace" speedrun in just a couple of months.
Their suggestion for reducing token use? "Ask Claude how you can use fewer tokens."
Corporate America is out of ideas, folks
@ricci
"Ask Claude how you can use fewer tokens." needs to be a t-shirt -
@ricci it’s called session analysis and is part of tokenomics. Model choice, agents, routing, wrapping tools etc all can all substantially reduce costs. Out of the box most of these systems are designed to maximize cost with more token creation and consumption. Simple things like a cheaper capable model can save 50%.
The spend as much as you can is literally ridiculous, but sadly not surprising to still see.@drewdaniels @ricci I've been spending time with Bifrost llm proxy lately. You can give engineers a virtual token and restrict them to "engineering-plan" and "engineering-build", which you map to whatever two models make the most sense at the time. It also has context based routing where it tries to switch between plan and build for you but I've not used that yet.
-
J jwcph@helvede.net shared this topic

