A buddy of mine works at a place where they did the "please!
-
@hiiamfrompoland @ricci it can be cringe. There are a great many people using the technology that choose cheaper defaults and consult benchmarks. Benchmarks are a hot topic for LLM’s. I’ve heard many people and companies default to cheaper models like Sonnet or even run their own models.
Open weight models are hot for a reason too. Qwen3-coder-next (a local MoE model) benchmarks almost at (many people’s old default) Sonnet 4.6 and can be quantized to run in 30gb.@drewdaniels or, and bear with us here: STOP USING THE FUCKING SLOP BOT
-
@drewdaniels or, and bear with us here: STOP USING THE FUCKING SLOP BOT
@atax1a that makes a lot of sense.
-
A buddy of mine works at a place where they did the "please! use as many llm tokes as you can!" to "oh shit this is way more expensive than the people and services we thought it was going to replace" speedrun in just a couple of months.
Their suggestion for reducing token use? "Ask Claude how you can use fewer tokens."
Corporate America is out of ideas, folks
@ricci
Working at a place where the leadership is falling deeper and deeper into full-on AI psychosis is fucking rough. -
@stib it was, but I left it on purpose after I spotted it
-
A buddy of mine works at a place where they did the "please! use as many llm tokes as you can!" to "oh shit this is way more expensive than the people and services we thought it was going to replace" speedrun in just a couple of months.
Their suggestion for reducing token use? "Ask Claude how you can use fewer tokens."
Corporate America is out of ideas, folks
@ricci
"Ask Claude how you can use fewer tokens." needs to be a t-shirt -
@ricci it’s called session analysis and is part of tokenomics. Model choice, agents, routing, wrapping tools etc all can all substantially reduce costs. Out of the box most of these systems are designed to maximize cost with more token creation and consumption. Simple things like a cheaper capable model can save 50%.
The spend as much as you can is literally ridiculous, but sadly not surprising to still see.@drewdaniels @ricci I've been spending time with Bifrost llm proxy lately. You can give engineers a virtual token and restrict them to "engineering-plan" and "engineering-build", which you map to whatever two models make the most sense at the time. It also has context based routing where it tries to switch between plan and build for you but I've not used that yet.
-
J jwcph@helvede.net shared this topic