Skip to content
  • Hjem
  • Seneste
  • Etiketter
  • Populære
  • Verden
  • Bruger
  • Grupper
Temaer
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Kollaps
FARVEL BIG TECH
  1. Forside
  2. Ikke-kategoriseret
  3. A buddy of mine works at a place where they did the "please!

A buddy of mine works at a place where they did the "please!

Planlagt Fastgjort Låst Flyttet Ikke-kategoriseret
26 Indlæg 21 Posters 0 Visninger
  • Ældste til nyeste
  • Nyeste til ældste
  • Most Votes
Svar
  • Svar som emne
Login for at svare
Denne tråd er blevet slettet. Kun brugere med emne behandlings privilegier kan se den.
  • drewdaniels@mastodon.onlineD drewdaniels@mastodon.online

    @hiiamfrompoland @ricci it can be cringe. There are a great many people using the technology that choose cheaper defaults and consult benchmarks. Benchmarks are a hot topic for LLM’s. I’ve heard many people and companies default to cheaper models like Sonnet or even run their own models.
    Open weight models are hot for a reason too. Qwen3-coder-next (a local MoE model) benchmarks almost at (many people’s old default) Sonnet 4.6 and can be quantized to run in 30gb.

    atax1a@infosec.exchangeA This user is from outside of this forum
    atax1a@infosec.exchangeA This user is from outside of this forum
    atax1a@infosec.exchange
    wrote sidst redigeret af
    #21

    @drewdaniels or, and bear with us here: STOP USING THE FUCKING SLOP BOT

    drewdaniels@mastodon.onlineD 1 Reply Last reply
    0
    • atax1a@infosec.exchangeA atax1a@infosec.exchange

      @drewdaniels or, and bear with us here: STOP USING THE FUCKING SLOP BOT

      drewdaniels@mastodon.onlineD This user is from outside of this forum
      drewdaniels@mastodon.onlineD This user is from outside of this forum
      drewdaniels@mastodon.online
      wrote sidst redigeret af
      #22

      @atax1a that makes a lot of sense.

      1 Reply Last reply
      0
      • ricci@discuss.systemsR ricci@discuss.systems

        A buddy of mine works at a place where they did the "please! use as many llm tokes as you can!" to "oh shit this is way more expensive than the people and services we thought it was going to replace" speedrun in just a couple of months.

        Their suggestion for reducing token use? "Ask Claude how you can use fewer tokens."

        Corporate America is out of ideas, folks

        jargoggles@kolektiva.socialJ This user is from outside of this forum
        jargoggles@kolektiva.socialJ This user is from outside of this forum
        jargoggles@kolektiva.social
        wrote sidst redigeret af
        #23

        @ricci
        Working at a place where the leadership is falling deeper and deeper into full-on AI psychosis is fucking rough.

        1 Reply Last reply
        0
        • ricci@discuss.systemsR This user is from outside of this forum
          ricci@discuss.systemsR This user is from outside of this forum
          ricci@discuss.systems
          wrote sidst redigeret af
          #24

          @stib it was, but I left it on purpose after I spotted it

          1 Reply Last reply
          0
          • ricci@discuss.systemsR ricci@discuss.systems

            A buddy of mine works at a place where they did the "please! use as many llm tokes as you can!" to "oh shit this is way more expensive than the people and services we thought it was going to replace" speedrun in just a couple of months.

            Their suggestion for reducing token use? "Ask Claude how you can use fewer tokens."

            Corporate America is out of ideas, folks

            lefou23@c.imL This user is from outside of this forum
            lefou23@c.imL This user is from outside of this forum
            lefou23@c.im
            wrote sidst redigeret af
            #25

            @ricci
            "Ask Claude how you can use fewer tokens." needs to be a t-shirt

            1 Reply Last reply
            0
            • drewdaniels@mastodon.onlineD drewdaniels@mastodon.online

              @ricci it’s called session analysis and is part of tokenomics. Model choice, agents, routing, wrapping tools etc all can all substantially reduce costs. Out of the box most of these systems are designed to maximize cost with more token creation and consumption. Simple things like a cheaper capable model can save 50%.
              The spend as much as you can is literally ridiculous, but sadly not surprising to still see.

              wren@discuss.systemsW This user is from outside of this forum
              wren@discuss.systemsW This user is from outside of this forum
              wren@discuss.systems
              wrote sidst redigeret af
              #26

              @drewdaniels @ricci I've been spending time with Bifrost llm proxy lately. You can give engineers a virtual token and restrict them to "engineering-plan" and "engineering-build", which you map to whatever two models make the most sense at the time. It also has context based routing where it tries to switch between plan and build for you but I've not used that yet.

              1 Reply Last reply
              0
              • jwcph@helvede.netJ jwcph@helvede.net shared this topic
              Svar
              • Svar som emne
              Login for at svare
              • Ældste til nyeste
              • Nyeste til ældste
              • Most Votes


              • Log ind

              • Login or register to search.
              Powered by NodeBB Contributors
              Graciously hosted by data.coop
              • First post
                Last post
              0
              • Hjem
              • Seneste
              • Etiketter
              • Populære
              • Verden
              • Bruger
              • Grupper