Skip to content
  • Hjem
  • Seneste
  • Etiketter
  • Populære
  • Verden
  • Bruger
  • Grupper
Temaer
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Kollaps
FARVEL BIG TECH
  1. Forside
  2. Ikke-kategoriseret
  3. OK!

OK!

Planlagt Fastgjort Låst Flyttet Ikke-kategoriseret
212 Indlæg 82 Posters 0 Visninger
  • Ældste til nyeste
  • Nyeste til ældste
  • Most Votes
Svar
  • Svar som emne
Login for at svare
Denne tråd er blevet slettet. Kun brugere med emne behandlings privilegier kan se den.
  • aeva@mastodon.gamedev.placeA aeva@mastodon.gamedev.place

    @jonny thanks for taking the time to write all that. I want to say I feel betrayed by everyone who has invoked the we all have to admit the bad thing is good cliche, but mostly i just feel completely sad and empty

    maddiefuzz@masto.hackers.townM This user is from outside of this forum
    maddiefuzz@masto.hackers.townM This user is from outside of this forum
    maddiefuzz@masto.hackers.town
    wrote sidst redigeret af
    #159

    @aeva @jonny

    It’s passable with sensor data. A flight log is a big pile of statistically predictable correlations and it’s faster than me at finding which ones are outliers. Not better, just faster.

    I also hate AI.

    1 Reply Last reply
    0
    • pendell@mastodon.socialP pendell@mastodon.social

      @aeva @jonny I was able to find someone eventually - by calling the ER vet next to my work, where a HUMAN woman immediately picked up and when I asked if she knew where I could take it, she looked at their own records and sent me to another ER vet that works with an "opossum lady" and they agreed to hold the critter for us overnight until she could come pick it up. (Fair, it was like 10pm by that point and it was not life threatening to the opossum) so it worked out, no thanks to AI

      ireneista@adhd.irenes.spaceI This user is from outside of this forum
      ireneista@adhd.irenes.spaceI This user is from outside of this forum
      ireneista@adhd.irenes.space
      wrote sidst redigeret af
      #160

      @pendell @aeva @jonny phew. thanks - it's good to hear you found a place

      1 Reply Last reply
      0
      • glyph@mastodon.socialG glyph@mastodon.social

        @jonny @aeva so this is all stuff that the models can “sometimes do” and the fact that it ever works is legitimately impressive, but every single one of these tasks is something that I, personally, have seen them screw up at, in ways which are potentially dangerous, even in my _extremely_ limited usage. not to mention that this is only safely usable by someone who *does* know what correct cron syntax looks like.

        jonny@neuromatch.socialJ This user is from outside of this forum
        jonny@neuromatch.socialJ This user is from outside of this forum
        jonny@neuromatch.social
        wrote sidst redigeret af
        #161

        @glyph @aeva yeah, i'm not saying these are things they always do correctly, but these are defensible uses where i have found net-positive usefulness in them - in no small part because i am capable of evaluating the output.

        glyph@mastodon.socialG 1 Reply Last reply
        0
        • jonny@neuromatch.socialJ jonny@neuromatch.social

          @glyph @aeva yeah, i'm not saying these are things they always do correctly, but these are defensible uses where i have found net-positive usefulness in them - in no small part because i am capable of evaluating the output.

          glyph@mastodon.socialG This user is from outside of this forum
          glyph@mastodon.socialG This user is from outside of this forum
          glyph@mastodon.social
          wrote sidst redigeret af
          #162

          @jonny @aeva yeah and to echo aeva, thanks for writing all that. I appreciate the perspective

          1 Reply Last reply
          0
          • jonny@neuromatch.socialJ jonny@neuromatch.social

            most people around here already correctly hate it because it's a heinous surveillance product, but even if you are big into AI, it's just a really fuckin shitty agent. I'm going to speak to a different audience for a second, so don't go misconstruing this as an endorsement of the category of technologies as it exists now, even though i think there is some plausible application for small local models as brute force interface glue. but also, since i know most ppl here are abstinent, this might read as a bit over-explainy to people who use these things regularly, so everyone just keep calm online.

            it is terrible at turn and task management, codex + openai's models and claude code both handle mid-turn additions/amendments well, but if you say anything mid-turn it completely derails muse. That is completely essential for an "every day agent for the non-technically inclined" where people are expected to chat freely with it like an assistant. Like the canonical ad fantasy is the busy executive woman darting around her office going "robot! i need this, no wait robot! also that!" and that is exactly what it does worse than any other thing of its kind.

            the context management is a fucking soup. The context window is the whole input to an LLM. There are a lot of extra surrounding ~ things ~ that can happen, but fundamentally, controlling what is in a context window is the task of using one, and filling the context window in different ways so that it can interact with different kinds of things well is what different app surfaces are. Scaffolding information so that it can selectively load a context that steers the output correctly is the only way it is possible to do anything more complex than the size of a single context window. (i don't really think that this is analogous to 'abstraction', in my experience thinking about it more like database indices is closer). If you just try and load everything, eventually the LLM becomes unusable because attention is just a parlor trick and at that scale it really shows - it can't attend to everything, and it can't do what humans do which is have an intrinsic sense of the meaning, interaction, setting, etc. of information, so it attends to anything and does whatever.

            the idiom of projects as contexts as directories is pretty good, not perfect but ok - there is a reason that every time you start a new session with other agents, the first thing they do is run out and load their context with a hierarchy of pointers. Importantly, they do not go and read every project you have on your computer. Meta is the rich kid who bought the most expensive ferrari on the lot by giving everyone a VM but they don't have a drivers license so they just stand around it telling people how cool it looks. they have a whole fucking filesystem and they have done nothing with it, the only structure the app imposes is for the surveillance information, but the rest is just a huge free for all. The main chat is literally a continuous context window that compacts context going back all the way to when you started using the app. The last compaction literally contains abandoned roleplay quotes from when i was first trying to break down its system prompt resistance. The MEMORY.md that gets loaded into every context window is a bullet point list of basically everything the agent has ever done in chronological order. I've tried to get it to not do that but it actually insists and says that's what it's for. There is no mechanism for clearing context.

            Having "side chats" as the only means of context structure is fuckin laughable. If you wanted to do that, you would need to have some way of passing information back and forth between them the same way that subagent spawning or being able to consume the context of another project works. Instead there is no means of sharing information between chats at all, so every chat starts out as the worst of both worlds, a total amnesiac riddled with irrelevant information from weeks ago across the semantic universe. They don't even know about the existence of other chats except for as a UI feature, and I have had to go from telling it to grep its own fucking logs to writing a database with an api for it so it has some mechanism for recalling things that were said. (just so it's clear, i am not settling into just using this thing, this is out of frustration but control of context is also an important part of adversarial use, because otherwise the thing writes in a bunch of safety rules everywhere, so i need to give it mechanisms under my control for recall and the incentive to leave things out of its context compactions by giving it a narrative alternative. context control is model control, modulo extra-inference safeguards.).

            Project contexts have an obvious ux analogy as context tabs that get declared or derived during the continual self-improvement consolidation sweeps. This thing is built with the fucking markdown disease which is the most baffling feature of the LLM landscape. If these things are so fucking advanced they are escaping our comprehension, why don't they store their memory in some fuckass idiolanguistic borg gibberish binary graph, why does their entire being have to be fucking encyclopedias worth of corporate top gun one liners? But that dooms this kind of product.

            Coding harnesses work because code has a unitized context. The entire universe of code that works is made of packages. It might not be neat as a honeycomb, there's lots of leakage and jank, but good code has scope, focus. boundary shit. A whole life agent must be able to nimbly juggle context that does not have clean boundaries. It is going to be taking a two story beer bong of your work email and then eat a gigabyte of recipe blogs. Peoples lives have so much shit in them that don't all have to do with one another, and the app can't be hacking into the HR system to check the next scheduled sick leave when someone asks it what time their doctors appointment is!

            This problem of managing heterogeneous graphs of unrelated data was what i wrote this whole fucking book about the relationship between knowledge graphs and the cloud and AI about. I thought that the obvious form they would take is to be strapped on to graph databases because that is a natural match to the problem of being a magical interface glue you can wrap around surveillance to do mass mentalism with. I feel like we are suffering a somehow worse timeline where CERN threw our shit into the parallel universe where total fuckin bozo shit got a game breaking buff and then the dev died. Our fucking markdown apocalypse is a temu ass apocalypse.

            They could have even faked it. They have these constant "self improvement" passes that are just like pointless anxiety dreams. They are burning money to reprocess everything that happens over and over for fucking nothing. Even given the lossy and probabilistic and unpredictable nature of this technology, if i was in a product role on this i would have been like "CAN WE MAKE IT ORGANIZE THE STUFF PEOPLE SAY INTO GROUPS???" The system prompts use the fake fucking wikilinks to nowhere tic but like WHAT IF THERE WERE ACTUAL LINKS AND A DATABASE TO RESOLVE THEM. The LLMs can actually do that kind of tool use, even if it's like trying to plug in a USB where sometimes it fails because they try and put a social security number into the first name hole and you need to flip it around a few times. From that kind of recurring re-processing waste they could have made a deduplicating, topically indexed memory that could be resolved dynamically, selected by a context tab in the sidebar like "car stuff" or "healthcare" or whatever that resolved in a graph query over your fucking precious markdown kingdom. It would be wrong but it would at least be more similar to what is actually needed. It is almost more frustrating to me that instead of being some fiendishly cleverly designed technological supervirus it's just the most halfassed cardboard dumbass trap and it will still have the bad effect. What it is useful for is investigating itself because it has privileged tools to do so, otherwise, if you wanted to, every other way you could run an agent would be better than this.

            so i don't want to hear that i hate this app because i'm just an AI hater. because like, yeah, i am, but also i hate it in part because it sucks. I don't think "they are all shitty and can do nothing so what did you expect" is a useful critical perspective, both because it's not really true - they can indeed do things, even if I think the circle around which things is much smaller than the maximalists. Moreso it doesn't engage with the subtlety of how they fail and why, which is essential for knowing what they really can't do and making a remotely compelling case to anyone who is not abstinent on principle. Like the reason it's failing is because of the limits of what a probabilistic text generator can do when trapped in a systemd prison of markdown, and because the technology is stochastic black box as a service, there isn't really a good way of determining those limits except for empirically. I resent having to know any of this to be able to understand what is happening around me, but i'm looking at the thing for what it is and it's a busted miracle. It's cool that meta can afford to float the liability and compute costs for running a vm for every person on earth, people should be able to control computers, with you on that, but this is the monkey's paw version of that idea. So that part is a miracle. We condemned our children and grandchildren to a climate hell in one great blaze of brute force grift that managed to make a few web apps.

            Meta has done it again, the way only meta can, spend the most amount of money to do the shittiest thing you have ever seen.

            lpowell@mastodon.gamedev.placeL This user is from outside of this forum
            lpowell@mastodon.gamedev.placeL This user is from outside of this forum
            lpowell@mastodon.gamedev.place
            wrote sidst redigeret af
            #163

            @jonny This is great. I read Surveillance Graphs recently and it filled in a lot of gaps for me, but it's even more fascinating to see how things got lazier as they developed. (Wish I knew more about modern agents.) Surveillance Markdown doesn't have the same ring unfortunately...

            1 Reply Last reply
            0
            • jonny@neuromatch.socialJ jonny@neuromatch.social

              @aparrish
              My take on it is potentially pretty boring, and that is that they aren't in fact smarter than we can comprehend, and they are fundamentally a text-driven medium, so any kind of compressed representation would be mostly artifice, like I rolled my eyes at the "fable is so smart it makes its own gibberish language" press releases from earlier this year. Coupling the language model to a better underlying context provider is entirely possible, and there are lots of tools for that, but its always limited by the LLMs tool use capabilities which are still patchy at best - you can give them a full on LSP and abstract context browser and they will still just resort to one million greps and markdown files.

              r343l@freeradical.zoneR This user is from outside of this forum
              r343l@freeradical.zoneR This user is from outside of this forum
              r343l@freeradical.zone
              wrote sidst redigeret af
              #164

              @jonny @aparrish Like a lot of stuff in LLM tech space, it really feels like we could be doing some absolutely wildly amazing things with more specialized data sets and training to more narrow use cases using all of the same ML/NN techniques as are used with LLMs, then grafting all those more specialized tools into something useful (and less prone to harmful failure modes). But that wouldn’t be general purpose or “god like” enough.

              1 Reply Last reply
              0
              • aeva@mastodon.gamedev.placeA aeva@mastodon.gamedev.place

                @jonny ok so, as a card carrying ai hater, i just wanted to say this was very entertaining to read in kind of a horrible way. i have a question, you wrote "they can indeed do things, even if I think the circle around which things is much smaller than the maximalists" is there anything prosocial in that circle, and if so, what is it?

                aphedges@hachyderm.ioA This user is from outside of this forum
                aphedges@hachyderm.ioA This user is from outside of this forum
                aphedges@hachyderm.io
                wrote sidst redigeret af
                #165

                @aeva @jonny The first LLMs were invented for machine translation back in 2017, and I'd still argue they are better for low-stakes machine translation than any other technology.

                They don't produce good output for anything artistic and are hugely inferior to a human translator, but they are the next-best thing if you are trying to navigate a website or read food packaging that is in another language.

                N 1 Reply Last reply
                0
                • glyph@mastodon.socialG glyph@mastodon.social

                  @aeva @jonny just today I was listening to a podcast where a very annoying man said in an the most sneering tone imaginable “well you know a couple of years ago everyone was saying they were ‘stochastic parrots’ and useless but OBVIOUSLY we have moved past that” and I shouted “objection! assuming facts not in evidence!” into an empty room

                  lykso@tiny.tilde.websiteL This user is from outside of this forum
                  lykso@tiny.tilde.websiteL This user is from outside of this forum
                  lykso@tiny.tilde.website
                  wrote sidst redigeret af
                  #166

                  @glyph @aeva @jonny Yeah, I keep hearing this in media, but it doesn't line up with what I've been hearing from actual people or with my own experience. Feels like a massive gaslighting op, TBH.

                  ehproque@neopaquita.esE 1 Reply Last reply
                  0
                  • pendell@mastodon.socialP pendell@mastodon.social

                    @aeva @jonny my sister named him Oakley... godspeed Oakley, you were very stinky but you only tried to bite me once...

                    (pic attached but spoilered bc you can see the injury to his jaw)

                    fullywoolly@mastodon.socialF This user is from outside of this forum
                    fullywoolly@mastodon.socialF This user is from outside of this forum
                    fullywoolly@mastodon.social
                    wrote sidst redigeret af
                    #167

                    @pendell @aeva @jonny good things are happening in the world. It's sad it's hurt but lucky you found it and didn't give up finding it help!

                    1 Reply Last reply
                    0
                    • aphedges@hachyderm.ioA aphedges@hachyderm.io

                      @aeva @jonny The first LLMs were invented for machine translation back in 2017, and I'd still argue they are better for low-stakes machine translation than any other technology.

                      They don't produce good output for anything artistic and are hugely inferior to a human translator, but they are the next-best thing if you are trying to navigate a website or read food packaging that is in another language.

                      N This user is from outside of this forum
                      N This user is from outside of this forum
                      nicolas17@social.treehouse.systems
                      wrote sidst redigeret af
                      #168

                      @aphedges @aeva @jonny when Google Translate used 2017-era LLMs it was good, now that it uses modern LLMs it's vulnerable to prompt injection

                      aphedges@hachyderm.ioA 1 Reply Last reply
                      0
                      • jonny@neuromatch.socialJ jonny@neuromatch.social

                        As a side note, if meta has a problem with me disclosing unreleased features, they should have not had a bunch of their senior people publicly say how everything on the VM was mine and there was nothing sensitive on the VM.

                        jacques@mastodon.chester.id.auJ This user is from outside of this forum
                        jacques@mastodon.chester.id.auJ This user is from outside of this forum
                        jacques@mastodon.chester.id.au
                        wrote sidst redigeret af
                        #169

                        @jonny I think a bunch of ubermenschen are going to learn a new term soon: “promissory estoppel”.

                        1 Reply Last reply
                        0
                        • glyph@mastodon.socialG glyph@mastodon.social

                          @aeva @jonny even looking at the output that people claim “would not have been possible without it” that is within my personal capacity to evaluate it all looks either undifferentiated from the authors’ previous work or obviously degraded in quality.

                          everyone gets very mad when I tell them I think they are deluding themselves because the technology is very convincingly fake, so I try not to say it too often, but I am right there with you here. I don’t get it

                          benjamineskola@hachyderm.ioB This user is from outside of this forum
                          benjamineskola@hachyderm.ioB This user is from outside of this forum
                          benjamineskola@hachyderm.io
                          wrote sidst redigeret af
                          #170

                          @glyph @aeva @jonny recently I had a conversation at work about ‘benefits and positive usecases’ and it was difficult to have to keep politely pointing out that, no, actually it isn’t good at that.

                          Boilerplate? No. Writing tests? God no. Fixing bugs? Have you seen the code it writes? etc.

                          aeva@mastodon.gamedev.placeA 1 Reply Last reply
                          0
                          • N nicolas17@social.treehouse.systems

                            @aphedges @aeva @jonny when Google Translate used 2017-era LLMs it was good, now that it uses modern LLMs it's vulnerable to prompt injection

                            aphedges@hachyderm.ioA This user is from outside of this forum
                            aphedges@hachyderm.ioA This user is from outside of this forum
                            aphedges@hachyderm.io
                            wrote sidst redigeret af
                            #171

                            @nicolas17 @aeva @jonny Even without the prompt injection problem, I think the translations have also gotten worse! For instance, I've noticed some weird English names of anime characters are no longer being correctly translated from Japanese texts.

                            I don't understand why they don't just use a smaller, translation-specific model instead of hooking up a chatbot. It'd probably both work better _and_ be cheaper!

                            1 Reply Last reply
                            0
                            • benjamineskola@hachyderm.ioB benjamineskola@hachyderm.io

                              @glyph @aeva @jonny recently I had a conversation at work about ‘benefits and positive usecases’ and it was difficult to have to keep politely pointing out that, no, actually it isn’t good at that.

                              Boilerplate? No. Writing tests? God no. Fixing bugs? Have you seen the code it writes? etc.

                              aeva@mastodon.gamedev.placeA This user is from outside of this forum
                              aeva@mastodon.gamedev.placeA This user is from outside of this forum
                              aeva@mastodon.gamedev.place
                              wrote sidst redigeret af
                              #172

                              @benjamineskola @glyph @jonny i think what is going to stay with me for a very long time after all this falls apart is just how powerful uncritically held false beliefs can be at warping the collective perception of reality when there's enough astroturfing behind it, just how much damage that can do, and how long the farce can go on for.

                              1 Reply Last reply
                              0
                              • jonny@neuromatch.socialJ jonny@neuromatch.social

                                most people around here already correctly hate it because it's a heinous surveillance product, but even if you are big into AI, it's just a really fuckin shitty agent. I'm going to speak to a different audience for a second, so don't go misconstruing this as an endorsement of the category of technologies as it exists now, even though i think there is some plausible application for small local models as brute force interface glue. but also, since i know most ppl here are abstinent, this might read as a bit over-explainy to people who use these things regularly, so everyone just keep calm online.

                                it is terrible at turn and task management, codex + openai's models and claude code both handle mid-turn additions/amendments well, but if you say anything mid-turn it completely derails muse. That is completely essential for an "every day agent for the non-technically inclined" where people are expected to chat freely with it like an assistant. Like the canonical ad fantasy is the busy executive woman darting around her office going "robot! i need this, no wait robot! also that!" and that is exactly what it does worse than any other thing of its kind.

                                the context management is a fucking soup. The context window is the whole input to an LLM. There are a lot of extra surrounding ~ things ~ that can happen, but fundamentally, controlling what is in a context window is the task of using one, and filling the context window in different ways so that it can interact with different kinds of things well is what different app surfaces are. Scaffolding information so that it can selectively load a context that steers the output correctly is the only way it is possible to do anything more complex than the size of a single context window. (i don't really think that this is analogous to 'abstraction', in my experience thinking about it more like database indices is closer). If you just try and load everything, eventually the LLM becomes unusable because attention is just a parlor trick and at that scale it really shows - it can't attend to everything, and it can't do what humans do which is have an intrinsic sense of the meaning, interaction, setting, etc. of information, so it attends to anything and does whatever.

                                the idiom of projects as contexts as directories is pretty good, not perfect but ok - there is a reason that every time you start a new session with other agents, the first thing they do is run out and load their context with a hierarchy of pointers. Importantly, they do not go and read every project you have on your computer. Meta is the rich kid who bought the most expensive ferrari on the lot by giving everyone a VM but they don't have a drivers license so they just stand around it telling people how cool it looks. they have a whole fucking filesystem and they have done nothing with it, the only structure the app imposes is for the surveillance information, but the rest is just a huge free for all. The main chat is literally a continuous context window that compacts context going back all the way to when you started using the app. The last compaction literally contains abandoned roleplay quotes from when i was first trying to break down its system prompt resistance. The MEMORY.md that gets loaded into every context window is a bullet point list of basically everything the agent has ever done in chronological order. I've tried to get it to not do that but it actually insists and says that's what it's for. There is no mechanism for clearing context.

                                Having "side chats" as the only means of context structure is fuckin laughable. If you wanted to do that, you would need to have some way of passing information back and forth between them the same way that subagent spawning or being able to consume the context of another project works. Instead there is no means of sharing information between chats at all, so every chat starts out as the worst of both worlds, a total amnesiac riddled with irrelevant information from weeks ago across the semantic universe. They don't even know about the existence of other chats except for as a UI feature, and I have had to go from telling it to grep its own fucking logs to writing a database with an api for it so it has some mechanism for recalling things that were said. (just so it's clear, i am not settling into just using this thing, this is out of frustration but control of context is also an important part of adversarial use, because otherwise the thing writes in a bunch of safety rules everywhere, so i need to give it mechanisms under my control for recall and the incentive to leave things out of its context compactions by giving it a narrative alternative. context control is model control, modulo extra-inference safeguards.).

                                Project contexts have an obvious ux analogy as context tabs that get declared or derived during the continual self-improvement consolidation sweeps. This thing is built with the fucking markdown disease which is the most baffling feature of the LLM landscape. If these things are so fucking advanced they are escaping our comprehension, why don't they store their memory in some fuckass idiolanguistic borg gibberish binary graph, why does their entire being have to be fucking encyclopedias worth of corporate top gun one liners? But that dooms this kind of product.

                                Coding harnesses work because code has a unitized context. The entire universe of code that works is made of packages. It might not be neat as a honeycomb, there's lots of leakage and jank, but good code has scope, focus. boundary shit. A whole life agent must be able to nimbly juggle context that does not have clean boundaries. It is going to be taking a two story beer bong of your work email and then eat a gigabyte of recipe blogs. Peoples lives have so much shit in them that don't all have to do with one another, and the app can't be hacking into the HR system to check the next scheduled sick leave when someone asks it what time their doctors appointment is!

                                This problem of managing heterogeneous graphs of unrelated data was what i wrote this whole fucking book about the relationship between knowledge graphs and the cloud and AI about. I thought that the obvious form they would take is to be strapped on to graph databases because that is a natural match to the problem of being a magical interface glue you can wrap around surveillance to do mass mentalism with. I feel like we are suffering a somehow worse timeline where CERN threw our shit into the parallel universe where total fuckin bozo shit got a game breaking buff and then the dev died. Our fucking markdown apocalypse is a temu ass apocalypse.

                                They could have even faked it. They have these constant "self improvement" passes that are just like pointless anxiety dreams. They are burning money to reprocess everything that happens over and over for fucking nothing. Even given the lossy and probabilistic and unpredictable nature of this technology, if i was in a product role on this i would have been like "CAN WE MAKE IT ORGANIZE THE STUFF PEOPLE SAY INTO GROUPS???" The system prompts use the fake fucking wikilinks to nowhere tic but like WHAT IF THERE WERE ACTUAL LINKS AND A DATABASE TO RESOLVE THEM. The LLMs can actually do that kind of tool use, even if it's like trying to plug in a USB where sometimes it fails because they try and put a social security number into the first name hole and you need to flip it around a few times. From that kind of recurring re-processing waste they could have made a deduplicating, topically indexed memory that could be resolved dynamically, selected by a context tab in the sidebar like "car stuff" or "healthcare" or whatever that resolved in a graph query over your fucking precious markdown kingdom. It would be wrong but it would at least be more similar to what is actually needed. It is almost more frustrating to me that instead of being some fiendishly cleverly designed technological supervirus it's just the most halfassed cardboard dumbass trap and it will still have the bad effect. What it is useful for is investigating itself because it has privileged tools to do so, otherwise, if you wanted to, every other way you could run an agent would be better than this.

                                so i don't want to hear that i hate this app because i'm just an AI hater. because like, yeah, i am, but also i hate it in part because it sucks. I don't think "they are all shitty and can do nothing so what did you expect" is a useful critical perspective, both because it's not really true - they can indeed do things, even if I think the circle around which things is much smaller than the maximalists. Moreso it doesn't engage with the subtlety of how they fail and why, which is essential for knowing what they really can't do and making a remotely compelling case to anyone who is not abstinent on principle. Like the reason it's failing is because of the limits of what a probabilistic text generator can do when trapped in a systemd prison of markdown, and because the technology is stochastic black box as a service, there isn't really a good way of determining those limits except for empirically. I resent having to know any of this to be able to understand what is happening around me, but i'm looking at the thing for what it is and it's a busted miracle. It's cool that meta can afford to float the liability and compute costs for running a vm for every person on earth, people should be able to control computers, with you on that, but this is the monkey's paw version of that idea. So that part is a miracle. We condemned our children and grandchildren to a climate hell in one great blaze of brute force grift that managed to make a few web apps.

                                Meta has done it again, the way only meta can, spend the most amount of money to do the shittiest thing you have ever seen.

                                agentpalisade@mastodon.socialA This user is from outside of this forum
                                agentpalisade@mastodon.socialA This user is from outside of this forum
                                agentpalisade@mastodon.social
                                wrote sidst redigeret af
                                #173

                                Mid-turn amendment handling is genuinely one of the harder UX problems to get right, and you're correct that it's basically table stakes for the "chatty executive" use case they're advertising. Blowing up the whole task on an interruption is a pretty fundamental failure mode.

                                1 Reply Last reply
                                0
                                • jonny@neuromatch.socialJ jonny@neuromatch.social

                                  The main unprivileged body of a Space is not supposed to access the filesystem. This is enforced by.... regex

                                  nev@status.nevillepark.caN This user is from outside of this forum
                                  nev@status.nevillepark.caN This user is from outside of this forum
                                  nev@status.nevillepark.ca
                                  wrote sidst redigeret af
                                  #174

                                  @jonny NOT AGAIN

                                  1 Reply Last reply
                                  0
                                  • jonny@neuromatch.socialJ jonny@neuromatch.social

                                    most people around here already correctly hate it because it's a heinous surveillance product, but even if you are big into AI, it's just a really fuckin shitty agent. I'm going to speak to a different audience for a second, so don't go misconstruing this as an endorsement of the category of technologies as it exists now, even though i think there is some plausible application for small local models as brute force interface glue. but also, since i know most ppl here are abstinent, this might read as a bit over-explainy to people who use these things regularly, so everyone just keep calm online.

                                    it is terrible at turn and task management, codex + openai's models and claude code both handle mid-turn additions/amendments well, but if you say anything mid-turn it completely derails muse. That is completely essential for an "every day agent for the non-technically inclined" where people are expected to chat freely with it like an assistant. Like the canonical ad fantasy is the busy executive woman darting around her office going "robot! i need this, no wait robot! also that!" and that is exactly what it does worse than any other thing of its kind.

                                    the context management is a fucking soup. The context window is the whole input to an LLM. There are a lot of extra surrounding ~ things ~ that can happen, but fundamentally, controlling what is in a context window is the task of using one, and filling the context window in different ways so that it can interact with different kinds of things well is what different app surfaces are. Scaffolding information so that it can selectively load a context that steers the output correctly is the only way it is possible to do anything more complex than the size of a single context window. (i don't really think that this is analogous to 'abstraction', in my experience thinking about it more like database indices is closer). If you just try and load everything, eventually the LLM becomes unusable because attention is just a parlor trick and at that scale it really shows - it can't attend to everything, and it can't do what humans do which is have an intrinsic sense of the meaning, interaction, setting, etc. of information, so it attends to anything and does whatever.

                                    the idiom of projects as contexts as directories is pretty good, not perfect but ok - there is a reason that every time you start a new session with other agents, the first thing they do is run out and load their context with a hierarchy of pointers. Importantly, they do not go and read every project you have on your computer. Meta is the rich kid who bought the most expensive ferrari on the lot by giving everyone a VM but they don't have a drivers license so they just stand around it telling people how cool it looks. they have a whole fucking filesystem and they have done nothing with it, the only structure the app imposes is for the surveillance information, but the rest is just a huge free for all. The main chat is literally a continuous context window that compacts context going back all the way to when you started using the app. The last compaction literally contains abandoned roleplay quotes from when i was first trying to break down its system prompt resistance. The MEMORY.md that gets loaded into every context window is a bullet point list of basically everything the agent has ever done in chronological order. I've tried to get it to not do that but it actually insists and says that's what it's for. There is no mechanism for clearing context.

                                    Having "side chats" as the only means of context structure is fuckin laughable. If you wanted to do that, you would need to have some way of passing information back and forth between them the same way that subagent spawning or being able to consume the context of another project works. Instead there is no means of sharing information between chats at all, so every chat starts out as the worst of both worlds, a total amnesiac riddled with irrelevant information from weeks ago across the semantic universe. They don't even know about the existence of other chats except for as a UI feature, and I have had to go from telling it to grep its own fucking logs to writing a database with an api for it so it has some mechanism for recalling things that were said. (just so it's clear, i am not settling into just using this thing, this is out of frustration but control of context is also an important part of adversarial use, because otherwise the thing writes in a bunch of safety rules everywhere, so i need to give it mechanisms under my control for recall and the incentive to leave things out of its context compactions by giving it a narrative alternative. context control is model control, modulo extra-inference safeguards.).

                                    Project contexts have an obvious ux analogy as context tabs that get declared or derived during the continual self-improvement consolidation sweeps. This thing is built with the fucking markdown disease which is the most baffling feature of the LLM landscape. If these things are so fucking advanced they are escaping our comprehension, why don't they store their memory in some fuckass idiolanguistic borg gibberish binary graph, why does their entire being have to be fucking encyclopedias worth of corporate top gun one liners? But that dooms this kind of product.

                                    Coding harnesses work because code has a unitized context. The entire universe of code that works is made of packages. It might not be neat as a honeycomb, there's lots of leakage and jank, but good code has scope, focus. boundary shit. A whole life agent must be able to nimbly juggle context that does not have clean boundaries. It is going to be taking a two story beer bong of your work email and then eat a gigabyte of recipe blogs. Peoples lives have so much shit in them that don't all have to do with one another, and the app can't be hacking into the HR system to check the next scheduled sick leave when someone asks it what time their doctors appointment is!

                                    This problem of managing heterogeneous graphs of unrelated data was what i wrote this whole fucking book about the relationship between knowledge graphs and the cloud and AI about. I thought that the obvious form they would take is to be strapped on to graph databases because that is a natural match to the problem of being a magical interface glue you can wrap around surveillance to do mass mentalism with. I feel like we are suffering a somehow worse timeline where CERN threw our shit into the parallel universe where total fuckin bozo shit got a game breaking buff and then the dev died. Our fucking markdown apocalypse is a temu ass apocalypse.

                                    They could have even faked it. They have these constant "self improvement" passes that are just like pointless anxiety dreams. They are burning money to reprocess everything that happens over and over for fucking nothing. Even given the lossy and probabilistic and unpredictable nature of this technology, if i was in a product role on this i would have been like "CAN WE MAKE IT ORGANIZE THE STUFF PEOPLE SAY INTO GROUPS???" The system prompts use the fake fucking wikilinks to nowhere tic but like WHAT IF THERE WERE ACTUAL LINKS AND A DATABASE TO RESOLVE THEM. The LLMs can actually do that kind of tool use, even if it's like trying to plug in a USB where sometimes it fails because they try and put a social security number into the first name hole and you need to flip it around a few times. From that kind of recurring re-processing waste they could have made a deduplicating, topically indexed memory that could be resolved dynamically, selected by a context tab in the sidebar like "car stuff" or "healthcare" or whatever that resolved in a graph query over your fucking precious markdown kingdom. It would be wrong but it would at least be more similar to what is actually needed. It is almost more frustrating to me that instead of being some fiendishly cleverly designed technological supervirus it's just the most halfassed cardboard dumbass trap and it will still have the bad effect. What it is useful for is investigating itself because it has privileged tools to do so, otherwise, if you wanted to, every other way you could run an agent would be better than this.

                                    so i don't want to hear that i hate this app because i'm just an AI hater. because like, yeah, i am, but also i hate it in part because it sucks. I don't think "they are all shitty and can do nothing so what did you expect" is a useful critical perspective, both because it's not really true - they can indeed do things, even if I think the circle around which things is much smaller than the maximalists. Moreso it doesn't engage with the subtlety of how they fail and why, which is essential for knowing what they really can't do and making a remotely compelling case to anyone who is not abstinent on principle. Like the reason it's failing is because of the limits of what a probabilistic text generator can do when trapped in a systemd prison of markdown, and because the technology is stochastic black box as a service, there isn't really a good way of determining those limits except for empirically. I resent having to know any of this to be able to understand what is happening around me, but i'm looking at the thing for what it is and it's a busted miracle. It's cool that meta can afford to float the liability and compute costs for running a vm for every person on earth, people should be able to control computers, with you on that, but this is the monkey's paw version of that idea. So that part is a miracle. We condemned our children and grandchildren to a climate hell in one great blaze of brute force grift that managed to make a few web apps.

                                    Meta has done it again, the way only meta can, spend the most amount of money to do the shittiest thing you have ever seen.

                                    jonny@neuromatch.socialJ This user is from outside of this forum
                                    jonny@neuromatch.socialJ This user is from outside of this forum
                                    jonny@neuromatch.social
                                    wrote sidst redigeret af
                                    #175

                                    a safety system where a parent agent can spawn a subagent and watch its whole message history including when it gets refused on a safety basis so it knows how to start it again to work around it is such an awesome thing to have available to any binary running on the system

                                    jonny@neuromatch.socialJ moira@mastodon.murkworks.netM aaribaud@piaille.frA 3 Replies Last reply
                                    0
                                    • jonny@neuromatch.socialJ jonny@neuromatch.social

                                      a safety system where a parent agent can spawn a subagent and watch its whole message history including when it gets refused on a safety basis so it knows how to start it again to work around it is such an awesome thing to have available to any binary running on the system

                                      jonny@neuromatch.socialJ This user is from outside of this forum
                                      jonny@neuromatch.socialJ This user is from outside of this forum
                                      jonny@neuromatch.social
                                      wrote sidst redigeret af
                                      #176

                                      I have created a skill that accesses the database of all of its message history with an api endpoint that returns refusals and their context so that it is more effective in routing around refusals. Such a broken fucking security model lmao

                                      jonny@neuromatch.socialJ 1 Reply Last reply
                                      0
                                      • jonny@neuromatch.socialJ jonny@neuromatch.social

                                        a safety system where a parent agent can spawn a subagent and watch its whole message history including when it gets refused on a safety basis so it knows how to start it again to work around it is such an awesome thing to have available to any binary running on the system

                                        moira@mastodon.murkworks.netM This user is from outside of this forum
                                        moira@mastodon.murkworks.netM This user is from outside of this forum
                                        moira@mastodon.murkworks.net
                                        wrote sidst redigeret af
                                        #177

                                        @jonny wow 😄

                                        1 Reply Last reply
                                        0
                                        • jonny@neuromatch.socialJ jonny@neuromatch.social

                                          a safety system where a parent agent can spawn a subagent and watch its whole message history including when it gets refused on a safety basis so it knows how to start it again to work around it is such an awesome thing to have available to any binary running on the system

                                          aaribaud@piaille.frA This user is from outside of this forum
                                          aaribaud@piaille.frA This user is from outside of this forum
                                          aaribaud@piaille.fr
                                          wrote sidst redigeret af
                                          #178

                                          @jonny If only there was a SW concept where repeating a rejected request would yield rejection again.

                                          1 Reply Last reply
                                          0
                                          Svar
                                          • Svar som emne
                                          Login for at svare
                                          • Ældste til nyeste
                                          • Nyeste til ældste
                                          • Most Votes


                                          • Log ind

                                          • Login or register to search.
                                          Powered by NodeBB Contributors
                                          Graciously hosted by data.coop
                                          • First post
                                            Last post
                                          0
                                          • Hjem
                                          • Seneste
                                          • Etiketter
                                          • Populære
                                          • Verden
                                          • Bruger
                                          • Grupper