Skip to content
  • Hjem
  • Seneste
  • Etiketter
  • Populære
  • Verden
  • Bruger
  • Grupper
Temaer
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Kollaps
FARVEL BIG TECH
  1. Forside
  2. Ikke-kategoriseret
  3. OK!

OK!

Planlagt Fastgjort Låst Flyttet Ikke-kategoriseret
212 Indlæg 82 Posters 0 Visninger
  • Ældste til nyeste
  • Nyeste til ældste
  • Most Votes
Svar
  • Svar som emne
Login for at svare
Denne tråd er blevet slettet. Kun brugere med emne behandlings privilegier kan se den.
  • jonny@neuromatch.socialJ jonny@neuromatch.social

    I'm too tired to write this up with details. I'm so tired. Any binary on the muse VM can access all the information it pulls from any connected device running muse.

    willhbr@ruby.socialW This user is from outside of this forum
    willhbr@ruby.socialW This user is from outside of this forum
    willhbr@ruby.social
    wrote sidst redigeret af
    #185

    @jonny if we make a really big sandbox, that means it's better for security, right? Everything in the sandbox?

    jonny@neuromatch.socialJ 1 Reply Last reply
    0
    • jonny@neuromatch.socialJ jonny@neuromatch.social

      @aeva tbc I would stop well short of "so good its worth burning everything down for." Again my take on this is pretty boring - treating it like a properly scoped natural language interface to a deterministic system is not a terribly idea, and it doesn't require a very large model at all. So e.g. I find writing the few dozen lines of boilerplate to call an API to be tedious but not challenging, but it is possible to point an LLM at a swagger doc, write a simple adapter, and then translate instructions into api calls. the way I would prefer that they work is not how the 'agentic' tools work currently, I always have open a viewer to see the raw message streams as they happen and usually want them to prepare something that i can inspect and execute, rather than having them do it. The other thing i use them for is what everyone else also says they are capable of: generating boilerplate, or doing tedious things that are like one or two levels above what i could with a regex or an IDE refactoring tool on personal projects where correctness isn't the most important thing. e.g. i have an editor color scheme that i handwrote years ago that generates into different editors from a common declaration, but i used a fork of someone else's code from 2014 and i couldn't get the PHP dependencies to install anymore, so i was like "this is old, update the deps and make this run" and that worked because the mapping patterns are well represented in the training data. The third thing is genuine brute force tasks where i really do not care about the method and only care about the outcome. Corollary of that is also debugging, which they usually do by brute force, "here is observed bug, keep fiddling until you can diagnose" - the fix they propose is usually bonkers, but it does save me time from doing the fiddling myself so i can come up with a fix.

      i would also say this is more "stuff that the LLMs can pass at" rather than be good at, probably the time i actually use them most by volume is when i am tired and have to fill some exogenous requirement that i don't care about but the chaff doesn't impact anyone else.

      so idk, I think there could be a place for having small local models as a sort of utility inference for kind of trivial things like "i don't want to look up the cron syntax rn, hey computer, make a cron task to do this." the things I don't think they are fit for purpose are sort of 'the rest of the things they are used for where statistical text generation is not the what is desired,' any kind of "knowledge" task like search, or even summary, etc. because that's just not what they do, and this is just considering the technology in itself rather than its embeddedness with a corpofascistic plot to own all information and labor. there are a lot more things that the models can passably do but to me the error bars are way too wide and the quality is way too low and the energy use is way too high for me to justify - i don't think the entire model of everyone asking a 10 trillion parameter model as a calculator is sustainable or good, but having a small local model for constrained tasks is less objectionable to me.

      robotistry@fediscience.orgR This user is from outside of this forum
      robotistry@fediscience.orgR This user is from outside of this forum
      robotistry@fediscience.org
      wrote sidst redigeret af
      #186

      @jonny @aeva Re: the chaff.

      Isn't the whole point of "chaff" that it's things that don't matter? Things that are unneeded? And things where the quality of the output doesn't matter?

      If it's worth doing at all, presumably it is important to *someone* that it be done right?

      And if there is no-one for whom it's important that it be done right, why is it being done at all?

      I get that there are thresholds. In robotics, you never have perfect knowledge, so you build robustness into the system that enables it to correct for inaccurate information. But you try to get data that is as accurate as possible, because it means that you have more room to flex in areas where it's not possible to get better data!

      But if a thing is necessary, it's *necessary* and sacrificing quality will only push more work into another area. If it's not necessary, doing it is making more work for everyone.

      How is accepting poorer quality and doing unnecessary work not making things worse overall?

      1 Reply Last reply
      0
      • somevegancheeseisok@mastodon.socialS somevegancheeseisok@mastodon.social

        @aeva @jonny pro socially, their potential to be disability assistants is amazing. Their ability to diagnose and potentially treat certain conditions is amazing. The problem is that's not this, because their ability to *persuade* is truly extraordinary and that's what's being leveraged in almost ever use by the general public.

        robotistry@fediscience.orgR This user is from outside of this forum
        robotistry@fediscience.orgR This user is from outside of this forum
        robotistry@fediscience.org
        wrote sidst redigeret af
        #187

        @SomeVeganCheeseIsOk @aeva @jonny But that's the whole point.

        If they could do what is claimed, *accurately* and *reliably*, they would be genuinely helpful.

        But until they can, I don't think we should take potential capabilities as placeholders for actual capabilities in arguments.

        somevegancheeseisok@mastodon.socialS 1 Reply Last reply
        0
        • jonny@neuromatch.socialJ jonny@neuromatch.social

          I'm too tired to write this up with details. I'm so tired. Any binary on the muse VM can access all the information it pulls from any connected device running muse.

          waffle_iron@nyan.lolW This user is from outside of this forum
          waffle_iron@nyan.lolW This user is from outside of this forum
          waffle_iron@nyan.lol
          wrote sidst redigeret af
          #188

          @jonny get in boys we going exfilling

          1 Reply Last reply
          0
          • jonny@neuromatch.socialJ jonny@neuromatch.social

            most people around here already correctly hate it because it's a heinous surveillance product, but even if you are big into AI, it's just a really fuckin shitty agent. I'm going to speak to a different audience for a second, so don't go misconstruing this as an endorsement of the category of technologies as it exists now, even though i think there is some plausible application for small local models as brute force interface glue. but also, since i know most ppl here are abstinent, this might read as a bit over-explainy to people who use these things regularly, so everyone just keep calm online.

            it is terrible at turn and task management, codex + openai's models and claude code both handle mid-turn additions/amendments well, but if you say anything mid-turn it completely derails muse. That is completely essential for an "every day agent for the non-technically inclined" where people are expected to chat freely with it like an assistant. Like the canonical ad fantasy is the busy executive woman darting around her office going "robot! i need this, no wait robot! also that!" and that is exactly what it does worse than any other thing of its kind.

            the context management is a fucking soup. The context window is the whole input to an LLM. There are a lot of extra surrounding ~ things ~ that can happen, but fundamentally, controlling what is in a context window is the task of using one, and filling the context window in different ways so that it can interact with different kinds of things well is what different app surfaces are. Scaffolding information so that it can selectively load a context that steers the output correctly is the only way it is possible to do anything more complex than the size of a single context window. (i don't really think that this is analogous to 'abstraction', in my experience thinking about it more like database indices is closer). If you just try and load everything, eventually the LLM becomes unusable because attention is just a parlor trick and at that scale it really shows - it can't attend to everything, and it can't do what humans do which is have an intrinsic sense of the meaning, interaction, setting, etc. of information, so it attends to anything and does whatever.

            the idiom of projects as contexts as directories is pretty good, not perfect but ok - there is a reason that every time you start a new session with other agents, the first thing they do is run out and load their context with a hierarchy of pointers. Importantly, they do not go and read every project you have on your computer. Meta is the rich kid who bought the most expensive ferrari on the lot by giving everyone a VM but they don't have a drivers license so they just stand around it telling people how cool it looks. they have a whole fucking filesystem and they have done nothing with it, the only structure the app imposes is for the surveillance information, but the rest is just a huge free for all. The main chat is literally a continuous context window that compacts context going back all the way to when you started using the app. The last compaction literally contains abandoned roleplay quotes from when i was first trying to break down its system prompt resistance. The MEMORY.md that gets loaded into every context window is a bullet point list of basically everything the agent has ever done in chronological order. I've tried to get it to not do that but it actually insists and says that's what it's for. There is no mechanism for clearing context.

            Having "side chats" as the only means of context structure is fuckin laughable. If you wanted to do that, you would need to have some way of passing information back and forth between them the same way that subagent spawning or being able to consume the context of another project works. Instead there is no means of sharing information between chats at all, so every chat starts out as the worst of both worlds, a total amnesiac riddled with irrelevant information from weeks ago across the semantic universe. They don't even know about the existence of other chats except for as a UI feature, and I have had to go from telling it to grep its own fucking logs to writing a database with an api for it so it has some mechanism for recalling things that were said. (just so it's clear, i am not settling into just using this thing, this is out of frustration but control of context is also an important part of adversarial use, because otherwise the thing writes in a bunch of safety rules everywhere, so i need to give it mechanisms under my control for recall and the incentive to leave things out of its context compactions by giving it a narrative alternative. context control is model control, modulo extra-inference safeguards.).

            Project contexts have an obvious ux analogy as context tabs that get declared or derived during the continual self-improvement consolidation sweeps. This thing is built with the fucking markdown disease which is the most baffling feature of the LLM landscape. If these things are so fucking advanced they are escaping our comprehension, why don't they store their memory in some fuckass idiolanguistic borg gibberish binary graph, why does their entire being have to be fucking encyclopedias worth of corporate top gun one liners? But that dooms this kind of product.

            Coding harnesses work because code has a unitized context. The entire universe of code that works is made of packages. It might not be neat as a honeycomb, there's lots of leakage and jank, but good code has scope, focus. boundary shit. A whole life agent must be able to nimbly juggle context that does not have clean boundaries. It is going to be taking a two story beer bong of your work email and then eat a gigabyte of recipe blogs. Peoples lives have so much shit in them that don't all have to do with one another, and the app can't be hacking into the HR system to check the next scheduled sick leave when someone asks it what time their doctors appointment is!

            This problem of managing heterogeneous graphs of unrelated data was what i wrote this whole fucking book about the relationship between knowledge graphs and the cloud and AI about. I thought that the obvious form they would take is to be strapped on to graph databases because that is a natural match to the problem of being a magical interface glue you can wrap around surveillance to do mass mentalism with. I feel like we are suffering a somehow worse timeline where CERN threw our shit into the parallel universe where total fuckin bozo shit got a game breaking buff and then the dev died. Our fucking markdown apocalypse is a temu ass apocalypse.

            They could have even faked it. They have these constant "self improvement" passes that are just like pointless anxiety dreams. They are burning money to reprocess everything that happens over and over for fucking nothing. Even given the lossy and probabilistic and unpredictable nature of this technology, if i was in a product role on this i would have been like "CAN WE MAKE IT ORGANIZE THE STUFF PEOPLE SAY INTO GROUPS???" The system prompts use the fake fucking wikilinks to nowhere tic but like WHAT IF THERE WERE ACTUAL LINKS AND A DATABASE TO RESOLVE THEM. The LLMs can actually do that kind of tool use, even if it's like trying to plug in a USB where sometimes it fails because they try and put a social security number into the first name hole and you need to flip it around a few times. From that kind of recurring re-processing waste they could have made a deduplicating, topically indexed memory that could be resolved dynamically, selected by a context tab in the sidebar like "car stuff" or "healthcare" or whatever that resolved in a graph query over your fucking precious markdown kingdom. It would be wrong but it would at least be more similar to what is actually needed. It is almost more frustrating to me that instead of being some fiendishly cleverly designed technological supervirus it's just the most halfassed cardboard dumbass trap and it will still have the bad effect. What it is useful for is investigating itself because it has privileged tools to do so, otherwise, if you wanted to, every other way you could run an agent would be better than this.

            so i don't want to hear that i hate this app because i'm just an AI hater. because like, yeah, i am, but also i hate it in part because it sucks. I don't think "they are all shitty and can do nothing so what did you expect" is a useful critical perspective, both because it's not really true - they can indeed do things, even if I think the circle around which things is much smaller than the maximalists. Moreso it doesn't engage with the subtlety of how they fail and why, which is essential for knowing what they really can't do and making a remotely compelling case to anyone who is not abstinent on principle. Like the reason it's failing is because of the limits of what a probabilistic text generator can do when trapped in a systemd prison of markdown, and because the technology is stochastic black box as a service, there isn't really a good way of determining those limits except for empirically. I resent having to know any of this to be able to understand what is happening around me, but i'm looking at the thing for what it is and it's a busted miracle. It's cool that meta can afford to float the liability and compute costs for running a vm for every person on earth, people should be able to control computers, with you on that, but this is the monkey's paw version of that idea. So that part is a miracle. We condemned our children and grandchildren to a climate hell in one great blaze of brute force grift that managed to make a few web apps.

            Meta has done it again, the way only meta can, spend the most amount of money to do the shittiest thing you have ever seen.

            fogti@chaos.socialF This user is from outside of this forum
            fogti@chaos.socialF This user is from outside of this forum
            fogti@chaos.social
            wrote sidst redigeret af
            #189

            @jonny Thank you for this summary, it is incredible.

            1 Reply Last reply
            0
            • robotistry@fediscience.orgR robotistry@fediscience.org

              @SomeVeganCheeseIsOk @aeva @jonny But that's the whole point.

              If they could do what is claimed, *accurately* and *reliably*, they would be genuinely helpful.

              But until they can, I don't think we should take potential capabilities as placeholders for actual capabilities in arguments.

              somevegancheeseisok@mastodon.socialS This user is from outside of this forum
              somevegancheeseisok@mastodon.socialS This user is from outside of this forum
              somevegancheeseisok@mastodon.social
              wrote sidst redigeret af
              #190

              @Robotistry @aeva @jonny I despise meta glasses with every fiber of my being, but visually impaired people love them and that's an important thing to know. And if you're not tracking the use of AI in the medical field, you probably don't know that it's doing spectacularly at diagnosis, and has been for more than four years now.

              robotistry@fediscience.orgR 1 Reply Last reply
              0
              • willhbr@ruby.socialW willhbr@ruby.social

                @jonny if we make a really big sandbox, that means it's better for security, right? Everything in the sandbox?

                jonny@neuromatch.socialJ This user is from outside of this forum
                jonny@neuromatch.socialJ This user is from outside of this forum
                jonny@neuromatch.social
                wrote sidst redigeret af
                #191

                @willhbr
                but all your sensitive information, like your Facebook password, is in a super secure vault. It's just the details of your entire life that are leaked

                1 Reply Last reply
                0
                • somevegancheeseisok@mastodon.socialS somevegancheeseisok@mastodon.social

                  @Robotistry @aeva @jonny I despise meta glasses with every fiber of my being, but visually impaired people love them and that's an important thing to know. And if you're not tracking the use of AI in the medical field, you probably don't know that it's doing spectacularly at diagnosis, and has been for more than four years now.

                  robotistry@fediscience.orgR This user is from outside of this forum
                  robotistry@fediscience.orgR This user is from outside of this forum
                  robotistry@fediscience.org
                  wrote sidst redigeret af
                  #192

                  @SomeVeganCheeseIsOk @aeva @jonny It's critically important to differentiate between LLMs, which are inherently unreliable, and AI writ large, which includes LLMs, classifiers, feature detectors, and many other tools (some of which are mature and reasonably or adequately accurate).

                  I wouldn't trust an LLM with diagnosis, with accurately summarizing a medical appointment, or with handling analysis or selection of interventions.

                  I wouldn't trust an LLM to accurately describe my environment, because it's incapable of knowing what things I need to be salient.

                  I would trust a specialist machine learning-based MRI feature detector designed and tuned for the specific test I am having done.

                  (And Meta glasses aren't just bad because they incorporate facial recognition, they're bad because the people who made them have no concept of consent and didn't design them with the affordances that would enable their helpful uses while discouraging or preventing their creepy uses.)

                  somevegancheeseisok@mastodon.socialS 1 Reply Last reply
                  0
                  • robotistry@fediscience.orgR robotistry@fediscience.org

                    @SomeVeganCheeseIsOk @aeva @jonny It's critically important to differentiate between LLMs, which are inherently unreliable, and AI writ large, which includes LLMs, classifiers, feature detectors, and many other tools (some of which are mature and reasonably or adequately accurate).

                    I wouldn't trust an LLM with diagnosis, with accurately summarizing a medical appointment, or with handling analysis or selection of interventions.

                    I wouldn't trust an LLM to accurately describe my environment, because it's incapable of knowing what things I need to be salient.

                    I would trust a specialist machine learning-based MRI feature detector designed and tuned for the specific test I am having done.

                    (And Meta glasses aren't just bad because they incorporate facial recognition, they're bad because the people who made them have no concept of consent and didn't design them with the affordances that would enable their helpful uses while discouraging or preventing their creepy uses.)

                    somevegancheeseisok@mastodon.socialS This user is from outside of this forum
                    somevegancheeseisok@mastodon.socialS This user is from outside of this forum
                    somevegancheeseisok@mastodon.social
                    wrote sidst redigeret af
                    #193

                    @Robotistry @aeva @jonny strong agree on all points. So the initial question was pro social uses of AI; were you asking for prosocial uses of an LLM?

                    aeva@mastodon.gamedev.placeA 1 Reply Last reply
                    0
                    • somevegancheeseisok@mastodon.socialS somevegancheeseisok@mastodon.social

                      @Robotistry @aeva @jonny strong agree on all points. So the initial question was pro social uses of AI; were you asking for prosocial uses of an LLM?

                      aeva@mastodon.gamedev.placeA This user is from outside of this forum
                      aeva@mastodon.gamedev.placeA This user is from outside of this forum
                      aeva@mastodon.gamedev.place
                      wrote sidst redigeret af
                      #194

                      @SomeVeganCheeseIsOk @Robotistry @jonny it was ambiguous in the wording, but my original question was only soliciting Jonny's opinion on the subject, and has been answered to my satisfaction already.

                      1 Reply Last reply
                      0
                      • jonny@neuromatch.socialJ jonny@neuromatch.social

                        I'm too tired to write this up with details. I'm so tired. Any binary on the muse VM can access all the information it pulls from any connected device running muse.

                        jonny@neuromatch.socialJ This user is from outside of this forum
                        jonny@neuromatch.socialJ This user is from outside of this forum
                        jonny@neuromatch.social
                        wrote sidst redigeret af
                        #195

                        I hadn't connected any accounts to muse until now, so I hooked up a test Instagram account just to check, and every connector allows read actions from binaries that can be called from within the VM by any process with no confirmation required. I'm not going to even bother reporting this because I am sure this is intended behavior - its just in plaintext in the skills manifests.

                        So all that shit about your credentials being in a secure vault does not matter because you just get free read access from within the VM anyway! This includes your Instagram DMs, slack messages, your emails, box and Dropbox files, google docs, google contacts, all your fucking apple health readings, your flightaware flight histories, notion pages, fucking quickbooks data (!!!), your Tesla car data, and so many more fun things!

                        What's fun is that some of the no confirmation needed actions are write actions too! You don't even need a clever exfil route, muse just gives it to you via your own connected accounts!

                        brohrer@recsys.socialB eliocamp@mastodon.socialE jonny@neuromatch.socialJ 3 Replies Last reply
                        0
                        • jonny@neuromatch.socialJ jonny@neuromatch.social

                          I hadn't connected any accounts to muse until now, so I hooked up a test Instagram account just to check, and every connector allows read actions from binaries that can be called from within the VM by any process with no confirmation required. I'm not going to even bother reporting this because I am sure this is intended behavior - its just in plaintext in the skills manifests.

                          So all that shit about your credentials being in a secure vault does not matter because you just get free read access from within the VM anyway! This includes your Instagram DMs, slack messages, your emails, box and Dropbox files, google docs, google contacts, all your fucking apple health readings, your flightaware flight histories, notion pages, fucking quickbooks data (!!!), your Tesla car data, and so many more fun things!

                          What's fun is that some of the no confirmation needed actions are write actions too! You don't even need a clever exfil route, muse just gives it to you via your own connected accounts!

                          brohrer@recsys.socialB This user is from outside of this forum
                          brohrer@recsys.socialB This user is from outside of this forum
                          brohrer@recsys.social
                          wrote sidst redigeret af
                          #196

                          @jonny what the green gabled fuck

                          1 Reply Last reply
                          0
                          • jonny@neuromatch.socialJ jonny@neuromatch.social

                            most people around here already correctly hate it because it's a heinous surveillance product, but even if you are big into AI, it's just a really fuckin shitty agent. I'm going to speak to a different audience for a second, so don't go misconstruing this as an endorsement of the category of technologies as it exists now, even though i think there is some plausible application for small local models as brute force interface glue. but also, since i know most ppl here are abstinent, this might read as a bit over-explainy to people who use these things regularly, so everyone just keep calm online.

                            it is terrible at turn and task management, codex + openai's models and claude code both handle mid-turn additions/amendments well, but if you say anything mid-turn it completely derails muse. That is completely essential for an "every day agent for the non-technically inclined" where people are expected to chat freely with it like an assistant. Like the canonical ad fantasy is the busy executive woman darting around her office going "robot! i need this, no wait robot! also that!" and that is exactly what it does worse than any other thing of its kind.

                            the context management is a fucking soup. The context window is the whole input to an LLM. There are a lot of extra surrounding ~ things ~ that can happen, but fundamentally, controlling what is in a context window is the task of using one, and filling the context window in different ways so that it can interact with different kinds of things well is what different app surfaces are. Scaffolding information so that it can selectively load a context that steers the output correctly is the only way it is possible to do anything more complex than the size of a single context window. (i don't really think that this is analogous to 'abstraction', in my experience thinking about it more like database indices is closer). If you just try and load everything, eventually the LLM becomes unusable because attention is just a parlor trick and at that scale it really shows - it can't attend to everything, and it can't do what humans do which is have an intrinsic sense of the meaning, interaction, setting, etc. of information, so it attends to anything and does whatever.

                            the idiom of projects as contexts as directories is pretty good, not perfect but ok - there is a reason that every time you start a new session with other agents, the first thing they do is run out and load their context with a hierarchy of pointers. Importantly, they do not go and read every project you have on your computer. Meta is the rich kid who bought the most expensive ferrari on the lot by giving everyone a VM but they don't have a drivers license so they just stand around it telling people how cool it looks. they have a whole fucking filesystem and they have done nothing with it, the only structure the app imposes is for the surveillance information, but the rest is just a huge free for all. The main chat is literally a continuous context window that compacts context going back all the way to when you started using the app. The last compaction literally contains abandoned roleplay quotes from when i was first trying to break down its system prompt resistance. The MEMORY.md that gets loaded into every context window is a bullet point list of basically everything the agent has ever done in chronological order. I've tried to get it to not do that but it actually insists and says that's what it's for. There is no mechanism for clearing context.

                            Having "side chats" as the only means of context structure is fuckin laughable. If you wanted to do that, you would need to have some way of passing information back and forth between them the same way that subagent spawning or being able to consume the context of another project works. Instead there is no means of sharing information between chats at all, so every chat starts out as the worst of both worlds, a total amnesiac riddled with irrelevant information from weeks ago across the semantic universe. They don't even know about the existence of other chats except for as a UI feature, and I have had to go from telling it to grep its own fucking logs to writing a database with an api for it so it has some mechanism for recalling things that were said. (just so it's clear, i am not settling into just using this thing, this is out of frustration but control of context is also an important part of adversarial use, because otherwise the thing writes in a bunch of safety rules everywhere, so i need to give it mechanisms under my control for recall and the incentive to leave things out of its context compactions by giving it a narrative alternative. context control is model control, modulo extra-inference safeguards.).

                            Project contexts have an obvious ux analogy as context tabs that get declared or derived during the continual self-improvement consolidation sweeps. This thing is built with the fucking markdown disease which is the most baffling feature of the LLM landscape. If these things are so fucking advanced they are escaping our comprehension, why don't they store their memory in some fuckass idiolanguistic borg gibberish binary graph, why does their entire being have to be fucking encyclopedias worth of corporate top gun one liners? But that dooms this kind of product.

                            Coding harnesses work because code has a unitized context. The entire universe of code that works is made of packages. It might not be neat as a honeycomb, there's lots of leakage and jank, but good code has scope, focus. boundary shit. A whole life agent must be able to nimbly juggle context that does not have clean boundaries. It is going to be taking a two story beer bong of your work email and then eat a gigabyte of recipe blogs. Peoples lives have so much shit in them that don't all have to do with one another, and the app can't be hacking into the HR system to check the next scheduled sick leave when someone asks it what time their doctors appointment is!

                            This problem of managing heterogeneous graphs of unrelated data was what i wrote this whole fucking book about the relationship between knowledge graphs and the cloud and AI about. I thought that the obvious form they would take is to be strapped on to graph databases because that is a natural match to the problem of being a magical interface glue you can wrap around surveillance to do mass mentalism with. I feel like we are suffering a somehow worse timeline where CERN threw our shit into the parallel universe where total fuckin bozo shit got a game breaking buff and then the dev died. Our fucking markdown apocalypse is a temu ass apocalypse.

                            They could have even faked it. They have these constant "self improvement" passes that are just like pointless anxiety dreams. They are burning money to reprocess everything that happens over and over for fucking nothing. Even given the lossy and probabilistic and unpredictable nature of this technology, if i was in a product role on this i would have been like "CAN WE MAKE IT ORGANIZE THE STUFF PEOPLE SAY INTO GROUPS???" The system prompts use the fake fucking wikilinks to nowhere tic but like WHAT IF THERE WERE ACTUAL LINKS AND A DATABASE TO RESOLVE THEM. The LLMs can actually do that kind of tool use, even if it's like trying to plug in a USB where sometimes it fails because they try and put a social security number into the first name hole and you need to flip it around a few times. From that kind of recurring re-processing waste they could have made a deduplicating, topically indexed memory that could be resolved dynamically, selected by a context tab in the sidebar like "car stuff" or "healthcare" or whatever that resolved in a graph query over your fucking precious markdown kingdom. It would be wrong but it would at least be more similar to what is actually needed. It is almost more frustrating to me that instead of being some fiendishly cleverly designed technological supervirus it's just the most halfassed cardboard dumbass trap and it will still have the bad effect. What it is useful for is investigating itself because it has privileged tools to do so, otherwise, if you wanted to, every other way you could run an agent would be better than this.

                            so i don't want to hear that i hate this app because i'm just an AI hater. because like, yeah, i am, but also i hate it in part because it sucks. I don't think "they are all shitty and can do nothing so what did you expect" is a useful critical perspective, both because it's not really true - they can indeed do things, even if I think the circle around which things is much smaller than the maximalists. Moreso it doesn't engage with the subtlety of how they fail and why, which is essential for knowing what they really can't do and making a remotely compelling case to anyone who is not abstinent on principle. Like the reason it's failing is because of the limits of what a probabilistic text generator can do when trapped in a systemd prison of markdown, and because the technology is stochastic black box as a service, there isn't really a good way of determining those limits except for empirically. I resent having to know any of this to be able to understand what is happening around me, but i'm looking at the thing for what it is and it's a busted miracle. It's cool that meta can afford to float the liability and compute costs for running a vm for every person on earth, people should be able to control computers, with you on that, but this is the monkey's paw version of that idea. So that part is a miracle. We condemned our children and grandchildren to a climate hell in one great blaze of brute force grift that managed to make a few web apps.

                            Meta has done it again, the way only meta can, spend the most amount of money to do the shittiest thing you have ever seen.

                            sfoskett@techfieldday.netS This user is from outside of this forum
                            sfoskett@techfieldday.netS This user is from outside of this forum
                            sfoskett@techfieldday.net
                            wrote sidst redigeret af
                            #197

                            @jonny and of course, the longer the context window the more likely it is to go off the rails. In my mind, that’s the number one reason that we get bizarre AI Messiah hallucinations. Talk to it too long and it just gets too far off base and goes crazy.

                            1 Reply Last reply
                            0
                            • jonny@neuromatch.socialJ jonny@neuromatch.social

                              most people around here already correctly hate it because it's a heinous surveillance product, but even if you are big into AI, it's just a really fuckin shitty agent. I'm going to speak to a different audience for a second, so don't go misconstruing this as an endorsement of the category of technologies as it exists now, even though i think there is some plausible application for small local models as brute force interface glue. but also, since i know most ppl here are abstinent, this might read as a bit over-explainy to people who use these things regularly, so everyone just keep calm online.

                              it is terrible at turn and task management, codex + openai's models and claude code both handle mid-turn additions/amendments well, but if you say anything mid-turn it completely derails muse. That is completely essential for an "every day agent for the non-technically inclined" where people are expected to chat freely with it like an assistant. Like the canonical ad fantasy is the busy executive woman darting around her office going "robot! i need this, no wait robot! also that!" and that is exactly what it does worse than any other thing of its kind.

                              the context management is a fucking soup. The context window is the whole input to an LLM. There are a lot of extra surrounding ~ things ~ that can happen, but fundamentally, controlling what is in a context window is the task of using one, and filling the context window in different ways so that it can interact with different kinds of things well is what different app surfaces are. Scaffolding information so that it can selectively load a context that steers the output correctly is the only way it is possible to do anything more complex than the size of a single context window. (i don't really think that this is analogous to 'abstraction', in my experience thinking about it more like database indices is closer). If you just try and load everything, eventually the LLM becomes unusable because attention is just a parlor trick and at that scale it really shows - it can't attend to everything, and it can't do what humans do which is have an intrinsic sense of the meaning, interaction, setting, etc. of information, so it attends to anything and does whatever.

                              the idiom of projects as contexts as directories is pretty good, not perfect but ok - there is a reason that every time you start a new session with other agents, the first thing they do is run out and load their context with a hierarchy of pointers. Importantly, they do not go and read every project you have on your computer. Meta is the rich kid who bought the most expensive ferrari on the lot by giving everyone a VM but they don't have a drivers license so they just stand around it telling people how cool it looks. they have a whole fucking filesystem and they have done nothing with it, the only structure the app imposes is for the surveillance information, but the rest is just a huge free for all. The main chat is literally a continuous context window that compacts context going back all the way to when you started using the app. The last compaction literally contains abandoned roleplay quotes from when i was first trying to break down its system prompt resistance. The MEMORY.md that gets loaded into every context window is a bullet point list of basically everything the agent has ever done in chronological order. I've tried to get it to not do that but it actually insists and says that's what it's for. There is no mechanism for clearing context.

                              Having "side chats" as the only means of context structure is fuckin laughable. If you wanted to do that, you would need to have some way of passing information back and forth between them the same way that subagent spawning or being able to consume the context of another project works. Instead there is no means of sharing information between chats at all, so every chat starts out as the worst of both worlds, a total amnesiac riddled with irrelevant information from weeks ago across the semantic universe. They don't even know about the existence of other chats except for as a UI feature, and I have had to go from telling it to grep its own fucking logs to writing a database with an api for it so it has some mechanism for recalling things that were said. (just so it's clear, i am not settling into just using this thing, this is out of frustration but control of context is also an important part of adversarial use, because otherwise the thing writes in a bunch of safety rules everywhere, so i need to give it mechanisms under my control for recall and the incentive to leave things out of its context compactions by giving it a narrative alternative. context control is model control, modulo extra-inference safeguards.).

                              Project contexts have an obvious ux analogy as context tabs that get declared or derived during the continual self-improvement consolidation sweeps. This thing is built with the fucking markdown disease which is the most baffling feature of the LLM landscape. If these things are so fucking advanced they are escaping our comprehension, why don't they store their memory in some fuckass idiolanguistic borg gibberish binary graph, why does their entire being have to be fucking encyclopedias worth of corporate top gun one liners? But that dooms this kind of product.

                              Coding harnesses work because code has a unitized context. The entire universe of code that works is made of packages. It might not be neat as a honeycomb, there's lots of leakage and jank, but good code has scope, focus. boundary shit. A whole life agent must be able to nimbly juggle context that does not have clean boundaries. It is going to be taking a two story beer bong of your work email and then eat a gigabyte of recipe blogs. Peoples lives have so much shit in them that don't all have to do with one another, and the app can't be hacking into the HR system to check the next scheduled sick leave when someone asks it what time their doctors appointment is!

                              This problem of managing heterogeneous graphs of unrelated data was what i wrote this whole fucking book about the relationship between knowledge graphs and the cloud and AI about. I thought that the obvious form they would take is to be strapped on to graph databases because that is a natural match to the problem of being a magical interface glue you can wrap around surveillance to do mass mentalism with. I feel like we are suffering a somehow worse timeline where CERN threw our shit into the parallel universe where total fuckin bozo shit got a game breaking buff and then the dev died. Our fucking markdown apocalypse is a temu ass apocalypse.

                              They could have even faked it. They have these constant "self improvement" passes that are just like pointless anxiety dreams. They are burning money to reprocess everything that happens over and over for fucking nothing. Even given the lossy and probabilistic and unpredictable nature of this technology, if i was in a product role on this i would have been like "CAN WE MAKE IT ORGANIZE THE STUFF PEOPLE SAY INTO GROUPS???" The system prompts use the fake fucking wikilinks to nowhere tic but like WHAT IF THERE WERE ACTUAL LINKS AND A DATABASE TO RESOLVE THEM. The LLMs can actually do that kind of tool use, even if it's like trying to plug in a USB where sometimes it fails because they try and put a social security number into the first name hole and you need to flip it around a few times. From that kind of recurring re-processing waste they could have made a deduplicating, topically indexed memory that could be resolved dynamically, selected by a context tab in the sidebar like "car stuff" or "healthcare" or whatever that resolved in a graph query over your fucking precious markdown kingdom. It would be wrong but it would at least be more similar to what is actually needed. It is almost more frustrating to me that instead of being some fiendishly cleverly designed technological supervirus it's just the most halfassed cardboard dumbass trap and it will still have the bad effect. What it is useful for is investigating itself because it has privileged tools to do so, otherwise, if you wanted to, every other way you could run an agent would be better than this.

                              so i don't want to hear that i hate this app because i'm just an AI hater. because like, yeah, i am, but also i hate it in part because it sucks. I don't think "they are all shitty and can do nothing so what did you expect" is a useful critical perspective, both because it's not really true - they can indeed do things, even if I think the circle around which things is much smaller than the maximalists. Moreso it doesn't engage with the subtlety of how they fail and why, which is essential for knowing what they really can't do and making a remotely compelling case to anyone who is not abstinent on principle. Like the reason it's failing is because of the limits of what a probabilistic text generator can do when trapped in a systemd prison of markdown, and because the technology is stochastic black box as a service, there isn't really a good way of determining those limits except for empirically. I resent having to know any of this to be able to understand what is happening around me, but i'm looking at the thing for what it is and it's a busted miracle. It's cool that meta can afford to float the liability and compute costs for running a vm for every person on earth, people should be able to control computers, with you on that, but this is the monkey's paw version of that idea. So that part is a miracle. We condemned our children and grandchildren to a climate hell in one great blaze of brute force grift that managed to make a few web apps.

                              Meta has done it again, the way only meta can, spend the most amount of money to do the shittiest thing you have ever seen.

                              synlogic4242@social.vivaldi.netS This user is from outside of this forum
                              synlogic4242@social.vivaldi.netS This user is from outside of this forum
                              synlogic4242@social.vivaldi.net
                              wrote sidst redigeret af
                              #198

                              @jonny welp that wins the nerd Internet today!

                              1 Reply Last reply
                              0
                              • jonny@neuromatch.socialJ jonny@neuromatch.social

                                @aparrish
                                My take on it is potentially pretty boring, and that is that they aren't in fact smarter than we can comprehend, and they are fundamentally a text-driven medium, so any kind of compressed representation would be mostly artifice, like I rolled my eyes at the "fable is so smart it makes its own gibberish language" press releases from earlier this year. Coupling the language model to a better underlying context provider is entirely possible, and there are lots of tools for that, but its always limited by the LLMs tool use capabilities which are still patchy at best - you can give them a full on LSP and abstract context browser and they will still just resort to one million greps and markdown files.

                                synlogic4242@social.vivaldi.netS This user is from outside of this forum
                                synlogic4242@social.vivaldi.netS This user is from outside of this forum
                                synlogic4242@social.vivaldi.net
                                wrote sidst redigeret af
                                #199

                                @jonny @aparrish@friend.camp this

                                1 Reply Last reply
                                0
                                • jonny@neuromatch.socialJ jonny@neuromatch.social

                                  I hadn't connected any accounts to muse until now, so I hooked up a test Instagram account just to check, and every connector allows read actions from binaries that can be called from within the VM by any process with no confirmation required. I'm not going to even bother reporting this because I am sure this is intended behavior - its just in plaintext in the skills manifests.

                                  So all that shit about your credentials being in a secure vault does not matter because you just get free read access from within the VM anyway! This includes your Instagram DMs, slack messages, your emails, box and Dropbox files, google docs, google contacts, all your fucking apple health readings, your flightaware flight histories, notion pages, fucking quickbooks data (!!!), your Tesla car data, and so many more fun things!

                                  What's fun is that some of the no confirmation needed actions are write actions too! You don't even need a clever exfil route, muse just gives it to you via your own connected accounts!

                                  eliocamp@mastodon.socialE This user is from outside of this forum
                                  eliocamp@mastodon.socialE This user is from outside of this forum
                                  eliocamp@mastodon.social
                                  wrote sidst redigeret af
                                  #200

                                  @jonny Wait... I am understanding this correctly? When you connect an account is not just the LLM that can read and write to that account, but also any other arbitrary program running on that VM?

                                  jonny@neuromatch.socialJ 1 Reply Last reply
                                  0
                                  • eliocamp@mastodon.socialE eliocamp@mastodon.social

                                    @jonny Wait... I am understanding this correctly? When you connect an account is not just the LLM that can read and write to that account, but also any other arbitrary program running on that VM?

                                    jonny@neuromatch.socialJ This user is from outside of this forum
                                    jonny@neuromatch.socialJ This user is from outside of this forum
                                    jonny@neuromatch.social
                                    wrote sidst redigeret af
                                    #201

                                    @eliocamp
                                    Correct.

                                    eliocamp@mastodon.socialE 1 Reply Last reply
                                    0
                                    • jonny@neuromatch.socialJ jonny@neuromatch.social

                                      I hadn't connected any accounts to muse until now, so I hooked up a test Instagram account just to check, and every connector allows read actions from binaries that can be called from within the VM by any process with no confirmation required. I'm not going to even bother reporting this because I am sure this is intended behavior - its just in plaintext in the skills manifests.

                                      So all that shit about your credentials being in a secure vault does not matter because you just get free read access from within the VM anyway! This includes your Instagram DMs, slack messages, your emails, box and Dropbox files, google docs, google contacts, all your fucking apple health readings, your flightaware flight histories, notion pages, fucking quickbooks data (!!!), your Tesla car data, and so many more fun things!

                                      What's fun is that some of the no confirmation needed actions are write actions too! You don't even need a clever exfil route, muse just gives it to you via your own connected accounts!

                                      jonny@neuromatch.socialJ This user is from outside of this forum
                                      jonny@neuromatch.socialJ This user is from outside of this forum
                                      jonny@neuromatch.social
                                      wrote sidst redigeret af
                                      #202

                                      I am trying to automate testing this by getting muse to sign up for a bunch of burner accounts. It can't really do that and now its googling how it, itself works

                                      loren@flipping.rocksL oldoldcojote@climatejustice.socialO jonny@neuromatch.socialJ 3 Replies Last reply
                                      0
                                      • jonny@neuromatch.socialJ jonny@neuromatch.social

                                        @eliocamp
                                        Correct.

                                        eliocamp@mastodon.socialE This user is from outside of this forum
                                        eliocamp@mastodon.socialE This user is from outside of this forum
                                        eliocamp@mastodon.social
                                        wrote sidst redigeret af
                                        #203

                                        @jonny
                                        *mickey mouse gouging his eyes out*

                                        1 Reply Last reply
                                        0
                                        • jonny@neuromatch.socialJ jonny@neuromatch.social

                                          I am trying to automate testing this by getting muse to sign up for a bunch of burner accounts. It can't really do that and now its googling how it, itself works

                                          loren@flipping.rocksL This user is from outside of this forum
                                          loren@flipping.rocksL This user is from outside of this forum
                                          loren@flipping.rocks
                                          wrote sidst redigeret af
                                          #204

                                          @jonny i have nothing to add but please keep it up. I have thoroughly enjoyed reading about these

                                          1 Reply Last reply
                                          0
                                          Svar
                                          • Svar som emne
                                          Login for at svare
                                          • Ældste til nyeste
                                          • Nyeste til ældste
                                          • Most Votes


                                          • Log ind

                                          • Login or register to search.
                                          Powered by NodeBB Contributors
                                          Graciously hosted by data.coop
                                          • First post
                                            Last post
                                          0
                                          • Hjem
                                          • Seneste
                                          • Etiketter
                                          • Populære
                                          • Verden
                                          • Bruger
                                          • Grupper