OK!
-
@jonny ok so, as a card carrying ai hater, i just wanted to say this was very entertaining to read in kind of a horrible way. i have a question, you wrote "they can indeed do things, even if I think the circle around which things is much smaller than the maximalists" is there anything prosocial in that circle, and if so, what is it?
@jonny i ask because people keep saying stuff to the effect of "now we all have to agree that it's useful and worth burning everything down for, but" as if i'm meant to understand *what* it is useful for and solemnly hang my head in agreement that they've got a point, but everything i see it getting used for is completely horrific and i haven't a god damn idea why people keep saying this and i feel like i'm taking crazy pills
-
@jonny i ask because people keep saying stuff to the effect of "now we all have to agree that it's useful and worth burning everything down for, but" as if i'm meant to understand *what* it is useful for and solemnly hang my head in agreement that they've got a point, but everything i see it getting used for is completely horrific and i haven't a god damn idea why people keep saying this and i feel like i'm taking crazy pills
@jonny (read that in as much of a "being calm online" tone as you can imagine, i don't actually have all that much energy right now)
-
@jonny i ask because people keep saying stuff to the effect of "now we all have to agree that it's useful and worth burning everything down for, but" as if i'm meant to understand *what* it is useful for and solemnly hang my head in agreement that they've got a point, but everything i see it getting used for is completely horrific and i haven't a god damn idea why people keep saying this and i feel like i'm taking crazy pills
@aeva @jonny even looking at the output that people claim “would not have been possible without it” that is within my personal capacity to evaluate it all looks either undifferentiated from the authors’ previous work or obviously degraded in quality.
everyone gets very mad when I tell them I think they are deluding themselves because the technology is very convincingly fake, so I try not to say it too often, but I am right there with you here. I don’t get it
-
@aeva @jonny even looking at the output that people claim “would not have been possible without it” that is within my personal capacity to evaluate it all looks either undifferentiated from the authors’ previous work or obviously degraded in quality.
everyone gets very mad when I tell them I think they are deluding themselves because the technology is very convincingly fake, so I try not to say it too often, but I am right there with you here. I don’t get it
@glyph @jonny this whole, uh, moment reminds me of that scene in FF7 rebirth where Barret and Cloud attempt to check into a hotel https://youtu.be/_FTYm8k7r1c?si=uQ6b8qGkYvnD2oJb&t=1215
-
@aeva @jonny even looking at the output that people claim “would not have been possible without it” that is within my personal capacity to evaluate it all looks either undifferentiated from the authors’ previous work or obviously degraded in quality.
everyone gets very mad when I tell them I think they are deluding themselves because the technology is very convincingly fake, so I try not to say it too often, but I am right there with you here. I don’t get it
@aeva @jonny just today I was listening to a podcast where a very annoying man said in an the most sneering tone imaginable “well you know a couple of years ago everyone was saying they were ‘stochastic parrots’ and useless but OBVIOUSLY we have moved past that” and I shouted “objection! assuming facts not in evidence!” into an empty room
-
@jonny i ask because people keep saying stuff to the effect of "now we all have to agree that it's useful and worth burning everything down for, but" as if i'm meant to understand *what* it is useful for and solemnly hang my head in agreement that they've got a point, but everything i see it getting used for is completely horrific and i haven't a god damn idea why people keep saying this and i feel like i'm taking crazy pills
@aeva @jonny Last night I was trying to find an animal ER, rehabilitator, or just ANYONE who could help an injured young opossum I found under a car, and one of these animal sanctuaries I called had me speak to a fucking AI voice bot that even did the thing where it waits and fakes typing noises after each response I gave. This is a place meant to deal with sick and injured wildlife, sometimes in emergency situations, and I had to talk to ChatGPT for 10 minutes and give it my name and ZIP code
-
@aeva @jonny Last night I was trying to find an animal ER, rehabilitator, or just ANYONE who could help an injured young opossum I found under a car, and one of these animal sanctuaries I called had me speak to a fucking AI voice bot that even did the thing where it waits and fakes typing noises after each response I gave. This is a place meant to deal with sick and injured wildlife, sometimes in emergency situations, and I had to talk to ChatGPT for 10 minutes and give it my name and ZIP code
-
@aeva @jonny Last night I was trying to find an animal ER, rehabilitator, or just ANYONE who could help an injured young opossum I found under a car, and one of these animal sanctuaries I called had me speak to a fucking AI voice bot that even did the thing where it waits and fakes typing noises after each response I gave. This is a place meant to deal with sick and injured wildlife, sometimes in emergency situations, and I had to talk to ChatGPT for 10 minutes and give it my name and ZIP code
@aeva @jonny All just so it could not connect me to a person and instead text me the number of a person it pulled off the same state website I'd already checked.
My own job has implemented their own version of an AI call assistant, so occasionally I get an email transcript of a call between it and one of our customers and it's always the most useless piece of shit imaginable. It's incapable of doing anything but wasting people's time.
-
-
@aeva @jonny I was able to find someone eventually - by calling the ER vet next to my work, where a HUMAN woman immediately picked up and when I asked if she knew where I could take it, she looked at their own records and sent me to another ER vet that works with an "opossum lady" and they agreed to hold the critter for us overnight until she could come pick it up. (Fair, it was like 10pm by that point and it was not life threatening to the opossum) so it worked out, no thanks to AI
-
@aeva @jonny I was able to find someone eventually - by calling the ER vet next to my work, where a HUMAN woman immediately picked up and when I asked if she knew where I could take it, she looked at their own records and sent me to another ER vet that works with an "opossum lady" and they agreed to hold the critter for us overnight until she could come pick it up. (Fair, it was like 10pm by that point and it was not life threatening to the opossum) so it worked out, no thanks to AI
-
@aeva @jonny All just so it could not connect me to a person and instead text me the number of a person it pulled off the same state website I'd already checked.
My own job has implemented their own version of an AI call assistant, so occasionally I get an email transcript of a call between it and one of our customers and it's always the most useless piece of shit imaginable. It's incapable of doing anything but wasting people's time.
@aeva @jonny I imagine my bosses are paying this AI company tens of thousands of dollars for this contract just so it can fail to understand what people are asking and then tell them it will "let the team know" no matter how simple the request. It can't even take a payment using a customer's card on file. It can't quote prices when they ask. All it can do is send me an email so I know I need to call them back to actually help them. It's just fucking voicemail with more steps.
-
@jonny i ask because people keep saying stuff to the effect of "now we all have to agree that it's useful and worth burning everything down for, but" as if i'm meant to understand *what* it is useful for and solemnly hang my head in agreement that they've got a point, but everything i see it getting used for is completely horrific and i haven't a god damn idea why people keep saying this and i feel like i'm taking crazy pills
@aeva tbc I would stop well short of "so good its worth burning everything down for." Again my take on this is pretty boring - treating it like a properly scoped natural language interface to a deterministic system is not a terribly idea, and it doesn't require a very large model at all. So e.g. I find writing the few dozen lines of boilerplate to call an API to be tedious but not challenging, but it is possible to point an LLM at a swagger doc, write a simple adapter, and then translate instructions into api calls. the way I would prefer that they work is not how the 'agentic' tools work currently, I always have open a viewer to see the raw message streams as they happen and usually want them to prepare something that i can inspect and execute, rather than having them do it. The other thing i use them for is what everyone else also says they are capable of: generating boilerplate, or doing tedious things that are like one or two levels above what i could with a regex or an IDE refactoring tool on personal projects where correctness isn't the most important thing. e.g. i have an editor color scheme that i handwrote years ago that generates into different editors from a common declaration, but i used a fork of someone else's code from 2014 and i couldn't get the PHP dependencies to install anymore, so i was like "this is old, update the deps and make this run" and that worked because the mapping patterns are well represented in the training data. The third thing is genuine brute force tasks where i really do not care about the method and only care about the outcome. Corollary of that is also debugging, which they usually do by brute force, "here is observed bug, keep fiddling until you can diagnose" - the fix they propose is usually bonkers, but it does save me time from doing the fiddling myself so i can come up with a fix.
i would also say this is more "stuff that the LLMs can pass at" rather than be good at, probably the time i actually use them most by volume is when i am tired and have to fill some exogenous requirement that i don't care about but the chaff doesn't impact anyone else.
so idk, I think there could be a place for having small local models as a sort of utility inference for kind of trivial things like "i don't want to look up the cron syntax rn, hey computer, make a cron task to do this." the things I don't think they are fit for purpose are sort of 'the rest of the things they are used for where statistical text generation is not the what is desired,' any kind of "knowledge" task like search, or even summary, etc. because that's just not what they do, and this is just considering the technology in itself rather than its embeddedness with a corpofascistic plot to own all information and labor. there are a lot more things that the models can passably do but to me the error bars are way too wide and the quality is way too low and the energy use is way too high for me to justify - i don't think the entire model of everyone asking a 10 trillion parameter model as a calculator is sustainable or good, but having a small local model for constrained tasks is less objectionable to me.
-
-
@aeva @jonny even looking at the output that people claim “would not have been possible without it” that is within my personal capacity to evaluate it all looks either undifferentiated from the authors’ previous work or obviously degraded in quality.
everyone gets very mad when I tell them I think they are deluding themselves because the technology is very convincingly fake, so I try not to say it too often, but I am right there with you here. I don’t get it
-
-
@aeva tbc I would stop well short of "so good its worth burning everything down for." Again my take on this is pretty boring - treating it like a properly scoped natural language interface to a deterministic system is not a terribly idea, and it doesn't require a very large model at all. So e.g. I find writing the few dozen lines of boilerplate to call an API to be tedious but not challenging, but it is possible to point an LLM at a swagger doc, write a simple adapter, and then translate instructions into api calls. the way I would prefer that they work is not how the 'agentic' tools work currently, I always have open a viewer to see the raw message streams as they happen and usually want them to prepare something that i can inspect and execute, rather than having them do it. The other thing i use them for is what everyone else also says they are capable of: generating boilerplate, or doing tedious things that are like one or two levels above what i could with a regex or an IDE refactoring tool on personal projects where correctness isn't the most important thing. e.g. i have an editor color scheme that i handwrote years ago that generates into different editors from a common declaration, but i used a fork of someone else's code from 2014 and i couldn't get the PHP dependencies to install anymore, so i was like "this is old, update the deps and make this run" and that worked because the mapping patterns are well represented in the training data. The third thing is genuine brute force tasks where i really do not care about the method and only care about the outcome. Corollary of that is also debugging, which they usually do by brute force, "here is observed bug, keep fiddling until you can diagnose" - the fix they propose is usually bonkers, but it does save me time from doing the fiddling myself so i can come up with a fix.
i would also say this is more "stuff that the LLMs can pass at" rather than be good at, probably the time i actually use them most by volume is when i am tired and have to fill some exogenous requirement that i don't care about but the chaff doesn't impact anyone else.
so idk, I think there could be a place for having small local models as a sort of utility inference for kind of trivial things like "i don't want to look up the cron syntax rn, hey computer, make a cron task to do this." the things I don't think they are fit for purpose are sort of 'the rest of the things they are used for where statistical text generation is not the what is desired,' any kind of "knowledge" task like search, or even summary, etc. because that's just not what they do, and this is just considering the technology in itself rather than its embeddedness with a corpofascistic plot to own all information and labor. there are a lot more things that the models can passably do but to me the error bars are way too wide and the quality is way too low and the energy use is way too high for me to justify - i don't think the entire model of everyone asking a 10 trillion parameter model as a calculator is sustainable or good, but having a small local model for constrained tasks is less objectionable to me.
@jonny thanks for taking the time to write all that. I want to say I feel betrayed by everyone who has invoked the we all have to admit the bad thing is good cliche, but mostly i just feel completely sad and empty
-
@jonny thanks for taking the time to write all that. I want to say I feel betrayed by everyone who has invoked the we all have to admit the bad thing is good cliche, but mostly i just feel completely sad and empty
@aeva yeah, that is my overriding feeling as well. the circle of plausible and defensible uses is very small to me, and the stuff outside that circle is the shit of a nightmare future that is depeopled and everything is broken all the time. The people that say "we just have to accept that they are good now" are dismissing the blanket claim that they are bad at everything, but in doing so make the opposite blanket claim. It is worth a bit of subtlety to try and bound which things they can do and why. However it is not worth the subtlety to try and rescue the background context they exist in, which is the largest companies in the world heaving great quantities of capital around to avoid stock prices dropping from targeted ad revenues having saturated and trying to become middlemen injected into literally every crack of human endeavor.
-
@aeva tbc I would stop well short of "so good its worth burning everything down for." Again my take on this is pretty boring - treating it like a properly scoped natural language interface to a deterministic system is not a terribly idea, and it doesn't require a very large model at all. So e.g. I find writing the few dozen lines of boilerplate to call an API to be tedious but not challenging, but it is possible to point an LLM at a swagger doc, write a simple adapter, and then translate instructions into api calls. the way I would prefer that they work is not how the 'agentic' tools work currently, I always have open a viewer to see the raw message streams as they happen and usually want them to prepare something that i can inspect and execute, rather than having them do it. The other thing i use them for is what everyone else also says they are capable of: generating boilerplate, or doing tedious things that are like one or two levels above what i could with a regex or an IDE refactoring tool on personal projects where correctness isn't the most important thing. e.g. i have an editor color scheme that i handwrote years ago that generates into different editors from a common declaration, but i used a fork of someone else's code from 2014 and i couldn't get the PHP dependencies to install anymore, so i was like "this is old, update the deps and make this run" and that worked because the mapping patterns are well represented in the training data. The third thing is genuine brute force tasks where i really do not care about the method and only care about the outcome. Corollary of that is also debugging, which they usually do by brute force, "here is observed bug, keep fiddling until you can diagnose" - the fix they propose is usually bonkers, but it does save me time from doing the fiddling myself so i can come up with a fix.
i would also say this is more "stuff that the LLMs can pass at" rather than be good at, probably the time i actually use them most by volume is when i am tired and have to fill some exogenous requirement that i don't care about but the chaff doesn't impact anyone else.
so idk, I think there could be a place for having small local models as a sort of utility inference for kind of trivial things like "i don't want to look up the cron syntax rn, hey computer, make a cron task to do this." the things I don't think they are fit for purpose are sort of 'the rest of the things they are used for where statistical text generation is not the what is desired,' any kind of "knowledge" task like search, or even summary, etc. because that's just not what they do, and this is just considering the technology in itself rather than its embeddedness with a corpofascistic plot to own all information and labor. there are a lot more things that the models can passably do but to me the error bars are way too wide and the quality is way too low and the energy use is way too high for me to justify - i don't think the entire model of everyone asking a 10 trillion parameter model as a calculator is sustainable or good, but having a small local model for constrained tasks is less objectionable to me.
@jonny @aeva so this is all stuff that the models can “sometimes do” and the fact that it ever works is legitimately impressive, but every single one of these tasks is something that I, personally, have seen them screw up at, in ways which are potentially dangerous, even in my _extremely_ limited usage. not to mention that this is only safely usable by someone who *does* know what correct cron syntax looks like.
-
@jonny @aeva so this is all stuff that the models can “sometimes do” and the fact that it ever works is legitimately impressive, but every single one of these tasks is something that I, personally, have seen them screw up at, in ways which are potentially dangerous, even in my _extremely_ limited usage. not to mention that this is only safely usable by someone who *does* know what correct cron syntax looks like.
@jonny @aeva now granted some of that danger is in the harness or in the current product design, but this still doesn’t compensate for all the times that you will ask them to do a basic translation task like this for you and get a massively time-wasting error out.
so this might be the same thing you’re saying with your distinction, but I cannot see these as things the models are “useful for”

