Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."
-
@futurebird apropos to absolutely nothing do you ever have to restrain yourself from using ant puns in daily speech
-
Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."
I keep hearing people say this and it's based on a huge misunderstanding of what training data is and how it works.
When books are scanned to train AI it's not like with "Internet Archive" they aren't preserved. It's not like you could ask the LLM to tell you what was on page 46 of the book and it could look that up.
1/
@futurebird I truly think if a considered effort about this got out without the typical anti-AI stuff attached (not that I’m saying that’s an incorrect point of view), the general public at large if they were informed about how there is no preservation happening and we have to basically buy information back that was already ours… they won’t like that.
I think this narrow area, which is critically important to our civilization is something that actually could be changed with amended laws.
-
S sebastian@social.itu.dk shared this topic
-
Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."
I keep hearing people say this and it's based on a huge misunderstanding of what training data is and how it works.
When books are scanned to train AI it's not like with "Internet Archive" they aren't preserved. It's not like you could ask the LLM to tell you what was on page 46 of the book and it could look that up.
1/
@futurebird
this
🤬 -
@futurebird I truly think if a considered effort about this got out without the typical anti-AI stuff attached (not that I’m saying that’s an incorrect point of view), the general public at large if they were informed about how there is no preservation happening and we have to basically buy information back that was already ours… they won’t like that.
I think this narrow area, which is critically important to our civilization is something that actually could be changed with amended laws.
@futurebird like for the most part I feel that it’s just a common sense thing. We shouldn’t be destroying knowledge as a means of intaking it due to a copyright rule. These companies should be required to preserve their own library or something and then if the company dissolves it should go into a public trust. This specific thing is a solvable problem outside of, you know, probably millions of books already being destroyed.
I just get the sense that this is a flock-like “duh” situation.
-
The mental image of greedy tech bros feeding precious rare books into the open flaming maw of some monstrous contraption is... kinda valid.
That is what is happening.
We could take action to change this. Just as you may be arrested for killing bald eagles, or setting a forest fire the law can protect things that no one person owns. Some of those things are the most valuable things that we have.
-
The mental image of greedy tech bros feeding precious rare books into the open flaming maw of some monstrous contraption is... kinda valid.
That is what is happening.
We could take action to change this. Just as you may be arrested for killing bald eagles, or setting a forest fire the law can protect things that no one person owns. Some of those things are the most valuable things that we have.
-
@futurebird apropos to absolutely nothing do you ever have to restrain yourself from using ant puns in daily speech
@xan There is hardly any topic for which @futurebird doesn't have an ant-based analogy or metaphor - and on the rare occasions she doesn't, she'll just shamelessly shoehorn in a mention of ants, anyway.
But that's why we love her. 🥰
-
@xan There is hardly any topic for which @futurebird doesn't have an ant-based analogy or metaphor - and on the rare occasions she doesn't, she'll just shamelessly shoehorn in a mention of ants, anyway.
But that's why we love her. 🥰
@ApostateEnglishman In meatspace, being as cat-aligned as I am, it is physically impossible for me to not curl my fingers into a paw shape in anticipation of uttering the word "pawsibly" in most interactions. I imagine the case is similar for @futurebird, otherwise I will eat my hat
-
Is this because it trained on these sites? Probably. But the LLM isn't designed to determine what part of the training data was significant for a particular output. Instead, after writing a response, like an undergraduate student who forgot to cite sources, it looks through the web for the sources that are the most similar to whatever it wrote.
If you know about research you know this is ass backwards.
But it's the best these kinds of systems can do.
3/
@futurebird I don't think this is true. Chatbot can search web first, put the stuff in its context window (i.e. read it) and generate answer based on that. There is no reason to do it the other way round.
It's closer to RAG than generation from memmory. Both Chatgpt and Claude, afaik, can do this.
LLM can still misinterpret what it read it or make up nonexistent sources, but I think it's designed better than as you describe it.
-
@futurebird I don't think this is true. Chatbot can search web first, put the stuff in its context window (i.e. read it) and generate answer based on that. There is no reason to do it the other way round.
It's closer to RAG than generation from memmory. Both Chatgpt and Claude, afaik, can do this.
LLM can still misinterpret what it read it or make up nonexistent sources, but I think it's designed better than as you describe it.
That could also work. But the fundamental issue is you don't really have a way to know what collection of 100s of documents created the text that you are reading.
And that's what I'd really like to see.
The rarebook that was fed to the machine may never come into this.
-
Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."
I keep hearing people say this and it's based on a huge misunderstanding of what training data is and how it works.
When books are scanned to train AI it's not like with "Internet Archive" they aren't preserved. It's not like you could ask the LLM to tell you what was on page 46 of the book and it could look that up.
1/
@futurebird The problem isn't with the books being disposed off. The problem is with the authors copyright being violated because a machine is using a media created for individual humans. This is clearly a mechanical way to skirt the copyright law.
-
Myth: "Public libraries and book stores have to dispose of books all the time so at least if these companies are scanning the books for AI they will be preserved for history."
I keep hearing people say this and it's based on a huge misunderstanding of what training data is and how it works.
When books are scanned to train AI it's not like with "Internet Archive" they aren't preserved. It's not like you could ask the LLM to tell you what was on page 46 of the book and it could look that up.
1/
@futurebird would it make sense to remove certain or some pages of books that are being sent to an AI feeding frenzy?
-
@futurebird would it make sense to remove certain or some pages of books that are being sent to an AI feeding frenzy?
LLMs work on big numbers and big data. To have an impact with any strategy that targets the training set you need to a spammer.
For example when and LLM reproduces entire pages of books verbatim that is often because those pages appeared in the training set multiple times. Probably 100s of times.
This is why "training set poisoning" strategies probably won't be as effective as people hope unless we're working on big scales.
-
LLMs work on big numbers and big data. To have an impact with any strategy that targets the training set you need to a spammer.
For example when and LLM reproduces entire pages of books verbatim that is often because those pages appeared in the training set multiple times. Probably 100s of times.
This is why "training set poisoning" strategies probably won't be as effective as people hope unless we're working on big scales.
I'm really interested in the customization and curation of training sets.
I think there is a lot of potential to make LLMs more useful with bespoke training sets and a totally different kind of output that isn't about chatting but rather about helping to organize and make the document set more useful.
-
LLMs work on big numbers and big data. To have an impact with any strategy that targets the training set you need to a spammer.
For example when and LLM reproduces entire pages of books verbatim that is often because those pages appeared in the training set multiple times. Probably 100s of times.
This is why "training set poisoning" strategies probably won't be as effective as people hope unless we're working on big scales.
@futurebird @MarkDW Can you explain why you think a large volume is required? I have seen papers to the contrary, eg:
https://arxiv.org/html/2510.07192v1 -
@futurebird I don't think this is true. Chatbot can search web first, put the stuff in its context window (i.e. read it) and generate answer based on that. There is no reason to do it the other way round.
It's closer to RAG than generation from memmory. Both Chatgpt and Claude, afaik, can do this.
LLM can still misinterpret what it read it or make up nonexistent sources, but I think it's designed better than as you describe it.
@janbogar @futurebird
Chatbot/ LLM /Generative AI is the worst Web search or document search ever invented.Just use a real search engine.
-
I'm really interested in the customization and curation of training sets.
I think there is a lot of potential to make LLMs more useful with bespoke training sets and a totally different kind of output that isn't about chatting but rather about helping to organize and make the document set more useful.
@futurebird @MarkDW
You don't need an LLM to do that better! -
It is likely that these companies *do* keep copies for future training. But they don't want to share this data or make it searchable by the public.
If someone destroys an old rare book to scan it the public should get a copy of the scan. This is about protecting our culture, history, and heritage.
It's about preserving research and science.
5/5
@futurebird Indeed. As I understand it, they specifically avoid keeping copies to avoid copyright issues.
-
J jwcph@helvede.net shared this topic
-
@janbogar @futurebird
Chatbot/ LLM /Generative AI is the worst Web search or document search ever invented.Just use a real search engine.
I was saying this two years ago however I don't think it's true anymore.
LLMs currently do a better job than web search engines filtering spam and returning the most relevant documents. This is in part because search engines have been made worse. But, we can't just tell people "use real search" when their experiences won't make that advice seem effective.
I agree that these systems should not be search engines.
But that is what they are becoming.
-
I was saying this two years ago however I don't think it's true anymore.
LLMs currently do a better job than web search engines filtering spam and returning the most relevant documents. This is in part because search engines have been made worse. But, we can't just tell people "use real search" when their experiences won't make that advice seem effective.
I agree that these systems should not be search engines.
But that is what they are becoming.
It's obvious that the makers of consumer LLMs want their systems to become the new primary interface for the web. When I look at what other people I know are doing in their office work, at colleges, they are using LLMs as search engines. Simply to avoid all of the junk and spam most search engines return. Many of them are bashful about it and don't like "AI" in general.
