We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.
-
We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.
«These booksellers suspected AI companies were behind these large bulk purchases because of the high number of books they were buying, the seemingly random choice of books, and the fact that these buyers, unlike libraries and universities, did not seem price sensitive at all.»
https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/I hope that booksellers raise prices for their rare books accordingly!
-
@404mediaco Destroying rare books is beyond fucked up! The material already exists in digital form for the vast majority of texts…and if you MUST destroy to scan, use modern reprints!
@JPummil @404mediaco ...and if you just wanna shred books, there's millions of copies of JKRowling that are ready...
-
That employees are scanning the barcodes or ISBNs on books — a unique serial number given to every published book — before scanning their content gives further credence to another theory put forth by booksellers: AI companies are trying to methodically scan every printed book in the world by working through the list of ISBNs. One bookseller told me they suspected this was the case because the very large orders they were getting never included very rare books that do not have ISBNs.
That's enlightening and ghastly. Great article, thank you.
@blobster @404mediaco Good job I never put an ISBN on my eBooks!
-
@404mediaco Destroying rare books is beyond fucked up! The material already exists in digital form for the vast majority of texts…and if you MUST destroy to scan, use modern reprints!
@JPummil @404mediaco Due to the recently-settled lawsuit, they can't reliably use digital sources because they can't guarantee the pedigree on the data. But thanks to first-sale doctrine, if they buy a physical copy and destructively scan it, the resulting data is theirs to use for training with no risk of a copyright violation.
-
@Bredroll @JPummil @404mediaco I think you misunderstand - the destruction is to get around Americain Copyright Law.
You can hate us more now.
@Epic_Null @Bredroll @JPummil @404mediaco
I guess it's a fringe benefit that it will also gradually diminish the availability of reliable information, thus increasing the value of such information in our brave dystopian future. The motto can be "They who can afford books get to possess true knowledge." -
That employees are scanning the barcodes or ISBNs on books — a unique serial number given to every published book — before scanning their content gives further credence to another theory put forth by booksellers: AI companies are trying to methodically scan every printed book in the world by working through the list of ISBNs. One bookseller told me they suspected this was the case because the very large orders they were getting never included very rare books that do not have ISBNs.
That's enlightening and ghastly. Great article, thank you.
@blobster @404mediaco Just for clarity, ISBNs track each edition of a book; hardbacks and paperbacks will get different ISBNs.
So this is a sorta dumb way to do what they want to do.
(Eg, 0374275637 and 0374533555 are both Kahneman's Thinking Fast and Slow, in hardback and paperback, respectively.) They may also be buying the same book with the ISBN-10 and ISBN-13, 978-0374533557 for the paperback.)
-
We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.
@404mediaco baller move and great story
-
@Epic_Null @Bredroll @JPummil @404mediaco And I guess it's also the reason why they'll never release the PDF or text files they now have in their possession. That, at least, would have been slightly comforting...
@jrdumas They do release it. It's in the agent and can be more or less extracted word-by-word. @Epic_Null @Bredroll @JPummil @404mediaco
-
@blobster @404mediaco Just for clarity, ISBNs track each edition of a book; hardbacks and paperbacks will get different ISBNs.
So this is a sorta dumb way to do what they want to do.
(Eg, 0374275637 and 0374533555 are both Kahneman's Thinking Fast and Slow, in hardback and paperback, respectively.) They may also be buying the same book with the ISBN-10 and ISBN-13, 978-0374533557 for the paperback.)
@adamshostack @blobster @404mediaco Both editions are being pulled from circulation and destroyed, though. Which I think might be the point.
-
@adamshostack @blobster @404mediaco Both editions are being pulled from circulation and destroyed, though. Which I think might be the point.
@drwho @blobster @404mediaco Why would you bother scanning the same book twice?
I believe that most of the rare books being gathered are probably self-published works, which are both rare, and most of those are obscure for good reasons. (The recent growth of self-publish with Kindle changes the ratio somewhat.)
-
@FediThing @Epic_Null @JPummil @404mediaco they are very much "do thing" and "ignore laws later"
-
We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.
@404mediaco
So they are not training AI, they want to be the gatekeeper of knowledge, with a fare.
We should buy books in bulk. -
@drwho @blobster @404mediaco Why would you bother scanning the same book twice?
I believe that most of the rare books being gathered are probably self-published works, which are both rare, and most of those are obscure for good reasons. (The recent growth of self-publish with Kindle changes the ratio somewhat.)
@adamshostack @blobster @404mediaco It gets one more copy out of circulation.
This feels uncomfortably like another step in the recapitulation of history.
-
«These booksellers suspected AI companies were behind these large bulk purchases because of the high number of books they were buying, the seemingly random choice of books, and the fact that these buyers, unlike libraries and universities, did not seem price sensitive at all.»
https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/I hope that booksellers raise prices for their rare books accordingly!
@nemobis I hope they stop fucking SELLING them to these cunts.
-
We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.
@404mediaco This just seems kind of desperate and futile. If their model can't perform well after it's been trained on every book that isn't out of print, what happens when they've scanned and destroyed every rare book too, and it still isn't good enough?
-
We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.
@404mediaco I saw a talk in 90s from the director of one of the copyright libraries (Cornell?) about the problems of preserving books. Some libraries chose de-acidification (slow, expensive) and some digitisation (destructive as we now all know).
- the question of what to digitise is not simple, it’s not just the text
- the purpose of digitisation is to make the content available to everyone
- presumably Amazon was unable to reach an agreement to access existing digitised content which is why it has to do its own? -
«These booksellers suspected AI companies were behind these large bulk purchases because of the high number of books they were buying, the seemingly random choice of books, and the fact that these buyers, unlike libraries and universities, did not seem price sensitive at all.»
https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/I hope that booksellers raise prices for their rare books accordingly!
Some qualifications needed: https://mamot.fr/@nemobis/117111964050668486
-
@JPummil @404mediaco they just cant be fucked to create nondestructive scanners, bunch of wankers
Create? The internet archive has been using nondestructive book scanners for years, it's nothing new.
It's also nowhere near as fast as just slicing off the spine and dropping the stack on a high speed ADF.
The machine is hungry. "Feed me, Seymour."
-
@404mediaco I saw a talk in 90s from the director of one of the copyright libraries (Cornell?) about the problems of preserving books. Some libraries chose de-acidification (slow, expensive) and some digitisation (destructive as we now all know).
- the question of what to digitise is not simple, it’s not just the text
- the purpose of digitisation is to make the content available to everyone
- presumably Amazon was unable to reach an agreement to access existing digitised content which is why it has to do its own?@sf105 @404mediaco woah, steady on there: digitisation is not necessarily destructive, right? Whatever happened to scanning page by page?
-
@Epic_Null @Bredroll @JPummil @404mediaco
I guess it's a fringe benefit that it will also gradually diminish the availability of reliable information, thus increasing the value of such information in our brave dystopian future. The motto can be "They who can afford books get to possess true knowledge."@GerardThornley @Epic_Null @Bredroll @JPummil @404mediaco everything old is new again