Skip to content
  • Hjem
  • Seneste
  • Etiketter
  • Populære
  • Verden
  • Bruger
  • Grupper
Temaer
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Kollaps
FARVEL BIG TECH
  1. Forside
  2. Ikke-kategoriseret
  3. We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.

We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.

Planlagt Fastgjort Låst Flyttet Ikke-kategoriseret
39 Indlæg 32 Posters 21 Visninger
  • Ældste til nyeste
  • Nyeste til ældste
  • Most Votes
Svar
  • Svar som emne
Login for at svare
Denne tråd er blevet slettet. Kun brugere med emne behandlings privilegier kan se den.
  • blobster@infosec.exchangeB blobster@infosec.exchange

    @404mediaco

    That employees are scanning the barcodes or ISBNs on books — a unique serial number given to every published book — before scanning their content gives further credence to another theory put forth by booksellers: AI companies are trying to methodically scan every printed book in the world by working through the list of ISBNs. One bookseller told me they suspected this was the case because the very large orders they were getting never included very rare books that do not have ISBNs.

    That's enlightening and ghastly. Great article, thank you.

    wisegreyowl@mastodonapp.ukW This user is from outside of this forum
    wisegreyowl@mastodonapp.ukW This user is from outside of this forum
    wisegreyowl@mastodonapp.uk
    wrote sidst redigeret af
    #13

    @blobster @404mediaco Good job I never put an ISBN on my eBooks!

    1 Reply Last reply
    0
    • jpummil@genomic.socialJ jpummil@genomic.social

      @404mediaco Destroying rare books is beyond fucked up! The material already exists in digital form for the vast majority of texts…and if you MUST destroy to scan, use modern reprints!

      mark@mastodon.fixermark.comM This user is from outside of this forum
      mark@mastodon.fixermark.comM This user is from outside of this forum
      mark@mastodon.fixermark.com
      wrote sidst redigeret af
      #14

      @JPummil @404mediaco Due to the recently-settled lawsuit, they can't reliably use digital sources because they can't guarantee the pedigree on the data. But thanks to first-sale doctrine, if they buy a physical copy and destructively scan it, the resulting data is theirs to use for training with no risk of a copyright violation.

      1 Reply Last reply
      0
      • epic_null@infosec.exchangeE epic_null@infosec.exchange

        @Bredroll @JPummil @404mediaco I think you misunderstand - the destruction is to get around Americain Copyright Law.

        You can hate us more now.

        G This user is from outside of this forum
        G This user is from outside of this forum
        gerardthornley@hachyderm.io
        wrote sidst redigeret af
        #15

        @Epic_Null @Bredroll @JPummil @404mediaco
        I guess it's a fringe benefit that it will also gradually diminish the availability of reliable information, thus increasing the value of such information in our brave dystopian future. The motto can be "They who can afford books get to possess true knowledge."

        S 1 Reply Last reply
        0
        • blobster@infosec.exchangeB blobster@infosec.exchange

          @404mediaco

          That employees are scanning the barcodes or ISBNs on books — a unique serial number given to every published book — before scanning their content gives further credence to another theory put forth by booksellers: AI companies are trying to methodically scan every printed book in the world by working through the list of ISBNs. One bookseller told me they suspected this was the case because the very large orders they were getting never included very rare books that do not have ISBNs.

          That's enlightening and ghastly. Great article, thank you.

          adamshostack@infosec.exchangeA This user is from outside of this forum
          adamshostack@infosec.exchangeA This user is from outside of this forum
          adamshostack@infosec.exchange
          wrote sidst redigeret af
          #16

          @blobster @404mediaco Just for clarity, ISBNs track each edition of a book; hardbacks and paperbacks will get different ISBNs.

          So this is a sorta dumb way to do what they want to do.

          (Eg, 0374275637 and 0374533555 are both Kahneman's Thinking Fast and Slow, in hardback and paperback, respectively.) They may also be buying the same book with the ISBN-10 and ISBN-13, 978-0374533557 for the paperback.)

          drwho@masto.hackers.townD 1 Reply Last reply
          0
          • 404mediaco@mastodon.social4 404mediaco@mastodon.social

            We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.

            https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/

            raven667@hachyderm.ioR This user is from outside of this forum
            raven667@hachyderm.ioR This user is from outside of this forum
            raven667@hachyderm.io
            wrote sidst redigeret af
            #17

            @404mediaco baller move and great story

            1 Reply Last reply
            0
            • jrdumas@piaille.frJ jrdumas@piaille.fr

              @Epic_Null @Bredroll @JPummil @404mediaco And I guess it's also the reason why they'll never release the PDF or text files they now have in their possession. That, at least, would have been slightly comforting...

              lostgen@det.socialL This user is from outside of this forum
              lostgen@det.socialL This user is from outside of this forum
              lostgen@det.social
              wrote sidst redigeret af
              #18

              @jrdumas They do release it. It's in the agent and can be more or less extracted word-by-word. @Epic_Null @Bredroll @JPummil @404mediaco

              astromancer5g@spore.socialA 1 Reply Last reply
              0
              • adamshostack@infosec.exchangeA adamshostack@infosec.exchange

                @blobster @404mediaco Just for clarity, ISBNs track each edition of a book; hardbacks and paperbacks will get different ISBNs.

                So this is a sorta dumb way to do what they want to do.

                (Eg, 0374275637 and 0374533555 are both Kahneman's Thinking Fast and Slow, in hardback and paperback, respectively.) They may also be buying the same book with the ISBN-10 and ISBN-13, 978-0374533557 for the paperback.)

                drwho@masto.hackers.townD This user is from outside of this forum
                drwho@masto.hackers.townD This user is from outside of this forum
                drwho@masto.hackers.town
                wrote sidst redigeret af
                #19

                @adamshostack @blobster @404mediaco Both editions are being pulled from circulation and destroyed, though. Which I think might be the point.

                adamshostack@infosec.exchangeA 1 Reply Last reply
                0
                • drwho@masto.hackers.townD drwho@masto.hackers.town

                  @adamshostack @blobster @404mediaco Both editions are being pulled from circulation and destroyed, though. Which I think might be the point.

                  adamshostack@infosec.exchangeA This user is from outside of this forum
                  adamshostack@infosec.exchangeA This user is from outside of this forum
                  adamshostack@infosec.exchange
                  wrote sidst redigeret af
                  #20

                  @drwho @blobster @404mediaco Why would you bother scanning the same book twice?

                  I believe that most of the rare books being gathered are probably self-published works, which are both rare, and most of those are obscure for good reasons. (The recent growth of self-publish with Kindle changes the ratio somewhat.)

                  drwho@masto.hackers.townD 1 Reply Last reply
                  0
                  • bredroll@mas.toB This user is from outside of this forum
                    bredroll@mas.toB This user is from outside of this forum
                    bredroll@mas.to
                    wrote sidst redigeret af
                    #21

                    @FediThing @Epic_Null @JPummil @404mediaco they are very much "do thing" and "ignore laws later"

                    1 Reply Last reply
                    0
                    • 404mediaco@mastodon.social4 404mediaco@mastodon.social

                      We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.

                      https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/

                      imprinted@mastodon.unoI This user is from outside of this forum
                      imprinted@mastodon.unoI This user is from outside of this forum
                      imprinted@mastodon.uno
                      wrote sidst redigeret af
                      #22

                      @404mediaco
                      So they are not training AI, they want to be the gatekeeper of knowledge, with a fare.
                      We should buy books in bulk.

                      1 Reply Last reply
                      0
                      • adamshostack@infosec.exchangeA adamshostack@infosec.exchange

                        @drwho @blobster @404mediaco Why would you bother scanning the same book twice?

                        I believe that most of the rare books being gathered are probably self-published works, which are both rare, and most of those are obscure for good reasons. (The recent growth of self-publish with Kindle changes the ratio somewhat.)

                        drwho@masto.hackers.townD This user is from outside of this forum
                        drwho@masto.hackers.townD This user is from outside of this forum
                        drwho@masto.hackers.town
                        wrote sidst redigeret af
                        #23

                        @adamshostack @blobster @404mediaco It gets one more copy out of circulation.

                        This feels uncomfortably like another step in the recapitulation of history.

                        1 Reply Last reply
                        0
                        • nemobis@mamot.frN nemobis@mamot.fr

                          «These booksellers suspected AI companies were behind these large bulk purchases because of the high number of books they were buying, the seemingly random choice of books, and the fact that these buyers, unlike libraries and universities, did not seem price sensitive at all.»
                          https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/

                          I hope that booksellers raise prices for their rare books accordingly!

                          penguinflight@mastodon.deP This user is from outside of this forum
                          penguinflight@mastodon.deP This user is from outside of this forum
                          penguinflight@mastodon.de
                          wrote sidst redigeret af
                          #24

                          @nemobis I hope they stop fucking SELLING them to these cunts.

                          marion_grau@climatejustice.socialM 1 Reply Last reply
                          0
                          • 404mediaco@mastodon.social4 404mediaco@mastodon.social

                            We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.

                            https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/

                            ids1024@mathstodon.xyzI This user is from outside of this forum
                            ids1024@mathstodon.xyzI This user is from outside of this forum
                            ids1024@mathstodon.xyz
                            wrote sidst redigeret af
                            #25

                            @404mediaco This just seems kind of desperate and futile. If their model can't perform well after it's been trained on every book that isn't out of print, what happens when they've scanned and destroyed every rare book too, and it still isn't good enough?

                            1 Reply Last reply
                            0
                            • 404mediaco@mastodon.social4 404mediaco@mastodon.social

                              We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.

                              https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/

                              sf105@mastodonapp.ukS This user is from outside of this forum
                              sf105@mastodonapp.ukS This user is from outside of this forum
                              sf105@mastodonapp.uk
                              wrote sidst redigeret af
                              #26

                              @404mediaco I saw a talk in 90s from the director of one of the copyright libraries (Cornell?) about the problems of preserving books. Some libraries chose de-acidification (slow, expensive) and some digitisation (destructive as we now all know).
                              - the question of what to digitise is not simple, it’s not just the text
                              - the purpose of digitisation is to make the content available to everyone
                              - presumably Amazon was unable to reach an agreement to access existing digitised content which is why it has to do its own?

                              misjavanlaatum@mastodon.gamedev.placeM 1 Reply Last reply
                              0
                              • nemobis@mamot.frN nemobis@mamot.fr

                                «These booksellers suspected AI companies were behind these large bulk purchases because of the high number of books they were buying, the seemingly random choice of books, and the fact that these buyers, unlike libraries and universities, did not seem price sensitive at all.»
                                https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/

                                I hope that booksellers raise prices for their rare books accordingly!

                                nemobis@mamot.frN This user is from outside of this forum
                                nemobis@mamot.frN This user is from outside of this forum
                                nemobis@mamot.fr
                                wrote sidst redigeret af
                                #27

                                Some qualifications needed: https://mamot.fr/@nemobis/117111964050668486

                                1 Reply Last reply
                                0
                                • bredroll@mas.toB bredroll@mas.to

                                  @JPummil @404mediaco they just cant be fucked to create nondestructive scanners, bunch of wankers

                                  dcatoffm@mastodon.socialD This user is from outside of this forum
                                  dcatoffm@mastodon.socialD This user is from outside of this forum
                                  dcatoffm@mastodon.social
                                  wrote sidst redigeret af
                                  #28

                                  @Bredroll
                                  @JPummil @404mediaco

                                  Create? The internet archive has been using nondestructive book scanners for years, it's nothing new.

                                  It's also nowhere near as fast as just slicing off the spine and dropping the stack on a high speed ADF.

                                  The machine is hungry. "Feed me, Seymour."

                                  1 Reply Last reply
                                  0
                                  • sf105@mastodonapp.ukS sf105@mastodonapp.uk

                                    @404mediaco I saw a talk in 90s from the director of one of the copyright libraries (Cornell?) about the problems of preserving books. Some libraries chose de-acidification (slow, expensive) and some digitisation (destructive as we now all know).
                                    - the question of what to digitise is not simple, it’s not just the text
                                    - the purpose of digitisation is to make the content available to everyone
                                    - presumably Amazon was unable to reach an agreement to access existing digitised content which is why it has to do its own?

                                    misjavanlaatum@mastodon.gamedev.placeM This user is from outside of this forum
                                    misjavanlaatum@mastodon.gamedev.placeM This user is from outside of this forum
                                    misjavanlaatum@mastodon.gamedev.place
                                    wrote sidst redigeret af
                                    #29

                                    @sf105 @404mediaco woah, steady on there: digitisation is not necessarily destructive, right? Whatever happened to scanning page by page?

                                    bluestarultor@tech.lgbtB 1 Reply Last reply
                                    0
                                    • G gerardthornley@hachyderm.io

                                      @Epic_Null @Bredroll @JPummil @404mediaco
                                      I guess it's a fringe benefit that it will also gradually diminish the availability of reliable information, thus increasing the value of such information in our brave dystopian future. The motto can be "They who can afford books get to possess true knowledge."

                                      S This user is from outside of this forum
                                      S This user is from outside of this forum
                                      shadsterling@mastodon.social
                                      wrote sidst redigeret af
                                      #30

                                      @GerardThornley @Epic_Null @Bredroll @JPummil @404mediaco everything old is new again

                                      1 Reply Last reply
                                      0
                                      • ids1024@mathstodon.xyzI This user is from outside of this forum
                                        ids1024@mathstodon.xyzI This user is from outside of this forum
                                        ids1024@mathstodon.xyz
                                        wrote sidst redigeret af
                                        #31

                                        @hakfoo @404mediaco Yeah. I'm not an expert in how training these models works, but I'd expect adding more input (to any already wildly massive set of training data) of old books that are inaccurate, outdated, or otherwise not generally of interest to modern readers will at best be a *very* marginal improvement to the model, and at worst bloat the model while make it produce more outdated and irrelevant information.

                                        1 Reply Last reply
                                        0
                                        • penguinflight@mastodon.deP penguinflight@mastodon.de

                                          @nemobis I hope they stop fucking SELLING them to these cunts.

                                          marion_grau@climatejustice.socialM This user is from outside of this forum
                                          marion_grau@climatejustice.socialM This user is from outside of this forum
                                          marion_grau@climatejustice.social
                                          wrote sidst redigeret af
                                          #32

                                          @penguinflight @nemobis this

                                          1 Reply Last reply
                                          0
                                          Svar
                                          • Svar som emne
                                          Login for at svare
                                          • Ældste til nyeste
                                          • Nyeste til ældste
                                          • Most Votes


                                          • Log ind

                                          • Har du ikke en konto? Tilmeld

                                          • Login or register to search.
                                          Powered by NodeBB Contributors
                                          Graciously hosted by data.coop
                                          • First post
                                            Last post
                                          0
                                          • Hjem
                                          • Seneste
                                          • Etiketter
                                          • Populære
                                          • Verden
                                          • Bruger
                                          • Grupper