Skip to content
  • Hjem
  • Seneste
  • Etiketter
  • Populære
  • Verden
  • Bruger
  • Grupper
Temaer
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Kollaps
FARVEL BIG TECH
  1. Forside
  2. Ikke-kategoriseret
  3. We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.

We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.

Planlagt Fastgjort Låst Flyttet Ikke-kategoriseret
39 Indlæg 32 Posters 21 Visninger
  • Ældste til nyeste
  • Nyeste til ældste
  • Most Votes
Svar
  • Svar som emne
Login for at svare
Denne tråd er blevet slettet. Kun brugere med emne behandlings privilegier kan se den.
  • jpummil@genomic.socialJ jpummil@genomic.social

    @404mediaco Destroying rare books is beyond fucked up! The material already exists in digital form for the vast majority of texts…and if you MUST destroy to scan, use modern reprints!

    the_turtle@polyglot.cityT This user is from outside of this forum
    the_turtle@polyglot.cityT This user is from outside of this forum
    the_turtle@polyglot.city
    wrote sidst redigeret af
    #12

    @JPummil @404mediaco ...and if you just wanna shred books, there's millions of copies of JKRowling that are ready...

    1 Reply Last reply
    0
    • blobster@infosec.exchangeB blobster@infosec.exchange

      @404mediaco

      That employees are scanning the barcodes or ISBNs on books — a unique serial number given to every published book — before scanning their content gives further credence to another theory put forth by booksellers: AI companies are trying to methodically scan every printed book in the world by working through the list of ISBNs. One bookseller told me they suspected this was the case because the very large orders they were getting never included very rare books that do not have ISBNs.

      That's enlightening and ghastly. Great article, thank you.

      wisegreyowl@mastodonapp.ukW This user is from outside of this forum
      wisegreyowl@mastodonapp.ukW This user is from outside of this forum
      wisegreyowl@mastodonapp.uk
      wrote sidst redigeret af
      #13

      @blobster @404mediaco Good job I never put an ISBN on my eBooks!

      1 Reply Last reply
      0
      • jpummil@genomic.socialJ jpummil@genomic.social

        @404mediaco Destroying rare books is beyond fucked up! The material already exists in digital form for the vast majority of texts…and if you MUST destroy to scan, use modern reprints!

        mark@mastodon.fixermark.comM This user is from outside of this forum
        mark@mastodon.fixermark.comM This user is from outside of this forum
        mark@mastodon.fixermark.com
        wrote sidst redigeret af
        #14

        @JPummil @404mediaco Due to the recently-settled lawsuit, they can't reliably use digital sources because they can't guarantee the pedigree on the data. But thanks to first-sale doctrine, if they buy a physical copy and destructively scan it, the resulting data is theirs to use for training with no risk of a copyright violation.

        1 Reply Last reply
        0
        • epic_null@infosec.exchangeE epic_null@infosec.exchange

          @Bredroll @JPummil @404mediaco I think you misunderstand - the destruction is to get around Americain Copyright Law.

          You can hate us more now.

          G This user is from outside of this forum
          G This user is from outside of this forum
          gerardthornley@hachyderm.io
          wrote sidst redigeret af
          #15

          @Epic_Null @Bredroll @JPummil @404mediaco
          I guess it's a fringe benefit that it will also gradually diminish the availability of reliable information, thus increasing the value of such information in our brave dystopian future. The motto can be "They who can afford books get to possess true knowledge."

          S 1 Reply Last reply
          0
          • blobster@infosec.exchangeB blobster@infosec.exchange

            @404mediaco

            That employees are scanning the barcodes or ISBNs on books — a unique serial number given to every published book — before scanning their content gives further credence to another theory put forth by booksellers: AI companies are trying to methodically scan every printed book in the world by working through the list of ISBNs. One bookseller told me they suspected this was the case because the very large orders they were getting never included very rare books that do not have ISBNs.

            That's enlightening and ghastly. Great article, thank you.

            adamshostack@infosec.exchangeA This user is from outside of this forum
            adamshostack@infosec.exchangeA This user is from outside of this forum
            adamshostack@infosec.exchange
            wrote sidst redigeret af
            #16

            @blobster @404mediaco Just for clarity, ISBNs track each edition of a book; hardbacks and paperbacks will get different ISBNs.

            So this is a sorta dumb way to do what they want to do.

            (Eg, 0374275637 and 0374533555 are both Kahneman's Thinking Fast and Slow, in hardback and paperback, respectively.) They may also be buying the same book with the ISBN-10 and ISBN-13, 978-0374533557 for the paperback.)

            drwho@masto.hackers.townD 1 Reply Last reply
            0
            • 404mediaco@mastodon.social4 404mediaco@mastodon.social

              We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.

              https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/

              raven667@hachyderm.ioR This user is from outside of this forum
              raven667@hachyderm.ioR This user is from outside of this forum
              raven667@hachyderm.io
              wrote sidst redigeret af
              #17

              @404mediaco baller move and great story

              1 Reply Last reply
              0
              • jrdumas@piaille.frJ jrdumas@piaille.fr

                @Epic_Null @Bredroll @JPummil @404mediaco And I guess it's also the reason why they'll never release the PDF or text files they now have in their possession. That, at least, would have been slightly comforting...

                lostgen@det.socialL This user is from outside of this forum
                lostgen@det.socialL This user is from outside of this forum
                lostgen@det.social
                wrote sidst redigeret af
                #18

                @jrdumas They do release it. It's in the agent and can be more or less extracted word-by-word. @Epic_Null @Bredroll @JPummil @404mediaco

                astromancer5g@spore.socialA 1 Reply Last reply
                0
                • adamshostack@infosec.exchangeA adamshostack@infosec.exchange

                  @blobster @404mediaco Just for clarity, ISBNs track each edition of a book; hardbacks and paperbacks will get different ISBNs.

                  So this is a sorta dumb way to do what they want to do.

                  (Eg, 0374275637 and 0374533555 are both Kahneman's Thinking Fast and Slow, in hardback and paperback, respectively.) They may also be buying the same book with the ISBN-10 and ISBN-13, 978-0374533557 for the paperback.)

                  drwho@masto.hackers.townD This user is from outside of this forum
                  drwho@masto.hackers.townD This user is from outside of this forum
                  drwho@masto.hackers.town
                  wrote sidst redigeret af
                  #19

                  @adamshostack @blobster @404mediaco Both editions are being pulled from circulation and destroyed, though. Which I think might be the point.

                  adamshostack@infosec.exchangeA 1 Reply Last reply
                  0
                  • drwho@masto.hackers.townD drwho@masto.hackers.town

                    @adamshostack @blobster @404mediaco Both editions are being pulled from circulation and destroyed, though. Which I think might be the point.

                    adamshostack@infosec.exchangeA This user is from outside of this forum
                    adamshostack@infosec.exchangeA This user is from outside of this forum
                    adamshostack@infosec.exchange
                    wrote sidst redigeret af
                    #20

                    @drwho @blobster @404mediaco Why would you bother scanning the same book twice?

                    I believe that most of the rare books being gathered are probably self-published works, which are both rare, and most of those are obscure for good reasons. (The recent growth of self-publish with Kindle changes the ratio somewhat.)

                    drwho@masto.hackers.townD 1 Reply Last reply
                    0
                    • bredroll@mas.toB This user is from outside of this forum
                      bredroll@mas.toB This user is from outside of this forum
                      bredroll@mas.to
                      wrote sidst redigeret af
                      #21

                      @FediThing @Epic_Null @JPummil @404mediaco they are very much "do thing" and "ignore laws later"

                      1 Reply Last reply
                      0
                      • 404mediaco@mastodon.social4 404mediaco@mastodon.social

                        We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.

                        https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/

                        imprinted@mastodon.unoI This user is from outside of this forum
                        imprinted@mastodon.unoI This user is from outside of this forum
                        imprinted@mastodon.uno
                        wrote sidst redigeret af
                        #22

                        @404mediaco
                        So they are not training AI, they want to be the gatekeeper of knowledge, with a fare.
                        We should buy books in bulk.

                        1 Reply Last reply
                        0
                        • adamshostack@infosec.exchangeA adamshostack@infosec.exchange

                          @drwho @blobster @404mediaco Why would you bother scanning the same book twice?

                          I believe that most of the rare books being gathered are probably self-published works, which are both rare, and most of those are obscure for good reasons. (The recent growth of self-publish with Kindle changes the ratio somewhat.)

                          drwho@masto.hackers.townD This user is from outside of this forum
                          drwho@masto.hackers.townD This user is from outside of this forum
                          drwho@masto.hackers.town
                          wrote sidst redigeret af
                          #23

                          @adamshostack @blobster @404mediaco It gets one more copy out of circulation.

                          This feels uncomfortably like another step in the recapitulation of history.

                          1 Reply Last reply
                          0
                          • nemobis@mamot.frN nemobis@mamot.fr

                            «These booksellers suspected AI companies were behind these large bulk purchases because of the high number of books they were buying, the seemingly random choice of books, and the fact that these buyers, unlike libraries and universities, did not seem price sensitive at all.»
                            https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/

                            I hope that booksellers raise prices for their rare books accordingly!

                            penguinflight@mastodon.deP This user is from outside of this forum
                            penguinflight@mastodon.deP This user is from outside of this forum
                            penguinflight@mastodon.de
                            wrote sidst redigeret af
                            #24

                            @nemobis I hope they stop fucking SELLING them to these cunts.

                            marion_grau@climatejustice.socialM 1 Reply Last reply
                            0
                            • 404mediaco@mastodon.social4 404mediaco@mastodon.social

                              We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.

                              https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/

                              ids1024@mathstodon.xyzI This user is from outside of this forum
                              ids1024@mathstodon.xyzI This user is from outside of this forum
                              ids1024@mathstodon.xyz
                              wrote sidst redigeret af
                              #25

                              @404mediaco This just seems kind of desperate and futile. If their model can't perform well after it's been trained on every book that isn't out of print, what happens when they've scanned and destroyed every rare book too, and it still isn't good enough?

                              1 Reply Last reply
                              0
                              • 404mediaco@mastodon.social4 404mediaco@mastodon.social

                                We placed a tracking device in a shipment of rare books to see which AI company was buying it, and found an Amazon facility where Amazon scans and destroys books.

                                https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/

                                sf105@mastodonapp.ukS This user is from outside of this forum
                                sf105@mastodonapp.ukS This user is from outside of this forum
                                sf105@mastodonapp.uk
                                wrote sidst redigeret af
                                #26

                                @404mediaco I saw a talk in 90s from the director of one of the copyright libraries (Cornell?) about the problems of preserving books. Some libraries chose de-acidification (slow, expensive) and some digitisation (destructive as we now all know).
                                - the question of what to digitise is not simple, it’s not just the text
                                - the purpose of digitisation is to make the content available to everyone
                                - presumably Amazon was unable to reach an agreement to access existing digitised content which is why it has to do its own?

                                misjavanlaatum@mastodon.gamedev.placeM 1 Reply Last reply
                                0
                                • nemobis@mamot.frN nemobis@mamot.fr

                                  «These booksellers suspected AI companies were behind these large bulk purchases because of the high number of books they were buying, the seemingly random choice of books, and the fact that these buyers, unlike libraries and universities, did not seem price sensitive at all.»
                                  https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/

                                  I hope that booksellers raise prices for their rare books accordingly!

                                  nemobis@mamot.frN This user is from outside of this forum
                                  nemobis@mamot.frN This user is from outside of this forum
                                  nemobis@mamot.fr
                                  wrote sidst redigeret af
                                  #27

                                  Some qualifications needed: https://mamot.fr/@nemobis/117111964050668486

                                  1 Reply Last reply
                                  0
                                  • bredroll@mas.toB bredroll@mas.to

                                    @JPummil @404mediaco they just cant be fucked to create nondestructive scanners, bunch of wankers

                                    dcatoffm@mastodon.socialD This user is from outside of this forum
                                    dcatoffm@mastodon.socialD This user is from outside of this forum
                                    dcatoffm@mastodon.social
                                    wrote sidst redigeret af
                                    #28

                                    @Bredroll
                                    @JPummil @404mediaco

                                    Create? The internet archive has been using nondestructive book scanners for years, it's nothing new.

                                    It's also nowhere near as fast as just slicing off the spine and dropping the stack on a high speed ADF.

                                    The machine is hungry. "Feed me, Seymour."

                                    1 Reply Last reply
                                    0
                                    • sf105@mastodonapp.ukS sf105@mastodonapp.uk

                                      @404mediaco I saw a talk in 90s from the director of one of the copyright libraries (Cornell?) about the problems of preserving books. Some libraries chose de-acidification (slow, expensive) and some digitisation (destructive as we now all know).
                                      - the question of what to digitise is not simple, it’s not just the text
                                      - the purpose of digitisation is to make the content available to everyone
                                      - presumably Amazon was unable to reach an agreement to access existing digitised content which is why it has to do its own?

                                      misjavanlaatum@mastodon.gamedev.placeM This user is from outside of this forum
                                      misjavanlaatum@mastodon.gamedev.placeM This user is from outside of this forum
                                      misjavanlaatum@mastodon.gamedev.place
                                      wrote sidst redigeret af
                                      #29

                                      @sf105 @404mediaco woah, steady on there: digitisation is not necessarily destructive, right? Whatever happened to scanning page by page?

                                      bluestarultor@tech.lgbtB 1 Reply Last reply
                                      0
                                      • G gerardthornley@hachyderm.io

                                        @Epic_Null @Bredroll @JPummil @404mediaco
                                        I guess it's a fringe benefit that it will also gradually diminish the availability of reliable information, thus increasing the value of such information in our brave dystopian future. The motto can be "They who can afford books get to possess true knowledge."

                                        S This user is from outside of this forum
                                        S This user is from outside of this forum
                                        shadsterling@mastodon.social
                                        wrote sidst redigeret af
                                        #30

                                        @GerardThornley @Epic_Null @Bredroll @JPummil @404mediaco everything old is new again

                                        1 Reply Last reply
                                        0
                                        • ids1024@mathstodon.xyzI This user is from outside of this forum
                                          ids1024@mathstodon.xyzI This user is from outside of this forum
                                          ids1024@mathstodon.xyz
                                          wrote sidst redigeret af
                                          #31

                                          @hakfoo @404mediaco Yeah. I'm not an expert in how training these models works, but I'd expect adding more input (to any already wildly massive set of training data) of old books that are inaccurate, outdated, or otherwise not generally of interest to modern readers will at best be a *very* marginal improvement to the model, and at worst bloat the model while make it produce more outdated and irrelevant information.

                                          1 Reply Last reply
                                          0
                                          Svar
                                          • Svar som emne
                                          Login for at svare
                                          • Ældste til nyeste
                                          • Nyeste til ældste
                                          • Most Votes


                                          • Log ind

                                          • Har du ikke en konto? Tilmeld

                                          • Login or register to search.
                                          Powered by NodeBB Contributors
                                          Graciously hosted by data.coop
                                          • First post
                                            Last post
                                          0
                                          • Hjem
                                          • Seneste
                                          • Etiketter
                                          • Populære
                                          • Verden
                                          • Bruger
                                          • Grupper