Skip to content
  • Hjem
  • Seneste
  • Etiketter
  • Populære
  • Verden
  • Bruger
  • Grupper
Temaer
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Kollaps
FARVEL BIG TECH
  1. Forside
  2. Ikke-kategoriseret
  3. We've been getting a lot of questions about how the Internet Archive digitizes books.

We've been getting a lot of questions about how the Internet Archive digitizes books.

Planlagt Fastgjort Låst Flyttet Ikke-kategoriseret
58 Indlæg 42 Posters 100 Visninger
  • Ældste til nyeste
  • Nyeste til ældste
  • Most Votes
Svar
  • Svar som emne
Login for at svare
Denne tråd er blevet slettet. Kun brugere med emne behandlings privilegier kan se den.
  • gloriouscow@oldbytes.spaceG gloriouscow@oldbytes.space

    @jj @Thorsted @internetarchive see my posts later. It sucks for archival too.

    jj@types.plJ This user is from outside of this forum
    jj@types.plJ This user is from outside of this forum
    jj@types.pl
    wrote sidst redigeret af
    #43

    @gloriouscow @Thorsted @internetarchive i saw them, yeah -- this is not fundamental to MRC, no? just the internet archive's particular use of it. were compression not at max it should be ok

    thorsted@digipres.clubT gloriouscow@oldbytes.spaceG 2 Replies Last reply
    0
    • gloriouscow@oldbytes.spaceG gloriouscow@oldbytes.space

      @jj @Thorsted @internetarchive see my posts later. It sucks for archival too.

      thorsted@digipres.clubT This user is from outside of this forum
      thorsted@digipres.clubT This user is from outside of this forum
      thorsted@digipres.club
      wrote sidst redigeret af
      #44

      @gloriouscow @jj @internetarchive agreed. Terrible for archiving.

      1 Reply Last reply
      0
      • nazokiyoubinbou@urusai.socialN nazokiyoubinbou@urusai.social

        @internetarchive Oh wow. That's a pretty neat scanner. And I find it impossible to believe it's more efficient to rip books apart when a thing like this exists... They just didn't care enough to look.

        marcel@waldvogel.familyM This user is from outside of this forum
        marcel@waldvogel.familyM This user is from outside of this forum
        marcel@waldvogel.family
        wrote sidst redigeret af
        #45

        @nazokiyoubinbou @internetarchive
        The other process is several times faster and tries to keep copyright lawyers at bay for some time. Win-win for Amazon etc. …
        https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/

        nazokiyoubinbou@urusai.socialN 1 Reply Last reply
        0
        • jj@types.plJ jj@types.pl

          @gloriouscow @Thorsted @internetarchive i saw them, yeah -- this is not fundamental to MRC, no? just the internet archive's particular use of it. were compression not at max it should be ok

          thorsted@digipres.clubT This user is from outside of this forum
          thorsted@digipres.clubT This user is from outside of this forum
          thorsted@digipres.club
          wrote sidst redigeret af
          #46

          @jj @gloriouscow @internetarchive my issue is when looking at the metadata there is an excessive amount of images to process.

          jj@types.plJ 1 Reply Last reply
          0
          • jj@types.plJ jj@types.pl

            @gloriouscow @Thorsted @internetarchive i saw them, yeah -- this is not fundamental to MRC, no? just the internet archive's particular use of it. were compression not at max it should be ok

            gloriouscow@oldbytes.spaceG This user is from outside of this forum
            gloriouscow@oldbytes.spaceG This user is from outside of this forum
            gloriouscow@oldbytes.space
            wrote sidst redigeret af
            #47

            @jj @Thorsted @internetarchive I just know I've never seen a jpeg change a letter

            jj@types.plJ 1 Reply Last reply
            0
            • marcel@waldvogel.familyM marcel@waldvogel.family

              @nazokiyoubinbou @internetarchive
              The other process is several times faster and tries to keep copyright lawyers at bay for some time. Win-win for Amazon etc. …
              https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/

              nazokiyoubinbou@urusai.socialN This user is from outside of this forum
              nazokiyoubinbou@urusai.socialN This user is from outside of this forum
              nazokiyoubinbou@urusai.social
              wrote sidst redigeret af
              #48

              @marcel @internetarchive I think actually ripping pages out takes longer than what I'm seeing here. Unless they have some super exact process to machine cut the books open without cutting out any letters, but what I had heard in the past was they employed people to tear them up and feed the pages in.

              I mean just look at that video. Flip, press, flip, press, flip, press. It's fast and accurate.

              But yes, they destroy them on purpose, which was... kind of my point. They could shred/etc even if they preserved them. Though I'm not clear why they can't give them to a charity or something (tax break!)

              BTW you shared a paywall article. I've already heard about that though and was specifically referring to it as an example, but if I hadn't that wouldn't be a particularly helpful link...

              marcel@waldvogel.familyM 1 Reply Last reply
              0
              • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

                We've been getting a lot of questions about how the Internet Archive digitizes books.

                The short answer: page by page, by hand.

                You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

                Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

                deathkitten@firetribe.orgD This user is from outside of this forum
                deathkitten@firetribe.orgD This user is from outside of this forum
                deathkitten@firetribe.org
                wrote sidst redigeret af
                #49

                @internetarchive@mastodon.archive.org As someone who used to work in a print shop, and has done my share scanning stuff from books, I desperately wish we'd had a scanner like that for the book copy jobs I did. Laying books open on a flatbed scanner is so tedious, and the repetitive motion of having to pick up the book to turn the page between each scan is exhausting. This would still be tedious, but a little easier only having to turn the page.

                You can tell this machine was designed by someone with care for both the book and the operator.

                N 1 Reply Last reply
                0
                • thorsted@digipres.clubT thorsted@digipres.club

                  @jj @gloriouscow @internetarchive my issue is when looking at the metadata there is an excessive amount of images to process.

                  jj@types.plJ This user is from outside of this forum
                  jj@types.plJ This user is from outside of this forum
                  jj@types.pl
                  wrote sidst redigeret af
                  #50

                  @Thorsted @gloriouscow @internetarchive this is not entirely my experience. for my
                  current work i'm looking at a bunch of (IA) MRC-encoded documents. pdfimages dumps the foreground/background/masks which gives you very clean pages of text (as images) upon inverting the masks. it's three images per page but this is Fine

                  what i do find difficult to do though is recovering inline images, since MRC (or at least IA's use of it) does not attempt to extract those at all

                  1 Reply Last reply
                  0
                  • gloriouscow@oldbytes.spaceG gloriouscow@oldbytes.space

                    @jj @Thorsted @internetarchive I just know I've never seen a jpeg change a letter

                    jj@types.plJ This user is from outside of this forum
                    jj@types.plJ This user is from outside of this forum
                    jj@types.pl
                    wrote sidst redigeret af
                    #51

                    @gloriouscow @Thorsted @internetarchive fair point!

                    1 Reply Last reply
                    0
                    • nazokiyoubinbou@urusai.socialN nazokiyoubinbou@urusai.social

                      @marcel @internetarchive I think actually ripping pages out takes longer than what I'm seeing here. Unless they have some super exact process to machine cut the books open without cutting out any letters, but what I had heard in the past was they employed people to tear them up and feed the pages in.

                      I mean just look at that video. Flip, press, flip, press, flip, press. It's fast and accurate.

                      But yes, they destroy them on purpose, which was... kind of my point. They could shred/etc even if they preserved them. Though I'm not clear why they can't give them to a charity or something (tax break!)

                      BTW you shared a paywall article. I've already heard about that though and was specifically referring to it as an example, but if I hadn't that wouldn't be a particularly helpful link...

                      marcel@waldvogel.familyM This user is from outside of this forum
                      marcel@waldvogel.familyM This user is from outside of this forum
                      marcel@waldvogel.family
                      wrote sidst redigeret af
                      #52

                      @nazokiyoubinbou @internetarchive
                      From the descriptions I've seen, I imagine, they just cut the spine abd a few Millimeters off. There is some 2 cm of leeway there, exact cutting not required.

                      A quick product search revealed scanners with 3 sheets per second, no human operator required. From a "scan things fast and cheap" standpoint, that is much "better" than what's done here. (Of course, the IA approach has several benefits to the AI approach, but I don't need to point them out here.)

                      1 Reply Last reply
                      0
                      • hammerwell@troet.cafeH hammerwell@troet.cafe

                        @gloriouscow @internetarchive Thanks. OCR should only be used as invisible 2nd layer. pdf/A shouldn't allow this. Version 1 had recognition errors which were corrected in v2. 3 and 4 allow for extra content, which is not advisable since it's unclear if it's accessible in the future. 2a should be choosen because of the accessibility requirement. Screen readers can read them properly.

                        textfiles@mastodon.archive.orgT This user is from outside of this forum
                        textfiles@mastodon.archive.orgT This user is from outside of this forum
                        textfiles@mastodon.archive.org
                        wrote sidst redigeret af
                        #53

                        @Hammerwell @gloriouscow @internetarchive OK, well, the file you are linking to (In https://archive.org/details/PCMAG) was not scanned by Internet Archive. It was uploaded by a user 5 years ago, using whatever PDF approach they decided to take, or, more likely, was taken by someone THEY duplicated the scans from.

                        Here's another non-IA-done scan, with proper work.

                        textfiles@mastodon.archive.orgT 1 Reply Last reply
                        0
                        • deathkitten@firetribe.orgD deathkitten@firetribe.org

                          @internetarchive@mastodon.archive.org As someone who used to work in a print shop, and has done my share scanning stuff from books, I desperately wish we'd had a scanner like that for the book copy jobs I did. Laying books open on a flatbed scanner is so tedious, and the repetitive motion of having to pick up the book to turn the page between each scan is exhausting. This would still be tedious, but a little easier only having to turn the page.

                          You can tell this machine was designed by someone with care for both the book and the operator.

                          N This user is from outside of this forum
                          N This user is from outside of this forum
                          nightclaw@mastodon.social
                          wrote sidst redigeret af
                          #54

                          @deathkitten @internetarchive original design by Daniel Reetz: https://diybookscanner.org/archivist/

                          Sadly IA has not released design docs for their upgraded version, despite being asked repeatedly.

                          1 Reply Last reply
                          0
                          • textfiles@mastodon.archive.orgT textfiles@mastodon.archive.org

                            @Hammerwell @gloriouscow @internetarchive OK, well, the file you are linking to (In https://archive.org/details/PCMAG) was not scanned by Internet Archive. It was uploaded by a user 5 years ago, using whatever PDF approach they decided to take, or, more likely, was taken by someone THEY duplicated the scans from.

                            Here's another non-IA-done scan, with proper work.

                            textfiles@mastodon.archive.orgT This user is from outside of this forum
                            textfiles@mastodon.archive.orgT This user is from outside of this forum
                            textfiles@mastodon.archive.org
                            wrote sidst redigeret af
                            #55

                            @Hammerwell @gloriouscow @internetarchive And here is the .TIFF file that PDF was generated from.

                            https://archive.org/download/pc_magazine-1984_04_17/pc_magazine-1984_04_17.cbz/pc_magazine-1984_04_17%2Fpc_magazine-1984_04_17.294.tiff

                            textfiles@mastodon.archive.orgT 1 Reply Last reply
                            0
                            • textfiles@mastodon.archive.orgT textfiles@mastodon.archive.org

                              @Hammerwell @gloriouscow @internetarchive And here is the .TIFF file that PDF was generated from.

                              https://archive.org/download/pc_magazine-1984_04_17/pc_magazine-1984_04_17.cbz/pc_magazine-1984_04_17%2Fpc_magazine-1984_04_17.294.tiff

                              textfiles@mastodon.archive.orgT This user is from outside of this forum
                              textfiles@mastodon.archive.orgT This user is from outside of this forum
                              textfiles@mastodon.archive.org
                              wrote sidst redigeret af
                              #56

                              @Hammerwell @gloriouscow @internetarchive So in terms of "stop using MRC in PDFs", I guess the answer is "Done"

                              1 Reply Last reply
                              0
                              • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

                                We've been getting a lot of questions about how the Internet Archive digitizes books.

                                The short answer: page by page, by hand.

                                You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

                                Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

                                ketumbra@infosec.exchangeK This user is from outside of this forum
                                ketumbra@infosec.exchangeK This user is from outside of this forum
                                ketumbra@infosec.exchange
                                wrote sidst redigeret af
                                #57

                                @internetarchive surely any book printed in the last 35 Years was made digitally in the first place? Do publishers really delete those? Or is there some weird copyright loophole for scanning a printed version?

                                1 Reply Last reply
                                0
                                • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

                                  We've been getting a lot of questions about how the Internet Archive digitizes books.

                                  The short answer: page by page, by hand.

                                  You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

                                  Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

                                  riggbeck@mastodon.socialR This user is from outside of this forum
                                  riggbeck@mastodon.socialR This user is from outside of this forum
                                  riggbeck@mastodon.social
                                  wrote sidst redigeret af
                                  #58

                                  @internetarchive

                                  There's no need to destroy books in the process. The techbros just want to make it as cheap as possible, with the added benefit of taking books out of circulation so people need digital technology to access them. It's a retreat from the egalitarian printed book, back into a world where information is controlled by the wealthy. The reverse of progress.

                                  1 Reply Last reply
                                  0
                                  • jwcph@helvede.netJ jwcph@helvede.net shared this topic
                                  Svar
                                  • Svar som emne
                                  Login for at svare
                                  • Ældste til nyeste
                                  • Nyeste til ældste
                                  • Most Votes


                                  • Log ind

                                  • Har du ikke en konto? Tilmeld

                                  • Login or register to search.
                                  Powered by NodeBB Contributors
                                  Graciously hosted by data.coop
                                  • First post
                                    Last post
                                  0
                                  • Hjem
                                  • Seneste
                                  • Etiketter
                                  • Populære
                                  • Verden
                                  • Bruger
                                  • Grupper