Skip to content
  • Hjem
  • Seneste
  • Etiketter
  • Populære
  • Verden
  • Bruger
  • Grupper
Temaer
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Kollaps
FARVEL BIG TECH
  1. Forside
  2. Ikke-kategoriseret
  3. We've been getting a lot of questions about how the Internet Archive digitizes books.

We've been getting a lot of questions about how the Internet Archive digitizes books.

Planlagt Fastgjort Låst Flyttet Ikke-kategoriseret
58 Indlæg 42 Posters 101 Visninger
  • Ældste til nyeste
  • Nyeste til ældste
  • Most Votes
Svar
  • Svar som emne
Login for at svare
Denne tråd er blevet slettet. Kun brugere med emne behandlings privilegier kan se den.
  • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

    We've been getting a lot of questions about how the Internet Archive digitizes books.

    The short answer: page by page, by hand.

    You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

    Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

    cornelia@plasmatrap.comC This user is from outside of this forum
    cornelia@plasmatrap.comC This user is from outside of this forum
    cornelia@plasmatrap.com
    wrote sidst redigeret af
    #32

    @internetarchive@mastodon.archive.org so cool

    1 Reply Last reply
    0
    • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

      We've been getting a lot of questions about how the Internet Archive digitizes books.

      The short answer: page by page, by hand.

      You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

      Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

      joscelyntransient@chaosfem.twJ This user is from outside of this forum
      joscelyntransient@chaosfem.twJ This user is from outside of this forum
      joscelyntransient@chaosfem.tw
      wrote sidst redigeret af
      #33

      @internetarchive as someone who has had to scan books many times in grad school…that machine is beautiful and I am immensely envious I never got to use one

      1 Reply Last reply
      0
      • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

        We've been getting a lot of questions about how the Internet Archive digitizes books.

        The short answer: page by page, by hand.

        You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

        Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

        utf_7@mastodon.socialU This user is from outside of this forum
        utf_7@mastodon.socialU This user is from outside of this forum
        utf_7@mastodon.social
        wrote sidst redigeret af
        #34

        @internetarchive

        weird, that there is no destructive way to scan automatically

        1 Reply Last reply
        0
        • gloriouscow@oldbytes.spaceG gloriouscow@oldbytes.space

          @internetarchive

          Amazing work but please please please stop using MRC in PDFs.

          hammerwell@troet.cafeH This user is from outside of this forum
          hammerwell@troet.cafeH This user is from outside of this forum
          hammerwell@troet.cafe
          wrote sidst redigeret af
          #35

          @gloriouscow @internetarchive Whats MRC? Shouldn't those pdf be in pdf/A? Preferably in pfd/A-2a.

          gloriouscow@oldbytes.spaceG 1 Reply Last reply
          0
          • hammerwell@troet.cafeH hammerwell@troet.cafe

            @gloriouscow @internetarchive Whats MRC? Shouldn't those pdf be in pdf/A? Preferably in pfd/A-2a.

            gloriouscow@oldbytes.spaceG This user is from outside of this forum
            gloriouscow@oldbytes.spaceG This user is from outside of this forum
            gloriouscow@oldbytes.space
            wrote sidst redigeret af
            #36

            @Hammerwell @internetarchive

            MRC is a compression technique ("Mixed Raster Content").

            It basically lifts the text off a graphic onto its own layer, stores it as 1bpp which can be compressed as such, then the underlying graphic, now complete with text-shaped holes in it, can then be compressed with jpeg or something.

            the problem is whatever workflow IA uses to do this dials compression to all the way to maximum which results in 1bpp text over a smeary, unreadable mess

            gloriouscow@oldbytes.spaceG 1 Reply Last reply
            0
            • gloriouscow@oldbytes.spaceG gloriouscow@oldbytes.space

              @Hammerwell @internetarchive

              MRC is a compression technique ("Mixed Raster Content").

              It basically lifts the text off a graphic onto its own layer, stores it as 1bpp which can be compressed as such, then the underlying graphic, now complete with text-shaped holes in it, can then be compressed with jpeg or something.

              the problem is whatever workflow IA uses to do this dials compression to all the way to maximum which results in 1bpp text over a smeary, unreadable mess

              gloriouscow@oldbytes.spaceG This user is from outside of this forum
              gloriouscow@oldbytes.spaceG This user is from outside of this forum
              gloriouscow@oldbytes.space
              wrote sidst redigeret af
              #37

              @Hammerwell @internetarchive

              here's an actual example.

              the left is from a Google Books scan. The right is from IA:

              imo this severely undermines IA's mission to preserve our history and there's a metric shit ton of content on IA that unfortunately probably needs to be reprocessed if not rescanned.

              gloriouscow@oldbytes.spaceG 1 Reply Last reply
              0
              • gloriouscow@oldbytes.spaceG gloriouscow@oldbytes.space

                @Hammerwell @internetarchive

                here's an actual example.

                the left is from a Google Books scan. The right is from IA:

                imo this severely undermines IA's mission to preserve our history and there's a metric shit ton of content on IA that unfortunately probably needs to be reprocessed if not rescanned.

                gloriouscow@oldbytes.spaceG This user is from outside of this forum
                gloriouscow@oldbytes.spaceG This user is from outside of this forum
                gloriouscow@oldbytes.space
                wrote sidst redigeret af
                #38

                @Hammerwell @internetarchive

                notice that the "D" in "dealer" has become an O because MRC has the "Xerox Bug"

                https://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres_are_switching_written_numbers_when_scanning

                gloriouscow@oldbytes.spaceG 1 Reply Last reply
                0
                • gloriouscow@oldbytes.spaceG gloriouscow@oldbytes.space

                  @Hammerwell @internetarchive

                  notice that the "D" in "dealer" has become an O because MRC has the "Xerox Bug"

                  https://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres_are_switching_written_numbers_when_scanning

                  gloriouscow@oldbytes.spaceG This user is from outside of this forum
                  gloriouscow@oldbytes.spaceG This user is from outside of this forum
                  gloriouscow@oldbytes.space
                  wrote sidst redigeret af
                  #39

                  @Hammerwell @internetarchive

                  so not only are the graphics ruined, but the text is ruined too. not a great situation.

                  hammerwell@troet.cafeH 1 Reply Last reply
                  0
                  • gloriouscow@oldbytes.spaceG gloriouscow@oldbytes.space

                    @Hammerwell @internetarchive

                    so not only are the graphics ruined, but the text is ruined too. not a great situation.

                    hammerwell@troet.cafeH This user is from outside of this forum
                    hammerwell@troet.cafeH This user is from outside of this forum
                    hammerwell@troet.cafe
                    wrote sidst redigeret af
                    #40

                    @gloriouscow @internetarchive Thanks. OCR should only be used as invisible 2nd layer. pdf/A shouldn't allow this. Version 1 had recognition errors which were corrected in v2. 3 and 4 allow for extra content, which is not advisable since it's unclear if it's accessible in the future. 2a should be choosen because of the accessibility requirement. Screen readers can read them properly.

                    textfiles@mastodon.archive.orgT 1 Reply Last reply
                    0
                    • thorsted@digipres.clubT thorsted@digipres.club

                      @gloriouscow @internetarchive I second this request.

                      jj@types.plJ This user is from outside of this forum
                      jj@types.plJ This user is from outside of this forum
                      jj@types.pl
                      wrote sidst redigeret af
                      #41

                      @Thorsted @gloriouscow @internetarchive i am of a mixed opinion here. MRC is good for archival but definitely sucks for the end user. it would be very nice if user-facing PDFs could be just the mask + get jbig2ified + get an OCR layer... but perhaps that is a task better suited for shadow libraries

                      gloriouscow@oldbytes.spaceG 1 Reply Last reply
                      0
                      • jj@types.plJ jj@types.pl

                        @Thorsted @gloriouscow @internetarchive i am of a mixed opinion here. MRC is good for archival but definitely sucks for the end user. it would be very nice if user-facing PDFs could be just the mask + get jbig2ified + get an OCR layer... but perhaps that is a task better suited for shadow libraries

                        gloriouscow@oldbytes.spaceG This user is from outside of this forum
                        gloriouscow@oldbytes.spaceG This user is from outside of this forum
                        gloriouscow@oldbytes.space
                        wrote sidst redigeret af
                        #42

                        @jj @Thorsted @internetarchive see my posts later. It sucks for archival too.

                        jj@types.plJ thorsted@digipres.clubT 2 Replies Last reply
                        0
                        • gloriouscow@oldbytes.spaceG gloriouscow@oldbytes.space

                          @jj @Thorsted @internetarchive see my posts later. It sucks for archival too.

                          jj@types.plJ This user is from outside of this forum
                          jj@types.plJ This user is from outside of this forum
                          jj@types.pl
                          wrote sidst redigeret af
                          #43

                          @gloriouscow @Thorsted @internetarchive i saw them, yeah -- this is not fundamental to MRC, no? just the internet archive's particular use of it. were compression not at max it should be ok

                          thorsted@digipres.clubT gloriouscow@oldbytes.spaceG 2 Replies Last reply
                          0
                          • gloriouscow@oldbytes.spaceG gloriouscow@oldbytes.space

                            @jj @Thorsted @internetarchive see my posts later. It sucks for archival too.

                            thorsted@digipres.clubT This user is from outside of this forum
                            thorsted@digipres.clubT This user is from outside of this forum
                            thorsted@digipres.club
                            wrote sidst redigeret af
                            #44

                            @gloriouscow @jj @internetarchive agreed. Terrible for archiving.

                            1 Reply Last reply
                            0
                            • nazokiyoubinbou@urusai.socialN nazokiyoubinbou@urusai.social

                              @internetarchive Oh wow. That's a pretty neat scanner. And I find it impossible to believe it's more efficient to rip books apart when a thing like this exists... They just didn't care enough to look.

                              marcel@waldvogel.familyM This user is from outside of this forum
                              marcel@waldvogel.familyM This user is from outside of this forum
                              marcel@waldvogel.family
                              wrote sidst redigeret af
                              #45

                              @nazokiyoubinbou @internetarchive
                              The other process is several times faster and tries to keep copyright lawyers at bay for some time. Win-win for Amazon etc. …
                              https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/

                              nazokiyoubinbou@urusai.socialN 1 Reply Last reply
                              0
                              • jj@types.plJ jj@types.pl

                                @gloriouscow @Thorsted @internetarchive i saw them, yeah -- this is not fundamental to MRC, no? just the internet archive's particular use of it. were compression not at max it should be ok

                                thorsted@digipres.clubT This user is from outside of this forum
                                thorsted@digipres.clubT This user is from outside of this forum
                                thorsted@digipres.club
                                wrote sidst redigeret af
                                #46

                                @jj @gloriouscow @internetarchive my issue is when looking at the metadata there is an excessive amount of images to process.

                                jj@types.plJ 1 Reply Last reply
                                0
                                • jj@types.plJ jj@types.pl

                                  @gloriouscow @Thorsted @internetarchive i saw them, yeah -- this is not fundamental to MRC, no? just the internet archive's particular use of it. were compression not at max it should be ok

                                  gloriouscow@oldbytes.spaceG This user is from outside of this forum
                                  gloriouscow@oldbytes.spaceG This user is from outside of this forum
                                  gloriouscow@oldbytes.space
                                  wrote sidst redigeret af
                                  #47

                                  @jj @Thorsted @internetarchive I just know I've never seen a jpeg change a letter

                                  jj@types.plJ 1 Reply Last reply
                                  0
                                  • marcel@waldvogel.familyM marcel@waldvogel.family

                                    @nazokiyoubinbou @internetarchive
                                    The other process is several times faster and tries to keep copyright lawyers at bay for some time. Win-win for Amazon etc. …
                                    https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/

                                    nazokiyoubinbou@urusai.socialN This user is from outside of this forum
                                    nazokiyoubinbou@urusai.socialN This user is from outside of this forum
                                    nazokiyoubinbou@urusai.social
                                    wrote sidst redigeret af
                                    #48

                                    @marcel @internetarchive I think actually ripping pages out takes longer than what I'm seeing here. Unless they have some super exact process to machine cut the books open without cutting out any letters, but what I had heard in the past was they employed people to tear them up and feed the pages in.

                                    I mean just look at that video. Flip, press, flip, press, flip, press. It's fast and accurate.

                                    But yes, they destroy them on purpose, which was... kind of my point. They could shred/etc even if they preserved them. Though I'm not clear why they can't give them to a charity or something (tax break!)

                                    BTW you shared a paywall article. I've already heard about that though and was specifically referring to it as an example, but if I hadn't that wouldn't be a particularly helpful link...

                                    marcel@waldvogel.familyM 1 Reply Last reply
                                    0
                                    • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

                                      We've been getting a lot of questions about how the Internet Archive digitizes books.

                                      The short answer: page by page, by hand.

                                      You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

                                      Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

                                      deathkitten@firetribe.orgD This user is from outside of this forum
                                      deathkitten@firetribe.orgD This user is from outside of this forum
                                      deathkitten@firetribe.org
                                      wrote sidst redigeret af
                                      #49

                                      @internetarchive@mastodon.archive.org As someone who used to work in a print shop, and has done my share scanning stuff from books, I desperately wish we'd had a scanner like that for the book copy jobs I did. Laying books open on a flatbed scanner is so tedious, and the repetitive motion of having to pick up the book to turn the page between each scan is exhausting. This would still be tedious, but a little easier only having to turn the page.

                                      You can tell this machine was designed by someone with care for both the book and the operator.

                                      N 1 Reply Last reply
                                      0
                                      • thorsted@digipres.clubT thorsted@digipres.club

                                        @jj @gloriouscow @internetarchive my issue is when looking at the metadata there is an excessive amount of images to process.

                                        jj@types.plJ This user is from outside of this forum
                                        jj@types.plJ This user is from outside of this forum
                                        jj@types.pl
                                        wrote sidst redigeret af
                                        #50

                                        @Thorsted @gloriouscow @internetarchive this is not entirely my experience. for my
                                        current work i'm looking at a bunch of (IA) MRC-encoded documents. pdfimages dumps the foreground/background/masks which gives you very clean pages of text (as images) upon inverting the masks. it's three images per page but this is Fine

                                        what i do find difficult to do though is recovering inline images, since MRC (or at least IA's use of it) does not attempt to extract those at all

                                        1 Reply Last reply
                                        0
                                        • gloriouscow@oldbytes.spaceG gloriouscow@oldbytes.space

                                          @jj @Thorsted @internetarchive I just know I've never seen a jpeg change a letter

                                          jj@types.plJ This user is from outside of this forum
                                          jj@types.plJ This user is from outside of this forum
                                          jj@types.pl
                                          wrote sidst redigeret af
                                          #51

                                          @gloriouscow @Thorsted @internetarchive fair point!

                                          1 Reply Last reply
                                          0
                                          Svar
                                          • Svar som emne
                                          Login for at svare
                                          • Ældste til nyeste
                                          • Nyeste til ældste
                                          • Most Votes


                                          • Log ind

                                          • Har du ikke en konto? Tilmeld

                                          • Login or register to search.
                                          Powered by NodeBB Contributors
                                          Graciously hosted by data.coop
                                          • First post
                                            Last post
                                          0
                                          • Hjem
                                          • Seneste
                                          • Etiketter
                                          • Populære
                                          • Verden
                                          • Bruger
                                          • Grupper