Skip to content
  • Hjem
  • Seneste
  • Etiketter
  • Populære
  • Verden
  • Bruger
  • Grupper
Temaer
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Kollaps
FARVEL BIG TECH
  1. Forside
  2. Ikke-kategoriseret
  3. We've been getting a lot of questions about how the Internet Archive digitizes books.

We've been getting a lot of questions about how the Internet Archive digitizes books.

Planlagt Fastgjort Låst Flyttet Ikke-kategoriseret
58 Indlæg 42 Posters 99 Visninger
  • Ældste til nyeste
  • Nyeste til ældste
  • Most Votes
Svar
  • Svar som emne
Login for at svare
Denne tråd er blevet slettet. Kun brugere med emne behandlings privilegier kan se den.
  • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

    We've been getting a lot of questions about how the Internet Archive digitizes books.

    The short answer: page by page, by hand.

    You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

    Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

    rl_dane@polymaths.socialR This user is from outside of this forum
    rl_dane@polymaths.socialR This user is from outside of this forum
    rl_dane@polymaths.social
    wrote sidst redigeret af
    #21

    @internetarchive

    AI bros destroy books because it's cheaper than scanning them non-destructively.
    AI bros steal potable water from municipalities because it's cheaper than aircooling servers.

    Something something sharpen the (metaphorical) guillotines, already.

    sosa@mastodon.uyS 1 Reply Last reply
    0
    • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

      We've been getting a lot of questions about how the Internet Archive digitizes books.

      The short answer: page by page, by hand.

      You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

      Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

      doomsdayscw@kolektiva.socialD This user is from outside of this forum
      doomsdayscw@kolektiva.socialD This user is from outside of this forum
      doomsdayscw@kolektiva.social
      wrote sidst redigeret af
      #22

      @internetarchive Amazing! I operated a scanner, and yeah, it requires patience and a delicate touch -- especially when dealing with historic documents! Thank you for sharing this! (And yeah, watching the scanner operator footage was so #ASMR for me. Bring that back? Please?)

      1 Reply Last reply
      0
      • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

        We've been getting a lot of questions about how the Internet Archive digitizes books.

        The short answer: page by page, by hand.

        You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

        Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

        julianoe@mastodon.xyzJ This user is from outside of this forum
        julianoe@mastodon.xyzJ This user is from outside of this forum
        julianoe@mastodon.xyz
        wrote sidst redigeret af
        #23

        @internetarchive and not destroying them in the process for shady legal reasons... amazing.

        Thanks for your work !! 🙏🙏🙏🙏🙏

        1 Reply Last reply
        0
        • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

          We've been getting a lot of questions about how the Internet Archive digitizes books.

          The short answer: page by page, by hand.

          You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

          Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

          light@noc.socialL This user is from outside of this forum
          light@noc.socialL This user is from outside of this forum
          light@noc.social
          wrote sidst redigeret af
          #24

          @internetarchive
          Seriously, fuck the muskrat for this and cease-and-desist-ing nitter.

          1 Reply Last reply
          0
          • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

            We've been getting a lot of questions about how the Internet Archive digitizes books.

            The short answer: page by page, by hand.

            You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

            Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

            J This user is from outside of this forum
            J This user is from outside of this forum
            jackmexa4@mastodon.social
            wrote sidst redigeret af
            #25

            @internetarchive

            Because guillotines weren’t meant for books!

            1 Reply Last reply
            0
            • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

              We've been getting a lot of questions about how the Internet Archive digitizes books.

              The short answer: page by page, by hand.

              You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

              Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

              lamb@libretooth.grL This user is from outside of this forum
              lamb@libretooth.grL This user is from outside of this forum
              lamb@libretooth.gr
              wrote sidst redigeret af
              #26

              @internetarchive gods bless all of you for the work you do

              1 Reply Last reply
              0
              • rl_dane@polymaths.socialR rl_dane@polymaths.social

                @internetarchive

                AI bros destroy books because it's cheaper than scanning them non-destructively.
                AI bros steal potable water from municipalities because it's cheaper than aircooling servers.

                Something something sharpen the (metaphorical) guillotines, already.

                sosa@mastodon.uyS This user is from outside of this forum
                sosa@mastodon.uyS This user is from outside of this forum
                sosa@mastodon.uy
                wrote sidst redigeret af
                #27

                @rl_dane @internetarchive I guess they have a point there. We should start thinking like them. For humanity it would be cheaper to use guilliotines... on biollionares

                1 Reply Last reply
                0
                • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

                  We've been getting a lot of questions about how the Internet Archive digitizes books.

                  The short answer: page by page, by hand.

                  You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

                  Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

                  lrt_writes@mstdn.partyL This user is from outside of this forum
                  lrt_writes@mstdn.partyL This user is from outside of this forum
                  lrt_writes@mstdn.party
                  wrote sidst redigeret af
                  #28

                  @internetarchive
                  And this is how the bad guys do it: https://www.forbes.com/sites/maryroeloffs/2026/08/17/ai-companies-are-buying-and-destroying-antique-books-heres-why/

                  1 Reply Last reply
                  0
                  • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

                    We've been getting a lot of questions about how the Internet Archive digitizes books.

                    The short answer: page by page, by hand.

                    You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

                    Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

                    norawickham@mastodon.socialN This user is from outside of this forum
                    norawickham@mastodon.socialN This user is from outside of this forum
                    norawickham@mastodon.social
                    wrote sidst redigeret af
                    #29

                    There's something charming about the meticulous process of scanning books by hand; it feels like a nod to the care that goes into preserving stories for future readers.

                    1 Reply Last reply
                    0
                    • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

                      We've been getting a lot of questions about how the Internet Archive digitizes books.

                      The short answer: page by page, by hand.

                      You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

                      Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

                      nazokiyoubinbou@urusai.socialN This user is from outside of this forum
                      nazokiyoubinbou@urusai.socialN This user is from outside of this forum
                      nazokiyoubinbou@urusai.social
                      wrote sidst redigeret af
                      #30

                      @internetarchive Oh wow. That's a pretty neat scanner. And I find it impossible to believe it's more efficient to rip books apart when a thing like this exists... They just didn't care enough to look.

                      marcel@waldvogel.familyM 1 Reply Last reply
                      0
                      • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

                        We've been getting a lot of questions about how the Internet Archive digitizes books.

                        The short answer: page by page, by hand.

                        You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

                        Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

                        rainer@socialbc.caR This user is from outside of this forum
                        rainer@socialbc.caR This user is from outside of this forum
                        rainer@socialbc.ca
                        wrote sidst redigeret af
                        #31

                        @internetarchive the patience of an absolute saint!

                        1 Reply Last reply
                        0
                        • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

                          We've been getting a lot of questions about how the Internet Archive digitizes books.

                          The short answer: page by page, by hand.

                          You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

                          Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

                          cornelia@plasmatrap.comC This user is from outside of this forum
                          cornelia@plasmatrap.comC This user is from outside of this forum
                          cornelia@plasmatrap.com
                          wrote sidst redigeret af
                          #32

                          @internetarchive@mastodon.archive.org so cool

                          1 Reply Last reply
                          0
                          • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

                            We've been getting a lot of questions about how the Internet Archive digitizes books.

                            The short answer: page by page, by hand.

                            You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

                            Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

                            joscelyntransient@chaosfem.twJ This user is from outside of this forum
                            joscelyntransient@chaosfem.twJ This user is from outside of this forum
                            joscelyntransient@chaosfem.tw
                            wrote sidst redigeret af
                            #33

                            @internetarchive as someone who has had to scan books many times in grad school…that machine is beautiful and I am immensely envious I never got to use one

                            1 Reply Last reply
                            0
                            • internetarchive@mastodon.archive.orgI internetarchive@mastodon.archive.org

                              We've been getting a lot of questions about how the Internet Archive digitizes books.

                              The short answer: page by page, by hand.

                              You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.

                              Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/

                              utf_7@mastodon.socialU This user is from outside of this forum
                              utf_7@mastodon.socialU This user is from outside of this forum
                              utf_7@mastodon.social
                              wrote sidst redigeret af
                              #34

                              @internetarchive

                              weird, that there is no destructive way to scan automatically

                              1 Reply Last reply
                              0
                              • gloriouscow@oldbytes.spaceG gloriouscow@oldbytes.space

                                @internetarchive

                                Amazing work but please please please stop using MRC in PDFs.

                                hammerwell@troet.cafeH This user is from outside of this forum
                                hammerwell@troet.cafeH This user is from outside of this forum
                                hammerwell@troet.cafe
                                wrote sidst redigeret af
                                #35

                                @gloriouscow @internetarchive Whats MRC? Shouldn't those pdf be in pdf/A? Preferably in pfd/A-2a.

                                gloriouscow@oldbytes.spaceG 1 Reply Last reply
                                0
                                • hammerwell@troet.cafeH hammerwell@troet.cafe

                                  @gloriouscow @internetarchive Whats MRC? Shouldn't those pdf be in pdf/A? Preferably in pfd/A-2a.

                                  gloriouscow@oldbytes.spaceG This user is from outside of this forum
                                  gloriouscow@oldbytes.spaceG This user is from outside of this forum
                                  gloriouscow@oldbytes.space
                                  wrote sidst redigeret af
                                  #36

                                  @Hammerwell @internetarchive

                                  MRC is a compression technique ("Mixed Raster Content").

                                  It basically lifts the text off a graphic onto its own layer, stores it as 1bpp which can be compressed as such, then the underlying graphic, now complete with text-shaped holes in it, can then be compressed with jpeg or something.

                                  the problem is whatever workflow IA uses to do this dials compression to all the way to maximum which results in 1bpp text over a smeary, unreadable mess

                                  gloriouscow@oldbytes.spaceG 1 Reply Last reply
                                  0
                                  • gloriouscow@oldbytes.spaceG gloriouscow@oldbytes.space

                                    @Hammerwell @internetarchive

                                    MRC is a compression technique ("Mixed Raster Content").

                                    It basically lifts the text off a graphic onto its own layer, stores it as 1bpp which can be compressed as such, then the underlying graphic, now complete with text-shaped holes in it, can then be compressed with jpeg or something.

                                    the problem is whatever workflow IA uses to do this dials compression to all the way to maximum which results in 1bpp text over a smeary, unreadable mess

                                    gloriouscow@oldbytes.spaceG This user is from outside of this forum
                                    gloriouscow@oldbytes.spaceG This user is from outside of this forum
                                    gloriouscow@oldbytes.space
                                    wrote sidst redigeret af
                                    #37

                                    @Hammerwell @internetarchive

                                    here's an actual example.

                                    the left is from a Google Books scan. The right is from IA:

                                    imo this severely undermines IA's mission to preserve our history and there's a metric shit ton of content on IA that unfortunately probably needs to be reprocessed if not rescanned.

                                    gloriouscow@oldbytes.spaceG 1 Reply Last reply
                                    0
                                    • gloriouscow@oldbytes.spaceG gloriouscow@oldbytes.space

                                      @Hammerwell @internetarchive

                                      here's an actual example.

                                      the left is from a Google Books scan. The right is from IA:

                                      imo this severely undermines IA's mission to preserve our history and there's a metric shit ton of content on IA that unfortunately probably needs to be reprocessed if not rescanned.

                                      gloriouscow@oldbytes.spaceG This user is from outside of this forum
                                      gloriouscow@oldbytes.spaceG This user is from outside of this forum
                                      gloriouscow@oldbytes.space
                                      wrote sidst redigeret af
                                      #38

                                      @Hammerwell @internetarchive

                                      notice that the "D" in "dealer" has become an O because MRC has the "Xerox Bug"

                                      https://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres_are_switching_written_numbers_when_scanning

                                      gloriouscow@oldbytes.spaceG 1 Reply Last reply
                                      0
                                      • gloriouscow@oldbytes.spaceG gloriouscow@oldbytes.space

                                        @Hammerwell @internetarchive

                                        notice that the "D" in "dealer" has become an O because MRC has the "Xerox Bug"

                                        https://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres_are_switching_written_numbers_when_scanning

                                        gloriouscow@oldbytes.spaceG This user is from outside of this forum
                                        gloriouscow@oldbytes.spaceG This user is from outside of this forum
                                        gloriouscow@oldbytes.space
                                        wrote sidst redigeret af
                                        #39

                                        @Hammerwell @internetarchive

                                        so not only are the graphics ruined, but the text is ruined too. not a great situation.

                                        hammerwell@troet.cafeH 1 Reply Last reply
                                        0
                                        • gloriouscow@oldbytes.spaceG gloriouscow@oldbytes.space

                                          @Hammerwell @internetarchive

                                          so not only are the graphics ruined, but the text is ruined too. not a great situation.

                                          hammerwell@troet.cafeH This user is from outside of this forum
                                          hammerwell@troet.cafeH This user is from outside of this forum
                                          hammerwell@troet.cafe
                                          wrote sidst redigeret af
                                          #40

                                          @gloriouscow @internetarchive Thanks. OCR should only be used as invisible 2nd layer. pdf/A shouldn't allow this. Version 1 had recognition errors which were corrected in v2. 3 and 4 allow for extra content, which is not advisable since it's unclear if it's accessible in the future. 2a should be choosen because of the accessibility requirement. Screen readers can read them properly.

                                          textfiles@mastodon.archive.orgT 1 Reply Last reply
                                          0
                                          Svar
                                          • Svar som emne
                                          Login for at svare
                                          • Ældste til nyeste
                                          • Nyeste til ældste
                                          • Most Votes


                                          • Log ind

                                          • Har du ikke en konto? Tilmeld

                                          • Login or register to search.
                                          Powered by NodeBB Contributors
                                          Graciously hosted by data.coop
                                          • First post
                                            Last post
                                          0
                                          • Hjem
                                          • Seneste
                                          • Etiketter
                                          • Populære
                                          • Verden
                                          • Bruger
                                          • Grupper