We've been getting a lot of questions about how the Internet Archive digitizes books.
-
We've been getting a lot of questions about how the Internet Archive digitizes books.
The short answer: page by page, by hand.
You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.
Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/
The difference between the Internet Archive and the AI jack offs, is that Brewster loves books.
-
We've been getting a lot of questions about how the Internet Archive digitizes books.
The short answer: page by page, by hand.
You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.
Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/
AI bros destroy books because it's cheaper than scanning them non-destructively.
AI bros steal potable water from municipalities because it's cheaper than aircooling servers.Something something sharpen the (metaphorical) guillotines, already.
-
We've been getting a lot of questions about how the Internet Archive digitizes books.
The short answer: page by page, by hand.
You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.
Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/
@internetarchive Amazing! I operated a scanner, and yeah, it requires patience and a delicate touch -- especially when dealing with historic documents! Thank you for sharing this! (And yeah, watching the scanner operator footage was so #ASMR for me. Bring that back? Please?)
-
We've been getting a lot of questions about how the Internet Archive digitizes books.
The short answer: page by page, by hand.
You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.
Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/
@internetarchive and not destroying them in the process for shady legal reasons... amazing.
Thanks for your work !!





-
We've been getting a lot of questions about how the Internet Archive digitizes books.
The short answer: page by page, by hand.
You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.
Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/
@internetarchive
Seriously, fuck the muskrat for this and cease-and-desist-ing nitter. -
We've been getting a lot of questions about how the Internet Archive digitizes books.
The short answer: page by page, by hand.
You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.
Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/
Because guillotines weren’t meant for books!
-
We've been getting a lot of questions about how the Internet Archive digitizes books.
The short answer: page by page, by hand.
You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.
Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/
@internetarchive gods bless all of you for the work you do
-
AI bros destroy books because it's cheaper than scanning them non-destructively.
AI bros steal potable water from municipalities because it's cheaper than aircooling servers.Something something sharpen the (metaphorical) guillotines, already.
@rl_dane @internetarchive
I guess they have a point there. We should start thinking like them. For humanity it would be cheaper to use guilliotines... on biollionares 
-
We've been getting a lot of questions about how the Internet Archive digitizes books.
The short answer: page by page, by hand.
You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.
Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/
@internetarchive
And this is how the bad guys do it: https://www.forbes.com/sites/maryroeloffs/2026/08/17/ai-companies-are-buying-and-destroying-antique-books-heres-why/ -
We've been getting a lot of questions about how the Internet Archive digitizes books.
The short answer: page by page, by hand.
You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.
Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/
There's something charming about the meticulous process of scanning books by hand; it feels like a nod to the care that goes into preserving stories for future readers.
-
We've been getting a lot of questions about how the Internet Archive digitizes books.
The short answer: page by page, by hand.
You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.
Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/
@internetarchive Oh wow. That's a pretty neat scanner. And I find it impossible to believe it's more efficient to rip books apart when a thing like this exists... They just didn't care enough to look.
-
We've been getting a lot of questions about how the Internet Archive digitizes books.
The short answer: page by page, by hand.
You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.
Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/
@internetarchive the patience of an absolute saint!
-
We've been getting a lot of questions about how the Internet Archive digitizes books.
The short answer: page by page, by hand.
You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.
Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/
-
We've been getting a lot of questions about how the Internet Archive digitizes books.
The short answer: page by page, by hand.
You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.
Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/
@internetarchive as someone who has had to scan books many times in grad school…that machine is beautiful and I am immensely envious I never got to use one
-
We've been getting a lot of questions about how the Internet Archive digitizes books.
The short answer: page by page, by hand.
You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.
Meet Eliza, and learn how we scan books: https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/
weird, that there is no destructive way to scan automatically
-
Amazing work but please please please stop using MRC in PDFs.
@gloriouscow @internetarchive Whats MRC? Shouldn't those pdf be in pdf/A? Preferably in pfd/A-2a.
-
@gloriouscow @internetarchive Whats MRC? Shouldn't those pdf be in pdf/A? Preferably in pfd/A-2a.
MRC is a compression technique ("Mixed Raster Content").
It basically lifts the text off a graphic onto its own layer, stores it as 1bpp which can be compressed as such, then the underlying graphic, now complete with text-shaped holes in it, can then be compressed with jpeg or something.
the problem is whatever workflow IA uses to do this dials compression to all the way to maximum which results in 1bpp text over a smeary, unreadable mess
-
MRC is a compression technique ("Mixed Raster Content").
It basically lifts the text off a graphic onto its own layer, stores it as 1bpp which can be compressed as such, then the underlying graphic, now complete with text-shaped holes in it, can then be compressed with jpeg or something.
the problem is whatever workflow IA uses to do this dials compression to all the way to maximum which results in 1bpp text over a smeary, unreadable mess
here's an actual example.
the left is from a Google Books scan. The right is from IA:
imo this severely undermines IA's mission to preserve our history and there's a metric shit ton of content on IA that unfortunately probably needs to be reprocessed if not rescanned.
-
here's an actual example.
the left is from a Google Books scan. The right is from IA:
imo this severely undermines IA's mission to preserve our history and there's a metric shit ton of content on IA that unfortunately probably needs to be reprocessed if not rescanned.
notice that the "D" in "dealer" has become an O because MRC has the "Xerox Bug"
-
notice that the "D" in "dealer" has become an O because MRC has the "Xerox Bug"
so not only are the graphics ruined, but the text is ruined too. not a great situation.