<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books.]]></title><description><![CDATA[<p>We've been getting a lot of questions about how the Internet Archive digitizes books.</p><p>The short answer: page by page, by hand.</p><p>You may remember our viral 2021 video of Eliza Zhang scanning a book. That's still how we do it.</p><p>Meet Eliza, and learn how we scan books: <a href="https://blog.archive.org/2021/02/09/meet-eliza-zhang-book-scanner-and-viral-video-star/" rel="nofollow noopener"><span>https://</span><span>blog.archive.org/2021/02/09/me</span><span>et-eliza-zhang-book-scanner-and-viral-video-star/</span></a></p>]]></description><link>https://forum.fedi.dk/topic/0fedc07c-9500-488c-82b8-33dc9d6786f9/we-ve-been-getting-a-lot-of-questions-about-how-the-internet-archive-digitizes-books.</link><generator>RSS for Node</generator><lastBuildDate>Fri, 11 Sep 2026 18:05:28 GMT</lastBuildDate><atom:link href="https://forum.fedi.dk/topic/0fedc07c-9500-488c-82b8-33dc9d6786f9.rss" rel="self" type="application/rss+xml"/><pubDate>Tue, 25 Aug 2026 20:56:34 GMT</pubDate><ttl>60</ttl><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 06:35:11 GMT]]></title><description><![CDATA[<p><span><a href="/user/internetarchive%40mastodon.archive.org">@<span>internetarchive</span></a></span> </p><p>There's no need to destroy books in the process. The techbros just want to make it as cheap as possible, with the added benefit of taking books out of circulation so people need digital technology to access them. It's a retreat from the egalitarian printed book, back into a world where information is controlled by the wealthy. The reverse of progress.</p>]]></description><link>https://forum.fedi.dk/post/https://mastodon.social/users/riggbeck/statuses/117160418440955651</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://mastodon.social/users/riggbeck/statuses/117160418440955651</guid><dc:creator><![CDATA[riggbeck@mastodon.social]]></dc:creator><pubDate>Wed, 26 Aug 2026 06:35:11 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 06:09:49 GMT]]></title><description><![CDATA[<p><span><a href="/user/internetarchive%40mastodon.archive.org">@<span>internetarchive</span></a></span> surely any book printed in the last 35 Years was made digitally in the first place? Do publishers really delete those? Or is there some weird copyright loophole for scanning a printed version?</p>]]></description><link>https://forum.fedi.dk/post/https://infosec.exchange/users/ketumbra/statuses/117160318687179790</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://infosec.exchange/users/ketumbra/statuses/117160318687179790</guid><dc:creator><![CDATA[ketumbra@infosec.exchange]]></dc:creator><pubDate>Wed, 26 Aug 2026 06:09:49 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 05:30:39 GMT]]></title><description><![CDATA[<p><span><a href="/user/hammerwell%40troet.cafe">@<span>Hammerwell</span></a></span> <span><a href="https://oldbytes.space/@gloriouscow">@<span>gloriouscow</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org">@<span>internetarchive</span></a></span> So in terms of "stop using MRC in PDFs", I guess the answer is "Done"</p>]]></description><link>https://forum.fedi.dk/post/https://mastodon.archive.org/users/textfiles/statuses/117160164696371049</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://mastodon.archive.org/users/textfiles/statuses/117160164696371049</guid><dc:creator><![CDATA[textfiles@mastodon.archive.org]]></dc:creator><pubDate>Wed, 26 Aug 2026 05:30:39 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 05:29:46 GMT]]></title><description><![CDATA[<p><span><a href="/user/hammerwell%40troet.cafe">@<span>Hammerwell</span></a></span> <span><a href="https://oldbytes.space/@gloriouscow">@<span>gloriouscow</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org">@<span>internetarchive</span></a></span> And here is the .TIFF file that PDF was generated from.</p><p><a href="https://archive.org/download/pc_magazine-1984_04_17/pc_magazine-1984_04_17.cbz/pc_magazine-1984_04_17%2Fpc_magazine-1984_04_17.294.tiff" rel="nofollow noopener"><span>https://</span><span>archive.org/download/pc_magazi</span><span>ne-1984_04_17/pc_magazine-1984_04_17.cbz/pc_magazine-1984_04_17%2Fpc_magazine-1984_04_17.294.tiff</span></a></p>]]></description><link>https://forum.fedi.dk/post/https://mastodon.archive.org/users/textfiles/statuses/117160161220658707</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://mastodon.archive.org/users/textfiles/statuses/117160161220658707</guid><dc:creator><![CDATA[textfiles@mastodon.archive.org]]></dc:creator><pubDate>Wed, 26 Aug 2026 05:29:46 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 05:27:48 GMT]]></title><description><![CDATA[<p><span><a href="/user/deathkitten%40firetribe.org">@<span>deathkitten</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org">@<span>internetarchive</span></a></span> original design by Daniel Reetz: <a href="https://diybookscanner.org/archivist/" rel="nofollow noopener"><span>https://</span><span>diybookscanner.org/archivist/</span><span></span></a></p><p>Sadly IA has not released design docs for their upgraded version, despite being asked repeatedly.</p>]]></description><link>https://forum.fedi.dk/post/https://mastodon.social/users/nightclaw/statuses/117160153480848206</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://mastodon.social/users/nightclaw/statuses/117160153480848206</guid><dc:creator><![CDATA[nightclaw@mastodon.social]]></dc:creator><pubDate>Wed, 26 Aug 2026 05:27:48 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 05:27:27 GMT]]></title><description><![CDATA[<p><span><a href="/user/hammerwell%40troet.cafe">@<span>Hammerwell</span></a></span> <span><a href="https://oldbytes.space/@gloriouscow">@<span>gloriouscow</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org">@<span>internetarchive</span></a></span> OK, well, the file you are linking to (In <a href="https://archive.org/details/PCMAG" rel="nofollow noopener"><span>https://</span><span>archive.org/details/PCMAG</span><span></span></a>) was not scanned by Internet Archive. It was uploaded by a user 5 years ago, using whatever PDF approach they decided to take, or, more likely, was taken by someone THEY duplicated the scans from.</p><p>Here's another non-IA-done scan, with proper work.</p>]]></description><link>https://forum.fedi.dk/post/https://mastodon.archive.org/users/textfiles/statuses/117160152079931740</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://mastodon.archive.org/users/textfiles/statuses/117160152079931740</guid><dc:creator><![CDATA[textfiles@mastodon.archive.org]]></dc:creator><pubDate>Wed, 26 Aug 2026 05:27:27 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 05:27:17 GMT]]></title><description><![CDATA[<p><span><a href="/user/nazokiyoubinbou%40urusai.social">@<span>nazokiyoubinbou</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org">@<span>internetarchive</span></a></span> <br />From the descriptions I've seen, I imagine, they just cut the spine abd a few Millimeters off. There is some 2 cm of leeway there, exact cutting not required.</p><p>A quick product search revealed scanners with 3 sheets per second, no human operator required. From a "scan things fast and cheap" standpoint, that is much "better" than what's done here. (Of course, the IA approach has several benefits to the AI approach, but I don't need to point them out here.)</p>]]></description><link>https://forum.fedi.dk/post/https://waldvogel.family/users/marcel/statuses/117160151481078145</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://waldvogel.family/users/marcel/statuses/117160151481078145</guid><dc:creator><![CDATA[marcel@waldvogel.family]]></dc:creator><pubDate>Wed, 26 Aug 2026 05:27:17 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 04:50:47 GMT]]></title><description><![CDATA[<p><span><a href="https://oldbytes.space/@gloriouscow" rel="nofollow noopener">@<span>gloriouscow</span></a></span> <span><a href="https://digipres.club/@Thorsted" rel="nofollow noopener">@<span>Thorsted</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org" rel="nofollow noopener">@<span>internetarchive</span></a></span> fair point!</p>]]></description><link>https://forum.fedi.dk/post/https://types.pl/users/jj/statuses/117160007925504927</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://types.pl/users/jj/statuses/117160007925504927</guid><dc:creator><![CDATA[jj@types.pl]]></dc:creator><pubDate>Wed, 26 Aug 2026 04:50:47 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 04:49:59 GMT]]></title><description><![CDATA[<p><span><a href="https://digipres.club/@Thorsted" rel="nofollow noopener">@<span>Thorsted</span></a></span> <span><a href="https://oldbytes.space/@gloriouscow" rel="nofollow noopener">@<span>gloriouscow</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org" rel="nofollow noopener">@<span>internetarchive</span></a></span> this is not <em>entirely</em> my experience. for my<br />current work i'm looking at a bunch of (IA) MRC-encoded documents. <code>pdfimages</code> dumps the foreground/background/masks which gives you very clean pages of text (as images) upon inverting the masks. it's three images per page but this is Fine</p><p>what i do find difficult to do though is recovering inline images, since MRC (or at least IA's use of it) does not attempt to extract those at all</p>]]></description><link>https://forum.fedi.dk/post/https://types.pl/users/jj/statuses/117160004756197080</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://types.pl/users/jj/statuses/117160004756197080</guid><dc:creator><![CDATA[jj@types.pl]]></dc:creator><pubDate>Wed, 26 Aug 2026 04:49:59 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 04:46:58 GMT]]></title><description><![CDATA[<p><a href="/user/internetarchive%40mastodon.archive.org">@internetarchive@mastodon.archive.org</a><span> As someone who used to work in a print shop, and has done my share scanning stuff from books, I desperately wish we'd had a scanner like that for the book copy jobs I did. Laying books open on a flatbed scanner is so tedious, and the repetitive motion of having to pick up the book to turn the page between each scan is exhausting. This would still be tedious, but a little easier only having to turn the page.<br /><br />You can tell this machine was designed by someone with care for both the book and the operator.</span></p>]]></description><link>https://forum.fedi.dk/post/https://firetribe.org/notes/aqd6ez4y76</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://firetribe.org/notes/aqd6ez4y76</guid><dc:creator><![CDATA[deathkitten@firetribe.org]]></dc:creator><pubDate>Wed, 26 Aug 2026 04:46:58 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 04:45:43 GMT]]></title><description><![CDATA[<p><span><a href="/user/marcel%40waldvogel.family" rel="nofollow noopener">@<span>marcel</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org" rel="nofollow noopener">@<span>internetarchive</span></a></span> I think actually ripping pages out takes longer than what I'm seeing here.  Unless they have some super exact process to machine cut the books open without cutting out any letters, but what I had heard in the past was they employed people to tear them up and feed the pages in.</p><p>I mean just look at that video.  Flip, press, flip, press, flip, press.  It's <em>fast and accurate.</em></p><p>But yes, they destroy them on purpose, which was...  kind of my point.  They could shred/etc even if they preserved them.  Though I'm not clear why they can't give them to a charity or something (tax break!)</p><p>BTW you shared a paywall article.  I've already heard about that though and was specifically referring to it as an example, but if I hadn't that wouldn't be a particularly helpful link...</p>]]></description><link>https://forum.fedi.dk/post/https://urusai.social/users/nazokiyoubinbou/statuses/117159987999548285</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://urusai.social/users/nazokiyoubinbou/statuses/117159987999548285</guid><dc:creator><![CDATA[nazokiyoubinbou@urusai.social]]></dc:creator><pubDate>Wed, 26 Aug 2026 04:45:43 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 04:43:22 GMT]]></title><description><![CDATA[<p><span><a href="https://types.pl/@jj" rel="nofollow noopener">@<span>jj</span></a></span> <span><a href="https://digipres.club/@Thorsted" rel="nofollow noopener">@<span>Thorsted</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org" rel="nofollow noopener">@<span>internetarchive</span></a></span> I just know I've never seen a jpeg change a letter</p>]]></description><link>https://forum.fedi.dk/post/https://oldbytes.space/users/gloriouscow/statuses/117159978764951375</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://oldbytes.space/users/gloriouscow/statuses/117159978764951375</guid><dc:creator><![CDATA[gloriouscow@oldbytes.space]]></dc:creator><pubDate>Wed, 26 Aug 2026 04:43:22 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 04:43:14 GMT]]></title><description><![CDATA[<p><span><a href="https://types.pl/@jj">@<span>jj</span></a></span> <span><a href="https://oldbytes.space/@gloriouscow">@<span>gloriouscow</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org">@<span>internetarchive</span></a></span> my issue is when looking at the metadata there is an excessive amount of images to process.</p>]]></description><link>https://forum.fedi.dk/post/https://digipres.club/users/Thorsted/statuses/117159978224470930</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://digipres.club/users/Thorsted/statuses/117159978224470930</guid><dc:creator><![CDATA[thorsted@digipres.club]]></dc:creator><pubDate>Wed, 26 Aug 2026 04:43:14 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 04:42:49 GMT]]></title><description><![CDATA[<p><span><a href="/user/nazokiyoubinbou%40urusai.social">@<span>nazokiyoubinbou</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org">@<span>internetarchive</span></a></span> <br />The other process is several times faster and tries to keep copyright lawyers at bay for some time. Win-win for Amazon etc. …<br /><a href="https://www.404media.co/we-tracked-a-shipment-of-rare-books-it-ended-at-an-amazon-ai-training-facility/" rel="nofollow noopener"><span>https://www.</span><span>404media.co/we-tracked-a-shipm</span><span>ent-of-rare-books-it-ended-at-an-amazon-ai-training-facility/</span></a></p>]]></description><link>https://forum.fedi.dk/post/https://waldvogel.family/users/marcel/statuses/117159976576734086</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://waldvogel.family/users/marcel/statuses/117159976576734086</guid><dc:creator><![CDATA[marcel@waldvogel.family]]></dc:creator><pubDate>Wed, 26 Aug 2026 04:42:49 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 04:41:44 GMT]]></title><description><![CDATA[<p><span><a href="https://oldbytes.space/@gloriouscow">@<span>gloriouscow</span></a></span> <span><a href="https://types.pl/@jj">@<span>jj</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org">@<span>internetarchive</span></a></span> agreed. Terrible for archiving.</p>]]></description><link>https://forum.fedi.dk/post/https://digipres.club/users/Thorsted/statuses/117159972346262023</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://digipres.club/users/Thorsted/statuses/117159972346262023</guid><dc:creator><![CDATA[thorsted@digipres.club]]></dc:creator><pubDate>Wed, 26 Aug 2026 04:41:44 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 04:40:58 GMT]]></title><description><![CDATA[<p><span><a href="https://oldbytes.space/@gloriouscow" rel="nofollow noopener">@<span>gloriouscow</span></a></span> <span><a href="https://digipres.club/@Thorsted" rel="nofollow noopener">@<span>Thorsted</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org" rel="nofollow noopener">@<span>internetarchive</span></a></span> i saw them, yeah -- this is not fundamental to MRC, no? just the internet archive's particular use of it. were compression not at max it should be ok</p>]]></description><link>https://forum.fedi.dk/post/https://types.pl/users/jj/statuses/117159969312926711</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://types.pl/users/jj/statuses/117159969312926711</guid><dc:creator><![CDATA[jj@types.pl]]></dc:creator><pubDate>Wed, 26 Aug 2026 04:40:58 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 04:39:00 GMT]]></title><description><![CDATA[<p><span><a href="https://types.pl/@jj" rel="nofollow noopener">@<span>jj</span></a></span> <span><a href="https://digipres.club/@Thorsted" rel="nofollow noopener">@<span>Thorsted</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org" rel="nofollow noopener">@<span>internetarchive</span></a></span> see my posts later. It sucks for archival too.</p>]]></description><link>https://forum.fedi.dk/post/https://oldbytes.space/users/gloriouscow/statuses/117159961594928786</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://oldbytes.space/users/gloriouscow/statuses/117159961594928786</guid><dc:creator><![CDATA[gloriouscow@oldbytes.space]]></dc:creator><pubDate>Wed, 26 Aug 2026 04:39:00 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 04:35:55 GMT]]></title><description><![CDATA[<p><span><a href="https://digipres.club/@Thorsted" rel="nofollow noopener">@<span>Thorsted</span></a></span> <span><a href="https://oldbytes.space/@gloriouscow" rel="nofollow noopener">@<span>gloriouscow</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org" rel="nofollow noopener">@<span>internetarchive</span></a></span> i am of a mixed opinion here. MRC is good for archival but definitely sucks for the end user. it would be very nice if user-facing PDFs could be just the mask + get jbig2ified + get an OCR layer... but perhaps that is a task better suited for shadow libraries</p>]]></description><link>https://forum.fedi.dk/post/https://types.pl/users/jj/statuses/117159949486620796</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://types.pl/users/jj/statuses/117159949486620796</guid><dc:creator><![CDATA[jj@types.pl]]></dc:creator><pubDate>Wed, 26 Aug 2026 04:35:55 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 04:24:18 GMT]]></title><description><![CDATA[<p><span><a href="https://oldbytes.space/@gloriouscow">@<span>gloriouscow</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org">@<span>internetarchive</span></a></span> Thanks. OCR should only be used as invisible 2nd layer. pdf/A shouldn't allow this. Version 1 had recognition errors which were corrected in v2. 3 and 4 allow for extra content, which is not advisable since it's unclear if it's accessible in the future. 2a should be choosen because of the accessibility requirement. Screen readers can read them properly.</p>]]></description><link>https://forum.fedi.dk/post/https://troet.cafe/users/Hammerwell/statuses/117159903806572936</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://troet.cafe/users/Hammerwell/statuses/117159903806572936</guid><dc:creator><![CDATA[hammerwell@troet.cafe]]></dc:creator><pubDate>Wed, 26 Aug 2026 04:24:18 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 04:16:45 GMT]]></title><description><![CDATA[<p><span><a href="/user/hammerwell%40troet.cafe" rel="nofollow noopener">@<span>Hammerwell</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org" rel="nofollow noopener">@<span>internetarchive</span></a></span> </p><p>so not only are the graphics ruined, but the text is ruined too. not a great situation.</p>]]></description><link>https://forum.fedi.dk/post/https://oldbytes.space/users/gloriouscow/statuses/117159874129740326</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://oldbytes.space/users/gloriouscow/statuses/117159874129740326</guid><dc:creator><![CDATA[gloriouscow@oldbytes.space]]></dc:creator><pubDate>Wed, 26 Aug 2026 04:16:45 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 04:14:57 GMT]]></title><description><![CDATA[<p><span><a href="/user/hammerwell%40troet.cafe" rel="nofollow noopener">@<span>Hammerwell</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org" rel="nofollow noopener">@<span>internetarchive</span></a></span> </p><p>notice that the "D" in "dealer" has become an O because MRC has the "Xerox Bug"</p><p><a href="https://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres_are_switching_written_numbers_when_scanning" rel="nofollow noopener"><span>https://www.</span><span>dkriesel.com/en/blog/2013/0802</span><span>_xerox-workcentres_are_switching_written_numbers_when_scanning</span></a></p>]]></description><link>https://forum.fedi.dk/post/https://oldbytes.space/users/gloriouscow/statuses/117159867049662099</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://oldbytes.space/users/gloriouscow/statuses/117159867049662099</guid><dc:creator><![CDATA[gloriouscow@oldbytes.space]]></dc:creator><pubDate>Wed, 26 Aug 2026 04:14:57 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 04:13:03 GMT]]></title><description><![CDATA[<p><span><a href="/user/hammerwell%40troet.cafe" rel="nofollow noopener">@<span>Hammerwell</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org" rel="nofollow noopener">@<span>internetarchive</span></a></span> </p><p>here's an actual example.</p><p>the left is from a Google Books scan. The right is from IA:</p><p>imo this severely undermines IA's mission to preserve our history and there's a metric shit ton of content on IA that unfortunately probably needs to be reprocessed if not rescanned.</p>]]></description><link>https://forum.fedi.dk/post/https://oldbytes.space/users/gloriouscow/statuses/117159859527547012</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://oldbytes.space/users/gloriouscow/statuses/117159859527547012</guid><dc:creator><![CDATA[gloriouscow@oldbytes.space]]></dc:creator><pubDate>Wed, 26 Aug 2026 04:13:03 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 04:10:29 GMT]]></title><description><![CDATA[<p><span><a href="/user/hammerwell%40troet.cafe" rel="nofollow noopener">@<span>Hammerwell</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org" rel="nofollow noopener">@<span>internetarchive</span></a></span> </p><p>MRC is a compression technique ("Mixed Raster Content").</p><p>It basically lifts the text off a graphic onto its own layer, stores it as 1bpp which can be compressed as such, then the underlying graphic, now complete with text-shaped holes in it, can then be compressed with jpeg or something.</p><p>the problem is whatever workflow IA uses to do this dials compression to all the way to maximum which results in 1bpp text over a smeary, unreadable mess</p>]]></description><link>https://forum.fedi.dk/post/https://oldbytes.space/users/gloriouscow/statuses/117159849482028151</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://oldbytes.space/users/gloriouscow/statuses/117159849482028151</guid><dc:creator><![CDATA[gloriouscow@oldbytes.space]]></dc:creator><pubDate>Wed, 26 Aug 2026 04:10:29 GMT</pubDate></item><item><title><![CDATA[Reply to We&#x27;ve been getting a lot of questions about how the Internet Archive digitizes books. on Wed, 26 Aug 2026 04:06:40 GMT]]></title><description><![CDATA[<p><span><a href="https://oldbytes.space/@gloriouscow">@<span>gloriouscow</span></a></span> <span><a href="/user/internetarchive%40mastodon.archive.org">@<span>internetarchive</span></a></span> Whats MRC? Shouldn't those pdf be in pdf/A? Preferably in pfd/A-2a.</p>]]></description><link>https://forum.fedi.dk/post/https://troet.cafe/users/Hammerwell/statuses/117159834441866728</link><guid isPermaLink="true">https://forum.fedi.dk/post/https://troet.cafe/users/Hammerwell/statuses/117159834441866728</guid><dc:creator><![CDATA[hammerwell@troet.cafe]]></dc:creator><pubDate>Wed, 26 Aug 2026 04:06:40 GMT</pubDate></item></channel></rss>