@gloriouscow @Thorsted @internetarchive fair point!
jj@types.pl
Indlæg
-
We've been getting a lot of questions about how the Internet Archive digitizes books. -
We've been getting a lot of questions about how the Internet Archive digitizes books.@Thorsted @gloriouscow @internetarchive this is not entirely my experience. for my
current work i'm looking at a bunch of (IA) MRC-encoded documents.pdfimagesdumps the foreground/background/masks which gives you very clean pages of text (as images) upon inverting the masks. it's three images per page but this is Finewhat i do find difficult to do though is recovering inline images, since MRC (or at least IA's use of it) does not attempt to extract those at all
-
We've been getting a lot of questions about how the Internet Archive digitizes books.@gloriouscow @Thorsted @internetarchive i saw them, yeah -- this is not fundamental to MRC, no? just the internet archive's particular use of it. were compression not at max it should be ok
-
We've been getting a lot of questions about how the Internet Archive digitizes books.@Thorsted @gloriouscow @internetarchive i am of a mixed opinion here. MRC is good for archival but definitely sucks for the end user. it would be very nice if user-facing PDFs could be just the mask + get jbig2ified + get an OCR layer... but perhaps that is a task better suited for shadow libraries