Skip to content
  • Hjem
  • Seneste
  • Etiketter
  • Populære
  • Verden
  • Bruger
  • Grupper
Temaer
  • Light
  • Brite
  • Cerulean
  • Cosmo
  • Flatly
  • Journal
  • Litera
  • Lumen
  • Lux
  • Materia
  • Minty
  • Morph
  • Pulse
  • Sandstone
  • Simplex
  • Sketchy
  • Spacelab
  • United
  • Yeti
  • Zephyr
  • Dark
  • Cyborg
  • Darkly
  • Quartz
  • Slate
  • Solar
  • Superhero
  • Vapor

  • Default (No Skin)
  • No Skin
Kollaps
FARVEL BIG TECH
jj@types.plJ

jj@types.pl

@jj@types.pl
About
Indlæg
4
Emner
0
Fremhævelser
0
Grupper
0
Følgere
0
Følger
0

Vis Original

Indlæg

Seneste Bedste Controversial

  • We've been getting a lot of questions about how the Internet Archive digitizes books.
    jj@types.plJ jj@types.pl

    @gloriouscow @Thorsted @internetarchive fair point!

    Ikke-kategoriseret

  • We've been getting a lot of questions about how the Internet Archive digitizes books.
    jj@types.plJ jj@types.pl

    @Thorsted @gloriouscow @internetarchive this is not entirely my experience. for my
    current work i'm looking at a bunch of (IA) MRC-encoded documents. pdfimages dumps the foreground/background/masks which gives you very clean pages of text (as images) upon inverting the masks. it's three images per page but this is Fine

    what i do find difficult to do though is recovering inline images, since MRC (or at least IA's use of it) does not attempt to extract those at all

    Ikke-kategoriseret

  • We've been getting a lot of questions about how the Internet Archive digitizes books.
    jj@types.plJ jj@types.pl

    @gloriouscow @Thorsted @internetarchive i saw them, yeah -- this is not fundamental to MRC, no? just the internet archive's particular use of it. were compression not at max it should be ok

    Ikke-kategoriseret

  • We've been getting a lot of questions about how the Internet Archive digitizes books.
    jj@types.plJ jj@types.pl

    @Thorsted @gloriouscow @internetarchive i am of a mixed opinion here. MRC is good for archival but definitely sucks for the end user. it would be very nice if user-facing PDFs could be just the mask + get jbig2ified + get an OCR layer... but perhaps that is a task better suited for shadow libraries

    Ikke-kategoriseret
  • Log ind

  • Har du ikke en konto? Tilmeld

  • Login or register to search.
Powered by NodeBB Contributors
Graciously hosted by data.coop
  • First post
    Last post
0
  • Hjem
  • Seneste
  • Etiketter
  • Populære
  • Verden
  • Bruger
  • Grupper