Extracting Images using pdfimages: Getting 3 Images Per Page: .jp2, .png, .jb2e

Extracting Images using pdfimages: Getting 3 Images Per Page: .jp2, .png, .jb2e 图片 1

Using pdfimages -all on a .pdf file, each page of which is text, I'm getting 3 images for each page in the pdf:

Foo-001-000.jp2

Foo-001-002.png

Foo-001-002.jb2e

The first file is mostly blank, but contains some ghostly background plus an occasional piece of text. The second file is black and white and appears to be some kind of mask, perhaps identifying where the text in the third file is located (?) The third file I am not able to view in Ubuntu's image viewer or gimp.

If I use -png I similarly get three images, but all are .png's. Most (almost all) of the pdf's text in in the third image.

pdfimages -list looks like this:

Could someone help me understand what I've got here, and how I might combine these three images to get a single images for each page. Or equivalently, just extract single images per page. They key issue for me is to keep as much information as is available in these images. I want to avoid degradation in quality.

i.sstatic.net/gYW4s.pngi.sstatic.net/ujoBu.pngi.sstatic.net/yRbwo.png

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论