PDF to Text, a challenging problem

The search engine has recently gained the ability to index the PDF file format. The change will deploy over a few months.

Extracting text information from PDFs is a significantly bigger challenge than it might seem. The crux of the problem is that the file format isn’t a text format at all, but a graphical format.

It doesn’t have text in the way you might think of it, but more of a mapping of glyphs to coordinates on “paper”. These glyphs may be rotated, overlap, and appear out of order, with very little semantic information attached to them.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论