How a PDF stores a page

A PDF page is a list of drawing instructions: put these characters at this position in this font, draw this image here. It usually has no idea that some lines form a paragraph, that a grid of words is a table, or which column should be read first. Word documents and spreadsheets store exactly that structure.

Reliable: extracting text

When a PDF was created from a document (not scanned), its text can be read back accurately. Columns and tables may come out in an unexpected order, and line breaks follow the page rather than the paragraph — TextVix can join broken lines into paragraphs for you.

Reliable: pages as images

Drawing each page as a JPG or PNG is faithful by nature, because it uses the same instructions a PDF viewer uses. The text in the images is no longer selectable.

Approximate: PDF to Word, Excel or PowerPoint

To rebuild an editable document, a converter must guess paragraphs, tables, columns and styles from positioned text. Even dedicated desktop software gets this wrong on complex pages. TextVix’s PDF to Word is deliberately text only: you get the words as clean paragraphs, and nothing pretends to be the original layout.

Needs OCR: scanned PDFs

A scanned PDF contains photos of pages, not text. Getting words out of it requires optical character recognition (OCR). TextVix doesn’t do OCR, so it tells you when a PDF has no text rather than returning an empty file.

Try PDF to Text

Extract the selectable text from PDF files and save it as a TXT file.

Open PDF to Text

Frequently asked questions

How can I tell if a PDF is scanned?

Try to select a word in your PDF viewer. If you can’t, or the whole page highlights as one block, it is an image.

Why is the text order wrong in my extracted text?

Multi-column layouts, sidebars and tables are stored in drawing order, which may not be reading order. Check those parts by hand.