← All articles

FileHero guide

How to extract plain text from a PDF

A step-by-step guide to extract plain text from a PDF, with advice on the settings and checks that matter.

Before you start

Check that the document contains selectable text. Text extraction discards typography and page layout.

How to do it, step by step

  1. Open the PDF and check that its words can be selected, rather than only the whole page image.
  2. Extract the text and download the TXT file.
  3. Open the file in a text editor. Read across column breaks and check repeated headers or broken words.

Check the saved file

Inspect reading order in columns and look for broken words or repeated headers.

Why PDF text extraction produces empty output

An image-only scan can look like a text document while containing no actual text layer. Plain-text extraction cannot read the pixels as words. Use an OCR workflow to create a text layer first. If the PDF already has text, compare its selection order with the extracted TXT.