How to Extract Text From a PDF, Including Scanned Documents

PDF text extraction works when a document contains actual text objects. A scanned page may only be an image, which requires OCR first. Learn how to identify each case, extract text from suitable files, and handle sensitive documents carefully.

8 min read

Getting words out of a PDF can mean two different things: copying text already stored in the document, or recognizing letters inside a scanned page image. The first is text extraction. The second requires optical character recognition, usually called OCR. Mixing up those cases leads to confusing results: a PDF may look perfectly readable on screen and still yield no selectable words.

This guide shows how to diagnose the document, extract selectable text, and decide when an OCR-capable workflow is needed. Kinsad's PDF tool can extract existing selectable text and download it as plain text; it does not perform OCR, make Word documents, or preserve the original page layout.

First determine whether the PDF contains selectable text

Open the document in a PDF viewer and try to drag over a sentence or use the viewer's find command for a word you can see. If individual characters or words can be selected, the PDF likely has a text layer that an extractor can read. If selection instead behaves like grabbing a whole page image, or search cannot find visible wording, the PDF may be a scan with no text layer.

This is a useful first test, not a guarantee that every page is the same. A document can mix digitally created pages with scanned inserts, or contain a page image plus a hidden OCR text layer. Check several pages, especially any that look different. Adobe describes the distinction in its explanation of recognizing text in scanned PDFs: a scan may contain image data rather than searchable text, and OCR can create a text layer.

Why a PDF page can look like text but not be text

A scanner or camera captures a visual page as pixels. Those pixels can depict printed letters clearly, but they do not inherently say which shapes are letters, words, or sentences. In contrast, many PDFs generated from office software include characters encoded as text objects along with instructions about where to draw them. A viewer can display both kinds of pages similarly, although the underlying data is different.

OCR analyzes an image and attempts to infer the words shown in it. Recognition can be affected by low resolution, skew, blur, unusual fonts, handwriting, stains, columns, or poor contrast. Even a successful OCR pass should be checked against the source, especially for names, amounts, dates, and instructions where a single misread character matters.

Extract existing text from a PDF

For a short passage, select the text in a trusted PDF viewer and copy it into the destination app. For longer documents, batch work, or a plain-text copy to keep, use a PDF text-extraction tool. The Kinsad PDF to Text tool reads selectable text in the browser and lets you preview extracted content before downloading a .txt file.

  1. Choose the right source. Open the PDF and verify that at least some of the relevant text can be selected. Keep the original unchanged.
  2. Open the extractor. Select the PDF in Kinsad's PDF to Text tool.
  3. Start extraction. The tool processes the document and assembles text page by page. Page labels are included where text is found.
  4. Review the preview. Look for missing passages, unusual spacing, and reading-order problems. Compare important details against the PDF.
  5. Download plain text if useful. Save the .txt output and use it as a working copy. Retain the source PDF for layout, images, and verification.

This tool is labeled “PDF to Word” in its route, but it does not create a Microsoft Word .docx file. The actual output is plain text. You can copy the text into a word processor and apply formatting yourself, but headings, columns, tables, font styling, and page layout should not be expected to transfer as they appear in the PDF.

A worked example: a selectable report with two columns

Suppose you have a 12-page report generated from a desktop publishing program. You can select a paragraph on page one, and search finds terms on later pages. The PDF likely contains text, so a text extractor may be useful. Extract it, then inspect the plain-text preview page by page. If a page has two columns, the resulting text may not follow the visual reading order: a line from the left column could be followed by one from the right, or text fragments might be separated in a way that is awkward to read.

For a quick quote, copying a paragraph directly from the PDF viewer may preserve the exact passage more reliably than extracting the whole document. For a full research draft or index, a .txt file can be a convenient starting point, but compare citations, numbers, and table data to the source before relying on them. Treat extraction as a transfer of text, not as a correctness audit.

If you need a formatted editable document, use a conversion workflow designed for that purpose and inspect the result. Plain-text extraction is appropriate when the text itself matters more than the visual styling; it is not a replacement for a faithful layout conversion.

What to do when the PDF is scanned

If you cannot select or search the words on a page, text extraction may return nothing for that page. Kinsad's tool reports when no selectable text is found and notes that scanned PDFs require OCR. It does not use OCR to turn the visible image into words. Choose an OCR-enabled application or service if you have permission to process the document there.

A typical OCR workflow takes the image-only PDF, recognizes the characters, and adds a text layer or produces recognized text. Afterward, test selection and search again. Then proofread the output against the original, checking page numbers and any critical text. If the scan is crooked or fuzzy, improving the scan quality first may make recognition more reliable; OCR cannot promise a perfect transcription from unreadable source material.

For a mixed document, identify the pages that lack selectable text rather than assuming one result applies to all. You may extract existing text from digital pages and OCR only the scanned portions, then combine and review the outputs. Keep page references so that every paragraph can be traced back to the visual original.

Formatting and reading order

A PDF describes a page, not necessarily a clean linear transcript. Columns, footnotes, and tables may be extracted in unexpected order or lose their relationships. Images, charts, signatures, and diagrams are not turned into descriptive words by ordinary text extraction. Check the original page whenever layout affects meaning, and keep a page reference with important passages.

Privacy and handling sensitive PDFs

Consider the document's contents and your organization's rules before using any online processor. Kinsad's current PDF-to-text implementation reads the selected file into browser memory and processes it with PDF.js in the browser; the source file is not uploaded to a Kinsad processing endpoint by this tool's extraction flow. This describes the current implementation, not a general guarantee about every website, browser extension, device, or future site change.

Local browser processing does not erase every privacy concern: the PDF exists on your device, extracted text appears on screen, and downloads may remain in shared or cloud-synced folders. For confidential or regulated records, follow your organization's approved process and protect the extracted file as carefully as the PDF. Do not send a document to an OCR service unless you are authorized and have reviewed its data practices.

Frequently Asked Questions

Q: Why can I see words but not extract them?

The visible page may be a scan or photo stored as an image rather than selectable text. OCR is needed to recognize those letters. Try selecting a sentence or searching for a visible word to check.

Q: Does Kinsad's PDF tool convert to a Word document or run OCR?

No. It extracts existing selectable text and downloads a plain .txt file. It does not create a .docx document or recognize text from scanned page images.

Q: Is the selected PDF uploaded for Kinsad text extraction?

The current tool implementation processes the file locally in the browser using PDF.js rather than sending it to a Kinsad extraction endpoint. Still follow your organization's rules for sensitive documents and protect both the original and extracted file.

The fastest way to choose a workflow is to test selectability first. Use text extraction for words already stored in the PDF, OCR for image-only pages, and careful review for either route. Keep the original document close at hand: it remains the reference for exact wording, structure, and context.