Guide · 7 min read
How to Get Text Out of a Scanned Image
This site ships no OCR tool — a deliberate choice. Learn to spot text PDFs vs scans, prepare clean images, and extract text you can count.
Last updated: 2026-10-03
Tools for this guide
Test whether your PDF already contains text
Before reaching for any converter, run the ten-second test: open the PDF, drag to select a sentence, and copy it into a text editor. If the words paste cleanly, your document already contains a real text layer and everything in this guide's extraction section will work. If selection grabs the whole page as a block, selects nothing, or pastes gibberish, you are looking at scanned images — pictures of text with no characters inside.
The PDF to Text tool settles the question definitively. Drop the file in and it reports words, characters, and pages extracted, shows the text grouped into lines, and offers a copy button plus a .txt download. An optional toggle re-extracts with page markers — separators reading Page N between pages — which keeps long documents navigable when you paste the result elsewhere. Extraction runs entirely in the browser, so the document is never uploaded.
When a file contains only scanned images, the tool says so plainly instead of inventing output: its error message explains the PDF may contain only scanned images. Treat that message as a diagnosis, not a failure — it tells you the document needs OCR, which is a different technology covered below, and saves you from trusting hallucinated text.
Understand what a scan actually is
A scanner or phone camera records a grid of pixels; it has no idea those pixels form letters. A text-based PDF, by contrast, stores each character with its position, font, and encoding — which is why selection, search, and extraction work there. Converting a scan between formats never bridges this gap: a PNG of a page and a PDF of the same page contain the identical unknowing pixels, just wrapped differently.
Many real documents are hybrids: a digitally generated contract with one scanned signature page, or a report where someone pasted photos of tables between text paragraphs. Extraction handles these gracefully, returning the real text pages and little or nothing for the image pages. Page markers earn their keep here, because the gaps in numbering reveal exactly which pages came back empty.
This distinction also explains mixed results people misread as bugs: three pages extract perfectly and the fourth returns blank. Nothing malfunctioned — the fourth page is a picture. Once you see documents as a mix of text pages and picture pages, every extraction result becomes predictable.
Why instant in-browser OCR is a red flag
This site offers no OCR tool, and that absence is deliberate. Real text recognition requires trained recognition engines with language data — downloads of around 17 MB or more before a single letter is recognized. A page claiming instant OCR with no engine download, no language pack, and no waiting time has skipped the step that does the actual work, and the only ways to skip it are uploading your file to somebody's server or returning low-grade guesses.
That matters twice over for scans, because scanned documents skew sensitive: contracts, IDs, medical forms, and financial statements are exactly what people scan. Uploading those to an unknown server to gain text you could get from an honest tool is a poor trade — the privacy cost is certain while the recognition quality is doubtful. Local-first processing exists precisely so private documents never need to leave the device.
So treat OCR claims the way you treat the extraction error message here: as information. A legitimate OCR product tells you which languages it supports, whether processing happens on your machine or its servers, and what output you get — searchable PDF, plain text, or both. Anything vaguer than that deserves skepticism, however polished the landing page looks.
Prepare scans a real OCR engine can read
Recognition quality is decided before the OCR engine ever runs: clean input produces clean text, and no engine rescues a crooked, shadowed phone photo of a creased page. Scan at a sensible resolution — around 300 dpi is the widely used working point for documents — and keep pages aligned so text lines run horizontally. Grayscale or color both work; what matters is contrast between letters and background, so disable any artistic filter and avoid heavy shadow across the page.
Crop away desk edges, fingers, and neighboring pages, and straighten skew before recognizing rather than after. Multi-page documents need consistent treatment: the same resolution, orientation, and cropping on every page, in reading order, so the recognized text flows without surprises. If a page is physically damaged — coffee stains, torn corners, faded thermal receipts — rescan it now rather than fighting the output later.
Photographing pages with a phone is acceptable when no scanner exists, but control the conditions: flat page, even overhead light, camera parallel to the paper, no flash glare. Two careful minutes per page at capture time save twenty minutes of correcting misrecognized characters afterward, particularly with names, numbers, and legal terms where a single wrong character changes meaning.
What to demand from a real OCR tool
When you do need recognition, choose the engine with open eyes. First, processing location: a desktop program or app that works offline keeps sensitive scans on your machine, while a website must receive your file — acceptable for a restaurant menu, questionable for a passport. Second, language support: the engine must explicitly list your document's language, since recognition models are language-specific and the wrong model produces confident nonsense.
Third, output format: prefer tools that return a searchable PDF (original image with an invisible text layer) or clean plain text, and avoid anything that returns only an image again. Fourth, verification aids: side-by-side views, confidence highlighting, and preserved layout all signal a serious product. Run one representative page through any candidate before committing a hundred-page archive, and proofread names, dates, and figures character by character — engines fail exactly where accuracy matters most.
Price and convenience cut both ways: free upload sites monetize somehow, usually through data retention or aggressive limits, while established offline programs cost money but keep documents home. Match the tool to the sensitivity of the paper, not just to the urgency of the moment.
Turn extracted text into something countable
Once extraction yields real text — from a text-based PDF, or from OCR output you produced elsewhere — the Word Counter turns it into usable numbers. Paste the text and it reports words, characters, characters without spaces, sentences, paragraphs, and lines, updating as you type or edit. Those six figures answer the practical questions: does this excerpt fit the application field, how long is the quoted passage, did the extraction drop whole paragraphs.
A solid workflow runs extraction, then counting, then cleanup in a loop. Extract with page markers so you can trace each passage to its source page; paste into the counter to check completeness against expectations — a ten-page contract returning forty words has clearly lost pages; trim headers, footers, and hyphenation artifacts from the scan era, watching the counts respond. The counter's live updates make this iterative tidying fast, because every deletion immediately shows its effect.
Keep the chain of custody short and labeled: source PDF name, extracted .txt download, counted draft. When a figure in the final document looks wrong weeks later, that trail lets you reopen the exact extraction and check the exact page instead of wondering which version anyone used.
Run a clean extraction pass on text PDFs
Extraction succeeds when the document cooperates, and cooperation is verifiable in under a minute. Drop one PDF onto the PDF to Text dropzone — files up to 100 MB are accepted — and processing starts immediately with a Reading document stage indicator, no convert button to hunt. The tool validates encryption first through the same loader the merge pipeline uses, so a locked file surfaces a clear password message instead of a cryptic stall. When the run completes, the header reports three figures together: words, characters, and pages extracted, each formatted with thousands separators so a lease reading 12,408 words across 31 pages is legible at a glance. Below sits the grouped text itself, reassembled line by line from glyph baselines and sorted left to right, which preserves multi-column reading order far more faithfully than naive copy-paste from a viewer.
Operate the output controls deliberately rather than skimming past them. The Copy text button places the entire extraction on the clipboard for pasting into editors, citation managers, or the Word Counter; the Download .txt button saves the same content under the source filename with a .txt extension, creating an auditable artifact you can archive beside the PDF. The Include page markers checkbox re-runs extraction with separators reading --- Page N --- between pages, so a quotation copied months later still points at page 14 instead of floating anonymously. Toggle markers on when the document exceeds a few pages, when collaborators will verify passages, or when hybrid content is suspected — the separators cost nothing and convert a wall of prose into a navigable transcript where empty stretches flag picture pages.
Treat the counts as a diagnostic instrument, not decoration. A ten-page services agreement returning forty words has plainly lost pages to scanned exhibits; a dissertation chapter whose character total collapses against expectations likely hides photographed tables between genuine paragraphs. Cross-check by selecting a sentence in the original viewer and comparing it against the extracted lines: matching phrasing confirms the pipeline, while garbled ordering or dropped diacritics mark sections to handle manually. Everything runs entirely in the tab with an explicit never-uploaded notice beneath the panel, so contracts, transcripts, and medical correspondence can be extracted on a shared workstation without routing sensitive prose through distant infrastructure. When extraction checks out, paste into the counter and begin cleanup; when it returns the scan-only message, pivot to the routing section below instead of hammering the same file.
Route scan-only files toward retyping or outside recognition
The scan-only verdict deserves respect rather than workarounds. When the extractor reports that text could not be pulled because the file may contain only scanned images, that sentence is the diagnosis: pixels arrived where characters were expected, and no slider, format shuffle, or re-save conjures letters from a photograph. Converting the PDF to images, printing to a fresh PDF, or renaming extensions merely rewraps identical pixels in fresh containers while the letterforms remain unrecognized grids. Corrupted files earn a separate message suggesting a re-save from the originating application, and password-locked files demand owner unlocking first — three distinct failures with three distinct remedies, none of which involve pressing the same button harder. Accepting the verdict early rescues hours otherwise spent laundering pictures through converters that promise transcription and deliver rearranged silence.
For short scans, disciplined retyping frequently outclasses every alternative on accuracy per minute. A single-page certificate, a half-page receipt, a vehicle card with seventeen fields: open the image beside a text editor, type what you see, then paste the result into the Word Counter to verify word, character, sentence, paragraph, and line totals against the receiving form. Numbers, proper nouns, and dates get a second glance because those are precisely the tokens recognition engines mangle and humans skim past — read serials aloud digit by digit, confirm accented names against the image, and expand hyphenated line-breaks that scans preserve as artifacts. Label the typed file with its source, such as invoice-2024-118-typed.txt, so future readers distinguish witnessed keystrokes from machine output and know which page to recheck.
Longer archives that justify outside recognition still benefit from preparation done here. Straighten rotation with a rotation pass first, since skewed baselines punish every downstream engine; crop desk edges, thumbs, and neighboring sheets so the page fills the frame; capture phone photographs flat under even overhead illumination with the lens parallel to the paper and flash disabled to suppress glare pools. Scan near 300 dpi where the hardware allows, hold resolution and orientation constant across the stack in reading order, and rescan stained or torn sheets now rather than proofreading around them later. Evaluate candidate recognition products with the skepticism this guide advocates: demand stated languages, a declared processing location that keeps sensitive matter on local hardware when sensitivity warrants it, searchable-PDF or plain-text output rather than another picture, and confidence highlighting for verification. Pilot one representative page, proofread names and figures character by character, and only then commit the archive.
Salvage tables, footnotes, and margin notes after extraction
Tabular matter and peripheral notes fail first under extraction, so budget a dedicated salvage pass for them. Ledger sheets, tariff schedules, laboratory grids, and timetable appendices rely on column alignment that linear text output cannot preserve: the extractor returns cell contents row by row in reading order, which keeps every figure present while surrendering the visual grid. Rebuild small tables manually by copying the extracted columnar sequence into a spreadsheet and re-slicing it against the page image open alongside — header row first, then quantities, then subtotals — watching the Word Counter respond as stray hyphenation and duplicated headers are pruned. Dagger symbols, asterisks, superscript numerals, and parenthetical addenda that annotate rows deserve relocation into a notes column rather than abandonment, since a tariff footnote can reverse the meaning of the figure it qualifies.
Footnotes, endnotes, marginalia, and errata slips need similar triage because positional cues evaporate outside the page. Extracted output appends footnote blocks where the baseline sorter placed them, often pages away from their anchors, so reunite each note with its numeral by searching the marker text and pasting the pair into the working draft together. Handwritten marginalia on otherwise typed memoranda — approver initials, circled paragraph numbers, arrow-linked insertions — never extract at all when they are ink on a scan; transcribe those strokes verbatim into bracketed editorial insertions so later readers distinguish author prose from annotator commentary. Errata slips tipped into bound volumes get the same treatment: type the correction, cite the slip location with a page-marker reference, and retire the loose sheet from the circulation copy.
Close with a fidelity ritual that catches silent omissions before they propagate into quotations. Compare the counter totals for the cleaned draft against a rough expectation from the source — chapter length, exhibit count, schedule rows — and investigate shortfalls page by page with markers enabled, since a blank stretch between Page 9 and Page 11 pinpoints the photographed exhibit that contributed pixels instead of prose. Read proper nouns, statute numbers, chemical names, and monetary figures aloud against the image, because those tokens carry the document weight and tolerate zero drift. Archive the extraction text, the counted draft, and a one-line provenance note naming the source file and extraction date; when a disputed figure resurfaces during review, that trio lets any colleague reopen the exact page and re-verify the exact token without relitigating the entire pipeline.
Frequently asked questions
Why does PDF to Text return nothing for my scan?
A scan is a picture of text, not text itself — there are no characters for the extractor to read. You need OCR software to recognize the letters first; this site deliberately does not pretend to do that in the browser.
Can I just convert the scan to a different format to get the text?
No. Converting a PDF to images or re-saving it rearranges the same picture data without recognizing any letters. Only OCR — or retyping — turns pictured words into real characters.
How do I know extraction actually worked?
The tool reports words, characters, and pages extracted, and shows the text with optional page markers. Skim the output for garbled ordering or missing sections before pasting it anywhere important.
Will extraction work on a password-protected PDF?
Only after it is unlocked by its owner. Password-protected files must be re-saved without the password first — no honest browser tool bypasses document encryption.
Why do page markers matter when I paste text elsewhere?
The optional toggle re-extracts with separators reading --- Page N --- between pages, so quoted passages stay traceable to their source page. Gaps in numbering also reveal which pages held only pictures and returned nothing.
Try it now — free, no signup
Get the words, skip the wrapper. Files stay on your device.
Open PDF to Text
