Skip to content
PDFStack

Scanning and OCR

How to scan documents well

The PDFStack team · 15 March 2026 · 4 min read

No amount of processing rescues a bad scan. Text recognition, compression and straightening all work on what the scanner gave them, and if that was a dim, crooked photograph of a page, the results will show it. Ten seconds of care at capture saves the lot.

Resolution: 300 dpi, almost always

Below roughly 200 dpi, the fine detail that distinguishes similar characters starts to disappear, and recognition accuracy falls off a cliff. Above 400 dpi you are storing detail nobody will ever see, at several times the file size.

300 dpi is the setting to use for ordinary printed text. Go higher only for small print, faint carbon copies, or anything with fine detail that matters.

Colour: greyscale for text

Scanning black ink on white paper in full colour triples the file size and adds nothing. It also adds noise: the sensor records slight colour variation across what should be flat white, which compression then has to encode.

Greyscale for ordinary documents. Colour only when the colour carries information. A signature in blue ink to distinguish it from a photocopy, a highlighted passage, a chart.

Get the page straight

A page fed in at an angle costs more than it looks. Line detection in recognition software assumes text runs horizontally; a few degrees of tilt breaks that assumption and words start merging across lines. Straightening afterwards helps, but it also means rasterising the page again, with another round of quality loss.

Square the stack against the guides. If you are photographing rather than scanning, get directly above the page rather than leaning over it, which is what produces the trapezoid shape no software fully corrects.

Light, if you are using a phone

Phone cameras are good enough for documents now, and the limiting factor is almost always lighting rather than the sensor. Two things to avoid: your own shadow falling across the page, and a single bright lamp to one side, which produces a gradient the compressor then treats as detail worth preserving.

Diffuse light from in front, or daylight from a window, beats an overhead bulb.

Check the first page before doing the rest

The most expensive scanning mistake is discovering the settings were wrong after forty pages. Scan one, open it, zoom to 100%, and read a line of the smallest text on it. If you can read it comfortably, the settings are right.

What to do afterwards

In order: straighten if anything is crooked, run recognition to add a text layer, then compress. That sequence matters. Recognition works better on a straight page, and compressing before recognition means recognising a degraded image.

Feeding a document scanner

Square the stack against the guides every time, not just at the start. Paper drifts as it feeds, and a stack that started straight can end crooked.

Remove staples and paperclips properly rather than flattening them. A folded corner produces a shadow that the compressor then treats as detail worth keeping, and a staple hole near text confuses recognition.

If pages are different sizes, scan them in groups. Mixed sizes in one pass produce a document whose pages are all different, which prints badly and looks careless.

Duplex scanning and blank backs

Scanning double-sided produces a blank page for every single-sided sheet in the stack. On a fifty-page bundle that is a lot of empty pages.

Remove them afterwards rather than trying to sort the stack beforehand. Detection by ink coverage is reliable when the threshold is set low, and it lists what it removed so you can check nothing with a single line on it disappeared.

What resolution actually costs

Doubling the resolution quadruples the pixel count and roughly quadruples the file size. Going from 300 to 600 dpi on a fifty-page document turns forty megabytes into a hundred and sixty, for detail that no screen will show and no reader will use.

The exception is small print. A footnote at six points has strokes thin enough that 300 dpi starts to lose them. If the document is dense legal text, going up is justified; for ordinary correspondence it is waste.

Checking a batch before filing it

Scroll the whole thing once at a readable zoom. You are looking for three things: pages that fed through twice, pages that are missing, and pages that came out upside down because the stack was turned.

All three are quick to fix while the paper is still in front of you and irritating to fix six months later.

The processing order that works

Straighten, remove blanks, run recognition, then compress. Each step benefits from the ones before it: straight pages read better, blank pages waste recognition time, and compressing last means recognition worked on the clean image rather than a degraded one.

Tools mentioned here