Skip to content
How-To

How to Redact a Scanned PDF

A scanned PDF looks like a document but behaves like a photograph. That changes how you find what to redact, what can go wrong, and how you check the result.

Why a scanned PDF is harder to redact

When a scanner or a phone app turns paper into a PDF, each page is stored as an image. The account number you can read on screen is a pattern of pixels, not a string of characters. Press Ctrl+F (or Cmd+F on a Mac) and search for it: nothing comes back.

Every “find and redact” feature depends on that search. If the file contains no text, the tool has nothing to match, and a search that finds nothing looks exactly like a document with nothing to redact. In desktop PDF software, recognising the text of a scan is often a separate command, to be run before searching — and easy to forget.

The fix is optical character recognition (OCR): software that reads the page image and works out which characters are where. Once the positions of the words are known, patterns can be detected and boxes placed over them.

The opposite trap: scans that already contain hidden text

Many office scanners run their own OCR and embed the result as an invisible text layer behind the image, so the file can be searched. It looks like a plain scan, but it is not.

That layer is where scanned documents leak. Paste a black rectangle over the image in a PDF editor or with an annotation tool, and the picture is covered — but the invisible text underneath is still there, and anyone can select it and copy it out.

Quick test: open your scan and try to select a word with the mouse. If a word highlights, the file has a text layer, and redaction has to remove that text too — not only cover the image.

How RedactOffline handles scanned pages

The editor decides page by page, so a PDF that mixes typed pages and scanned annexes is handled in one pass.

Page with a text layer

The existing text is used as-is and scanned for sensitive patterns. No OCR is run. On a scanner-made layer, detection is only as good as the scanner's own recognition.

Page that is only an image

A quick check looks for readable text first. If there is some, full OCR runs in a background worker in your browser, and detection runs on the recognised words.

On export, every page is redrawn from scratch: the page is rendered to an image, the redacted areas are painted over in the pixels, and a new PDF is assembled from those images. The original file's structure, including any old text layer, is not copied. On pages that had a text layer, the words lying outside the boxes are added back as invisible text so they remain selectable, and the words inside a box are dropped. On image-only pages, the output is image-only.

One caveat for scanner-made layers: a word is dropped when its position in the text layer falls inside a box. Scanner OCR is occasionally misaligned with the image, so a word can sit visually under a box while its hidden copy is recorded a little to the side. The search step in the checklist below catches exactly that.

None of this involves a server. The OCR engine, the detection patterns and the export all run in the browser tab. See the security architecture for how that is enforced.

Step by step: redacting a scan

1

Open the scanned PDF in the editor

Drop the file on the editor at redactoffline.com. It is read from your disk into the browser tab; nothing is uploaded. JPEG, PNG and WebP images of documents work the same way.

2

Let each page be analysed

Pages that already carry a text layer are read directly. Pages that are only an image go through text recognition (OCR) in a background worker in your browser. A progress bar shows “Recognizing text (OCR)…” while it runs.

3

Review what was detected

Email addresses, phone numbers, IBANs, card numbers, national ID and passport numbers, dates, postal addresses and similar patterns are boxed on the page. Untick anything that should stay visible.

4

Add your own terms

Type a word or phrase into the custom terms list — a person's name, a case number, a company — and every occurrence the OCR recognised on every analysed page is boxed.

5

Draw boxes over everything OCR cannot read

Signatures, handwriting, stamps, photos, and any word the scan made illegible. A drawn box removes that area just as permanently as a detected one.

6

Export and verify

Export the PDF, then open the downloaded copy and check it page by page with the checklist below before sending it anywhere.

On the free plan, automatic analysis covers the first 5 pages of a document. Premium removes the limit.

What OCR gets wrong — and what to do about it

OCR is a best guess, and every miss is something left in the clear. Plan for it rather than trusting the detection count.

Poor scans produce misread characters

Blur, skew, low contrast, coffee stains and faint photocopies turn a 0 into an O or a 1 into an l. A number with one misread digit may no longer match its pattern, so it is not boxed. Rescan at a higher resolution if you can; otherwise read each page yourself.

Handwriting is not reliably recognised

Signatures, margin notes and filled-in paper forms are usually beyond OCR. Draw boxes over them by hand.

Only English and French models are loaded

Numbers and Latin-alphabet words in other languages are often recognised, but accented characters can be misread. Arabic, Cyrillic, Chinese and other non-Latin scripts are not recognised at all.

Names are never detected automatically

There is no person-name detector. Put known names in the custom terms list, then look for the ones you did not know about — a colleague cc'd at the bottom, a name in a letterhead.

Images inside the scan

ID photos, logos, QR codes and barcodes can identify a person or an account. OCR does not flag them. Box them manually.

Checklist: verify the redacted scan before sharing

  1. Open the exported file, not the one in the editor, and scroll through every page at full size.
  2. Search the exported file for a few of the values you redacted. None of them may be found. On a page that was a plain image there is no text to find at all.
  3. Try to select text inside each black box. Nothing should highlight.
  4. Zoom in on the edges of the boxes: no ascenders, descenders or partial digits should show.
  5. Check the pages you did not expect to contain anything — covers, annexes, stamped back pages.
  6. Open the document properties and check the title and author fields. Clear them in the export dialog if needed (see removing PDF metadata).

Frequently asked questions

Can you redact a scanned PDF without OCR?

Yes, by drawing boxes by hand over every sensitive area — OCR is not needed to black out pixels. What OCR adds is the ability to find things: without it, the PDF contains no text, so no tool can search it or detect patterns in it.

Why does searching my scanned PDF find nothing?

A scan stores each page as an image. The words you see are pixels, not characters, so there is nothing for Ctrl+F or a search-and-redact tool to match until text recognition has been run.

Does RedactOffline detect names on scanned documents?

Not automatically — the detection engine has no person-name category. Add the names you know to the custom terms list so every recognised occurrence is boxed, and draw boxes over any others. A drawn box removes the content just as permanently.

Which languages does the OCR recognise?

The recognition models loaded are English and French. Other languages written in the Latin alphabet are often partly recognised, especially numbers, but accented characters may be misread. Non-Latin scripts are not supported by the OCR; redact those by drawing boxes.

Is the redacted scan searchable after export?

No. Pages that were images on the way in are images on the way out, with the redacted areas burned into the pixels. If you need a searchable copy, run OCR on the redacted file afterwards — the black areas contain no text to recover.

Is my scanned document uploaded for OCR?

No. The OCR engine runs inside your browser, in a background worker. The page images it reads never leave your device.

Redact a scanned PDF in your browser

OCR runs on your device · Free plan for documents up to 5 pages · Free account, no credit card

Open the editor

Related guides