OCR

Make a scan searchable without sending it anywhere.

Half of what a client sends is a photograph or a scan with no text in it at all. Oxofolio puts a searchable text layer underneath the image they sent — leaving the picture itself untouched — so the page can be found, read and extracted from like any other.

Runs offline · opens no sockets
The problem

Why this is worth a tool at all.

A folder of scanned statements is invisible to search, to extraction and to redaction. Anything that needs to read the words has nothing to read.

Every convenient way to fix that involves uploading the documents. Free OCR services are free because the documents are the product.

The check you can run yourself

Turn off networking, then OCR a scan. It works, because the engine was already on your disk.

Every claim on this page is about a mechanism rather than an outcome, so each one is something you can verify rather than something you have to believe.

How it works

How it runs with the network unplugged.

The engine ships inside the installer Tesseract is bundled on all three platforms. Nothing is downloaded on first use, so this works on a machine that has never had a network connection.
The image stays as the page The original scan remains the visible document; the recognised text goes underneath it. Nothing about the page a reviewer looks at changes.
One extraction path, not two It emits a searchable PDF, so everything downstream — find, extract, redaction, privacy scanning — treats an OCR'd scan exactly like any other document.
Files that already have text are skipped Rather than processed twice, which would layer a second, worse text layer under a perfectly good one.
The boundaries

What it will not do.

  • English is the only language pack bundled. Roman-script documents in other languages are recognised acceptably; a different script is not recognised at all, and the failure looks like a poor scan rather than a missing pack. Additional packs are a build change, not a download.
  • It reports no confidence percentage. “87% confident” is not a working paper — where a value matters, the snippet it came from is shown instead.
  • It does not correct, straighten or clean up the image. What you scanned is what stays on the page.
Read the security pack

Why the limits are on the page

A tool that claims everything is a tool nobody can check. Each of these is a decision rather than a gap, and stating them here means the download contains exactly what this page described.

The same discipline runs through the security pack, where the unfavourable answers sit on the first page rather than the fourth.

Straight answers

What people ask about this.

Does OCR work on Windows?

Yes, on all three platforms. The Windows bundle is verified during the build by running the engine with PATH cut back to System32, so a missing dependency fails the build rather than the customer.

Which languages are supported?

English only, today. This is stated plainly because the failure mode for an unsupported script is silence rather than an error.

Does it change my original scan?

No. A new searchable PDF is written beside it; the file the client sent is untouched.

Can I then extract figures from the scan?

Yes — once there is a text layer, extraction, search and redaction all work on it the same way they work on a born-digital PDF.

Next

The tools people use alongside it.

PDF to Excel

Every “PDF to Excel” button is a confident guess. This one shows its working.

Fourteen days, everything unlocked

Try it on your own files.

Download Oxofolio →