1Upload the scanned PDF documents; a clean 300 DPI grayscale scan gives the recogniser the most to work with.
2Tell it which language to expect — naming the language beats letting it guess by a wide margin.
3Run the recognition pass; the words are written back as an invisible layer sitting over the original page image.
4Download the searchable PDF file — it looks identical, but you can now search and select the text in it.
OCR PDF FAQ
Does the source format affect recognition?
+
It does. PDF fixes the page: fonts, vectors, raster images and text coordinates are frozen so every reader sees identical layout. How the page is stored decides what resolution and colour information the recogniser has to work with.
Anything specific to PDF here?
+
Yes — a PDF is an object graph, not a page image, so text stays selectable and vectors stay sharp no matter what happens to the raster content inside it. It affects what the recogniser can see.
How does OCR PDF turn a scan into text?
+
Concretely, each page is rendered, run through a Tesseract text-recognition pass in over 100 languages, and the recognised words are written back as an invisible text layer positioned over the original image — so the page looks unchanged but is searchable and selectable. The page still looks exactly as it did — the recognised text sits invisibly behind the image so search and selection work without changing the appearance.
Which languages can it recognise?
+
Over 100, including Latin, Cyrillic, Greek, Arabic, Hebrew, Chinese, Japanese, Korean and the Indic scripts. Telling it which language to expect materially improves accuracy over letting it guess.
Do bookmarks, links and form fields survive?
+
Outlines, internal links and annotations are preserved wherever the operation allows it. Digital signatures are the exception: any change to the file necessarily invalidates a signature, because that is precisely what a signature is for.
What can I upload to OCR PDF?
+
Any standard PDF, including ones produced by a scanner, by a word processor or by a print driver. Files with an owner password that only restricts editing are handled; files that need a user password to open are not, and we do not attempt to break encryption.
Is there a file size limit on OCR PDF?
+
Yes: free accounts process documents up to 25 MB each, which covers most reports and contracts but not a long scanned document at 600 DPI; Ghostscript and qpdf do the work. Scanned PDFs are the usual thing that exceeds it, and compressing the document first is normally enough to bring it back under.
Will OCR PDF lower the quality of my PDF documents?
+
Text and vector artwork are objects rather than pixels, so they stay perfectly sharp at any zoom no matter what happens. Only the embedded raster images can degrade, and only if the operation you chose resamples them.
Why would a modern-format site run OCR PDF?
+
WEBP.to is built around the modern replacement for both of the old defaults: one format that does lossy, lossless, alpha and animation, and reliably lands smaller than the JPEG or PNG it came from. People arrive here having decided the old defaults are costing them bandwidth, and that decision immediately raises the rest: what size, what quality, what happens to the transparency. OCR PDF is on the same upload and the same account so those questions get answered in one place.
Once OCR PDF has run, what is worth doing next?
+
The converter on this site moves images to and from WebP, PNG, JPG and AVIF, which is where you decide whether a smaller file is worth losing the older browsers. Doing it in that order matters: pick the pixels first and the container afterwards, because the container is the cheap decision and the pixels are the expensive one.
Do the sibling sites run something different?
+
The engines are identical — same libraries, same workers, same caps. What is different is the opinion attached, and the opinion here rests on a property of the format this site is named after: the same .webp extension can hold a lossy VP8 frame, a lossless one, or an animation, so the sensible processing path is decided per file rather than per extension.
What happens to my file, and is an account required?
+
No account, and nothing is retained: uploads are deleted from the workers shortly after the job finishes, and nothing is inspected, listed or indexed. Accounts exist for history and batch size only.