Choose a PDF
Select a PDF up to 40 MB. The browser checks its page count and whether it already contains selectable text.
Extract editable text from scanned or image-based PDF documents without sending the file to a server. Umi OCR processes up to 40 pages locally in your browser.
Small prioritizes accuracy. Tiny is faster with a smaller download; use Small for Japanese or to recheck important text. Models are saved for future visits when browser storage is available.
Files are processed locally in your browser. Images, PDFs, and extracted text are not uploaded to Umiocr.
A PDF can contain real text, page images, or a combination of both. When text is selectable, it can usually be extracted directly. When a scanner or camera created the document, each page behaves like a picture and needs optical character recognition. This OCR PDF to text tool inspects the document and chooses the appropriate local extraction path.
For scanned pages, the browser renders each page and runs OCR against the rendered image. For text-based PDFs, it reads the existing text layer instead. The result is assembled with page separators, making it easier to review long documents and trace a passage back to its original page.
The PDF never needs to leave your device. That matters for contracts, financial records, academic scans, medical documents, and internal reports. Processing speed depends on the number of pages, image resolution, document complexity, and the available memory and CPU on your device.
Go from a local file to editable text in three steps.
Select a PDF up to 40 MB. The browser checks its page count and whether it already contains selectable text.
Text PDFs are read directly; scanned pages are rendered and processed with OCR locally on your device.
Edit the combined output, copy it, or download a TXT file with clear page separators.
Image-only pages are rendered for recognition so printed text can be recovered from old scans, photocopies, and camera-created PDFs.
When a usable text layer exists, the tool avoids unnecessary OCR and reads the original characters directly for a faster result.
The selected PDF and extracted content stay in your browser. There is no document upload, account, watermark, or daily quota.
Convert individual photos, scans, PNG, JPG, WebP, and BMP images to text.
Open page →Compare PDF OCR methods and learn how to prepare difficult scans.
Open page →See how WebAssembly can recognize documents locally without an upload.
Open page →