PDFAreUs کا لوگوPDFAreUs

اس صفحے کا ابھی اردو میں ترجمہ نہیں ہوا، اس لیے یہ انگریزی میں دکھایا جا رہا ہے۔

PDF Guides

How to OCR a Scanned PDF and Get the Text Out

A scanned PDF is a picture of a page. You cannot search it, select a sentence or copy a paragraph out of it. OCR (optical character recognition) reads the letters in that picture and gives you the words as real text. This guide shows how to do it with OCR PDF, what you get at the end, and how to get a clean result.

By GaryLast updated 5 min read

Open OCR PDF

Why scanned PDFs are hard to work with

Press Ctrl+F in a scanned PDF and nothing is found, because there is no text for the viewer to search. Each page is a single image. That makes scanned documents slow to search, impossible to copy from, and unreadable to screen readers.

A quick test tells you which kind of PDF you have: try to select a line with your cursor. If individual words highlight, the PDF already contains text and you do not need OCR. PDF to Text or PDF to Word will get the words out faster and without recognition errors.

What OCR PDF gives you

OCR PDF reads every page and gives you the recognised text as a file: plain text (.txt) or a Word document (.docx). You can edit it, search it, paste it into an email or keep it next to the scan.

It does not produce a searchable PDF. Some desktop programs place an invisible layer of text behind the page image so the PDF itself becomes searchable; this tool does not do that. If a searchable PDF is what you have been asked for, you will need desktop OCR software for that step. If you need the words, this is the quicker route.

How to OCR a scanned PDF

  1. Open OCR PDF and choose your scanned file.
  2. Start the recognition. The first time, your browser downloads the recognition engine, which takes a moment.
  3. Wait while each page is drawn and read. Long documents take a while, because the work is done on your own device.
  4. Read through the recognised text on screen.
  5. Download it as a .txt file, or as .docx if you want to carry on editing in Word.

Getting an accurate result

OCR is only as good as the image it reads. Most errors come from the scan, not from the recognition, so a few minutes of preparation pays off.

  • Pages must be upright. Turn sideways or upside-down pages with Rotate PDF first.
  • Straighten tilted pages with Deskew Scanned PDF. Slanted lines of text are a common cause of garbled output.
  • Use a sharp scan. Blurry phone photos, shadows across the page and creases all cost accuracy. Rescan if you can.
  • Do not run OCR on a heavily compressed copy. Use the original scan, and compress afterwards if you need to.
  • Dark text on a light, plain background reads best. Coloured or patterned backgrounds confuse the engine.

What it reads well, and what it does not

The tool recognises English text. Printed and typed documents in ordinary fonts, such as letters, contracts, reports and book pages, come out well from a clear scan.

Expect weaker results from handwriting, very small print, stylised fonts, stamps over text, and tables or multi-column layouts, where the words may be right but the reading order is not. Text in other languages and scripts will not be recognised correctly.

Always proofread the output

Even a good OCR pass makes small mistakes, and they tend to fall where they matter most. Check these before you rely on the text:

  • Numbers: amounts, dates, phone numbers and reference codes
  • Look-alike characters: the letter O and zero, lowercase l and the number 1, “rn” read as “m”
  • Names and email addresses
  • Line breaks in the middle of sentences, which you may want to remove

What people use the text for

Common uses are quoting from a scanned contract or letter without retyping it, bringing an old paper document back into Word to update it, pulling figures from receipts and invoices into a spreadsheet, and making notes from scanned book chapters or handouts. For a long document you have to search often, keep the .txt file beside the scan with the same name: search the text file to find the passage, then open the scan to see it in context.

Why it is slower than other tools

Recognition runs in your browser, on your device, using the open-source Tesseract engine. Nothing is sent to a server, which is the right way to handle contracts, medical letters and ID documents, but it means your computer or phone does the heavy work. A few pages take seconds; a long report can take several minutes. Keep the tab open until it finishes.

Frequently asked questions

Does this make my PDF searchable?

No. OCR PDF gives you the recognised text as a .txt or .docx file. It does not add a text layer to the PDF itself.

Can I copy and paste the text after OCR?

Yes. The result is ordinary text that you can copy, edit and search in any text editor or in Word.

Does OCR work on handwriting?

Not reliably. It is designed for printed or typed text and is much less accurate on handwriting.

Which languages does it recognise?

English. Text in other languages or scripts will not be recognised correctly.

Is my scanned document uploaded?

No. Recognition runs in your browser, so the file stays on your device.

Related tools and guides

About the author

Gary

Gary builds and runs PDFAreUs and writes its guides, which explain how to prepare PDFs for email, job applications, school and government portals using tools that run in your browser.

More about Gary and PDFAreUs