Urdu OCR — Extract Urdu Text From Any Image
Upload an image containing Urdu and get editable Urdu text back in the original script. The model is told the source is Urdu, which matters because Urdu shares most letters with Arabic but adds its own and is typically set in a very different calligraphic style.
Drop your image here
or click to browse — JPG, PNG, WEBP, GIF…
Upload an image and click Extract Text
All readable text from the image will appear here.
What Makes Urdu OCR Difficult
Urdu is the hardest of the Perso-Arabic scripts to OCR, because it is traditionally typeset in Nastaliq rather than the flat Naskh style used for Arabic. Nastaliq slopes each word diagonally downward from right to left, so words overlap vertically and a single line of text has no consistent baseline — the segmentation assumptions most OCR engines make simply don't hold. Urdu also adds letters Arabic lacks, including the retroflex ٹ ڈ ڑ marked with small ط above, plus ہ and ے, and these distinguishing marks are tiny. Newspapers compound the problem with narrow columns and tight line spacing that let descenders from one line collide with the next.
Getting Accurate Urdu Results
- Naskh-set Urdu (common in textbooks and digital text) extracts far better than Nastaliq. If you have a choice of source, pick Naskh.
- Give the capture generous line spacing — in Nastaliq, overlapping lines are the main cause of garbled output.
- Check the retroflex letters ٹ ڈ ڑ against their plain counterparts ت د ر; the distinguishing mark is small and often lost.
- For newspaper scans, crop a single column at a time rather than the full page.
Who Uses Urdu OCR
Literature and poetry
Digitise Urdu poetry, prose and scanned books so the text becomes searchable and quotable.
News archives
Extract text from Urdu newspaper clippings and magazine scans for research or reference.
Documents and correspondence
Pull text off Urdu letters, certificates and official paperwork without retyping right-to-left script.
Urdu OCR — Frequently Asked Questions
Does it work on Nastaliq script?▼
It attempts Nastaliq and often produces usable text, but accuracy is noticeably lower than for Naskh-style Urdu. Nastaliq's sloping, overlapping word shapes are the single biggest obstacle in Urdu OCR. Expect to proofread.
Is Urdu different from Arabic OCR?▼
Yes, which is why there is a separate page. Urdu uses additional letters such as ٹ, ڈ, ڑ, ہ and ے that don't exist in Arabic, and its typical typeface style is entirely different. Telling the model it is reading Urdu rather than Arabic improves results.
Does the output keep the Urdu script?▼
Yes, the text comes back in Urdu's Perso-Arabic script, not romanised. You can paste it straight into a document or translation tool.
Can it read handwritten Urdu?▼
Rarely well. Handwritten Urdu inherits all of Nastaliq's difficulties and adds personal variation on top, so results are unreliable. Printed or digitally typeset Urdu is strongly recommended.
OCR in Other Languages
Or use the general image to text converter if your image contains several languages at once.