Optical Character Recognition of Urdu Text using Histogram of Oriented Gradient Features
preprint
OA: closed
CC-BY-4.0
Abstract
Optical character recognition has received significant research focus to digitize the text in images. Urdu OCR is a difficult task as compared to English and similar languages due to its complex nature where a character can have multiple inflections depending upon its position in the word. The proposed research work presents segmentation-free approach (i.e. holistic approach) for offline Urdu printed text detection. To extract text lines in an image, horizontal histogram projection is employed whereas for ligature segmentation in extracted image text line, proposed technique has used connected components labelling. In this model, set of 14 statistical features along with HOG features are extracted for each sub-word/ligature and used for the training of the proposed model. An open-source dataset UPTI [10] has been used to train and test the proposed algorithm. SVM with RBF kernel function is used for the classification of ligatures. The proposed algorithm has achieved 97.3% character recognition rate on given dataset.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00
- unpaywall
- last seen: 2026-05-27T02:00:06.600101+00:00
License: CC-BY-4.0