Printed Ottoman Text Recognition Using Synthetic Data and Data Augmentation
preprint
OA: closed
Abstract
Ottoman script is an Arabic alphabet-basedscript as well. It was a writing system of theTurkish language for several centuries until it was replaced with the modern Turkish script,which is based on the Latin alphabet, in 1928. With the ever increasing digitization campaigns, millions of Ottoman documents are coming to light. But, their contents are not directly accessible, nor they are digitally editable and searchable. OCR and text recognition technologies can bea solution to this problem in the form of auto-mated and semi-automated conversion systems. This study presents a DL-based characterrecognition system for the printed Ottoman script. We first generate a synthetic text image dataset from a text corpus, and then augment it using some image processing methods. We develop a hybrid Convolutional Neural Network-Bidirectional Long Short Term Memory recognizer and train it with the original and the augmented datasets. Finally we apply a Transfer Learning procedure for adapting the system to real image data. The proposed system obtains 0.16 CER on a test set containing line images from a historical printed Ottoman book.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00