OCRTurk: A Comprehensive OCR Benchmark for Turkish
2nd Workshop on Natural Language Processing for Turkic Languages, SIGTURK 2026, Rabat, Morocco, 29 March 2026, pp.197-208, (Full Text)
- Publication Type: Conference Paper / Full Text
- Doi Number: 10.18653/v1/2026.sigturk-1.16
- City: Rabat
- Country: Morocco
- Page Numbers: pp.197-208
- Open Archive Collection: AVESIS Open Access Collection
- Middle East Technical University Affiliated: Yes
Abstract
Document parsing is now widely used in applications, such as large-scale document digitization, retrieval-augmented generation, and domain-specific pipelines in healthcare and education. Benchmarking these models is crucial for assessing their reliability and practical robustness. Existing benchmarks mostly target high-resource languages and provide limited coverage for low-resource settings, such as Turkish. Moreover, existing studies on Turkish document parsing lack a standardized benchmark that reflects real-world scenarios and document diversity. To address this gap, we introduce OCRTurk, a Turkish document parsing benchmark covering multiple layout elements and document categories at three difficulty levels. OCRTurk consists of 180 Turkish documents drawn from academic articles, theses, slide decks, and non-academic articles. We evaluate seven OCR models on OCRTurk using element-wise metrics. Across difficulty levels, PaddleOCR achieves the strongest overall results, leading most element-wise metrics except figures and attaining the best Normalized Edit Distance scores in easy, medium, and hard subsets. We also observe performance variation by document type: models perform well on non-academic documents, while slideshows become the most challenging.