OCRTurk: A Comprehensive OCR Benchmark for Turkish


Creative Commons License

Yılmaz D., Munis E. A., TORAMAN Ç., Köse S. K., Aktaş B., Baytekin M. C., ...Daha Fazla

2nd Workshop on Natural Language Processing for Turkic Languages, SIGTURK 2026, Rabat, Fas, 29 Mart 2026, ss.197-208, (Tam Metin Bildiri)

Özet

Document parsing is now widely used in applications, such as large-scale document digitization, retrieval-augmented generation, and domain-specific pipelines in healthcare and education. Benchmarking these models is crucial for assessing their reliability and practical robustness. Existing benchmarks mostly target high-resource languages and provide limited coverage for low-resource settings, such as Turkish. Moreover, existing studies on Turkish document parsing lack a standardized benchmark that reflects real-world scenarios and document diversity. To address this gap, we introduce OCRTurk, a Turkish document parsing benchmark covering multiple layout elements and document categories at three difficulty levels. OCRTurk consists of 180 Turkish documents drawn from academic articles, theses, slide decks, and non-academic articles. We evaluate seven OCR models on OCRTurk using element-wise metrics. Across difficulty levels, PaddleOCR achieves the strongest overall results, leading most element-wise metrics except figures and attaining the best Normalized Edit Distance scores in easy, medium, and hard subsets. We also observe performance variation by document type: models perform well on non-academic documents, while slideshows become the most challenging.