A Dataset and Benchmark for Embedded Systems Information Retrieval and Generation Bilgi Erişimi Destekli Üretim ?Için Gömülü Sistemler Veri Kümesi ve De?gerlendirme Protokolü
34th Signal Processing and Communications Applications Conference, SIU 2026, İstanbul, Türkiye, 7 - 10 Temmuz 2026, (Tam Metin Bildiri)
- Yayın Türü: Bildiri / Tam Metin Bildiri
- Doi Numarası: 10.1109/siu71813.2026.11636984
- Basıldığı Şehir: İstanbul
- Basıldığı Ülke: Türkiye
- Anahtar Kelimeler: benchmark, dataset, embedding models, embedding systems, evaluation, information retrieval, large language models, text generation, tokens
- Orta Doğu Teknik Üniversitesi Adresli: Evet
Özet
This study presents a dataset and a benchmark framework developed to evaluate retrieval-augmented generation pipelines in the field of embedded systems. Since large language models are pretrained on general-purpose data, they may struggle to produce contextually appropriate and reliable responses in niche and highly specialized domains. The proposed structure aims to systematically investigate this problem by combining article-chunk-question information derived from open-source documents, books, and Wikipedia content. The benchmark framework integrates retrieval quality, retrieval-augmented generation quality, and system efficiency within a single experimental setting. In this respect, the study provides a practical foundation for reproducible and comparable evaluation in the field of embedded systems by jointly presenting a domain-specific dataset and a benchmark protocol.