Improving Fairness in Doubly Imbalanced Datasets


Yalcin A., ÖZTÜRK A. U., SEVER Y., Pauw V., Hachinger S., TOROSLU İ. H., ...Daha Fazla

Scientific Reports, cilt.16, sa.1, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 16 Sayı: 1
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1038/s41598-026-54702-x
  • Dergi Adı: Scientific Reports
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, BIOSIS, Chemical Abstracts Core, EMBASE, MEDLINE, Directory of Open Access Journals, Zoological Record, Academic Search Ultimate (EBSCO), Natural Science Collection (ProQuest), Biological Science Database (ProQuest), Biomedical Reference Collection: Corporate Edition (EBSCO), Health Research Premium Collection (ProQuest)
  • Anahtar Kelimeler: Algorithmic fairness, Fair AI, Fraud detection, Imbalanced data, Multi-parameter optimization
  • Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
  • Orta Doğu Teknik Üniversitesi Adresli: Evet

Özet

Fairness has been identified as an important aspect of Machine Learning and Artificial Intelligence solutions for decision making. Recent literature offers a variety of approaches for debiasing, however many of them fall short when the data collection is imbalanced. In this paper, we focus on a particular case, fairness in doubly imbalanced datasets, such that the data collection is imbalanced both for the label and the groups in the sensitive attribute. Firstly, we present an exploratory analysis to illustrate limitations in debiasing on a doubly imbalanced dataset. Then, a multi-criteria based solution is proposed for finding the most suitable sampling and distribution for the label and the sensitive attribute, in terms of fairness and classification accuracy.