Adaptive Oversampling for Imbalanced Data Classification


Ertekin Ş.

28th International Symposium on Computer and Information Sciences (ISCIS), Paris, Fransa, 28 - 29 Ekim 2013, cilt.264, ss.261-269 identifier identifier

  • Cilt numarası: 264
  • Doi Numarası: 10.1007/978-3-319-01604-7_26
  • Basıldığı Şehir: Paris
  • Basıldığı Ülke: Fransa
  • Sayfa Sayıları: ss.261-269

Özet

Data imbalance is known to significantly hinder the generalization performance of supervised learning algorithms. A common strategy to overcome this challenge is synthetic oversampling, where synthetic minority class examples are generated to balance the distribution between the examples of the majority and minority classes. We present a novel adaptive oversampling algorithm, Virtual, that combines the benefits of oversampling and active learning. Unlike traditional resampling methods which require preprocessing of the data, Virtual generates synthetic examples for the minority class during the training process, therefore it removes the need for an extra preprocessing stage. In the context of learning with Support Vector Machines, we demonstrate that Virtual outperforms competitive oversampling techniques both in terms of generalization performance and computational complexity.