From robust mixed models to temporal deep learning: a five-paradigm robustness benchmark on contaminated fNIRS data


Çakar S., Yozgatligil C., Gökalp Yavuz F.

The International Conference on Robust Statistics (ICORS) , İstanbul, Türkiye, 20 - 24 Temmuz 2026, ss.1, (Özet Bildiri)

  • Yayın Türü: Bildiri / Özet Bildiri
  • Basıldığı Şehir: İstanbul
  • Basıldığı Ülke: Türkiye
  • Sayfa Sayıları: ss.1
  • Orta Doğu Teknik Üniversitesi Adresli: Evet

Özet

Functional Near-Infrared Spectroscopy (fNIRS) provides a non-invasive window into brain activity but is highly vulnerable to motion artifacts, sensor displacement, and physiological noise. These sources of contamination pose significant challenges for predictive modeling, particularly in real-world settings where signal quality is often compromised. Moreover, fNIRS data exhibit a complex hierarchical structure, with measurements collected across multiple channels nested within subjects and repeated over time. Subject-specific modeling and robust settings are therefore not merely a statistical refinement but a practical necessity for reliable inference in this domain. This study presents a comprehensive cross-paradigm robustness benchmark evaluating several models across five methodological families: pure statistical models for repeated measures (including Linear Mixed Model (LMM) and its robust counterpart), hybrid statistical-machine learning (ML) models, pure ML algorithms, deep learning architectures, and a temporal deep learning model. The benchmark targets prediction of standardized oxygenated hemoglobin change (ΔHbO) using an fNIRS dataset comprising 30 subjects and 20 channels across three experimental conditions [1]. LMM analysis established a significant hierarchical structure with nested random intercepts for Channel within Subject: the overall nested random-effects structure was highly significant. The nested channel-within-subject random intercept specifically contributed an additional significant improvement, while fixed effects explained negligible variance. This result motivates models capable of capturing individual-level temporal dynamics. Robustness was assessed via a Contaminate-Then-Split (CTS) protocol in which synthetic outliers of two magnitudes (±10 SD, ±15 SD) were injected at three contamination rates prior to temporal train-test splitting. Results reveal three principal findings. First, contamination magnitude governs predictive degradation more strongly than rate. Second, the Robust LMM demonstrates the most stable performance across all scenarios, validating M-estimation theory in practice [2]. Among pure ML models, Support Vector Regression proves most robust owing to the implicit outlier tolerance of its ε−insensitive loss function, whereas Extreme Gradient Boosting exhibits systematic overfitting in every scenario, achieving the lowest training error yet among the highest test errors across all conditions. Third, the proposed Dilated Temporal Network with Learned Random Effects (DilTNet-RE; see Figure 1), which combines dilated 1-D convolutions with learned Subject, Channel, Subject and Channel interaction embeddings, achieves the lowest test MAE on clean data, outperforming competing models by explicitly capturing slow hemodynamic autocorrelation across sequential trials that static models discard by treating observations independently [3]. DilTNet-RE maintains this ranking across all contaminated scenarios with a training-test error pattern consistent with good generalisation, though its margin over competing models remains modest. [] Figure 1: Architecture of the Dilated Temporal Network with Learned Random Effects (DilTNet-RE). Input sequences of 16 time steps with 6 features per step pass through positional encoding, a 6-block dilated convolutional stack (dilations 1–32), and global average pooling. The temporal summary is concatenated with learned Subject, Channel, and Subject×Channel embeddings before a two-layer prediction head produces the standardized HbO estimate.