Skeleton-Based Activity Recognition for Children with Autism Using Graph Convolutional Networks


Ay B., ÖZTÜRK M. A., Aydin G.

SENSORS, cilt.26, sa.14, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 26 Sayı: 14
  • Basım Tarihi: 2026
  • Doi Numarası: 10.3390/s26144638
  • Dergi Adı: SENSORS
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Compendex, EMBASE, INSPEC, MEDLINE, Directory of Open Access Journals, Academic Search Ultimate (EBSCO), Biomedical Reference Collection: Corporate Edition (EBSCO), Health Research Premium Collection (ProQuest)
  • Orta Doğu Teknik Üniversitesi Adresli: Evet

Özet

Movement-based and physical activity programs are central tools in autism intervention, so recognizing the activities a child performs during therapy is valuable for objective progress tracking. Manual monitoring of these sessions is time-consuming and subjective, and raw videos raise privacy concerns because it shows identifiable children. We address autism therapeutic activity recognition from privacy-preserving 2D skeletons, and we focus on the practical difficulty of how several therapeutic activities differ only in subtle motion details. As a backbone, we adopt ProtoGCN, a graph convolutional network that represents each action as a combination of learnable motion prototypes. However, this contrastive backbone organizes all classes at once, so it does not enforce a margin between the few pairs that remain entangled after training. We therefore introduce a Refine-Confusable (RC) module, a training-only regularizer that pushes apart the empirically most-confused class pairs using a hinge-margin loss over momentum-updated class centroids. The module changes neither the backbone nor the inference cost. On the MMASD dataset, restricted to the ten-class 2D-skeleton configuration, the RC module improves the base model across random, session-independent, and subject-independent evaluation. The gain is largest on the strictest subject-independent split and a clip-level analysis confirms that this improvement is statistically significant. Under the protocol-matched holdout, the method reaches 96.30% accuracy with 0.959 macro-F1, surpassing recent 2D-skeleton baselines while keeping a lightweight and privacy-preserving modality. The improvements are modest, as expected on a small clinical dataset, and t-SNE and prototype visualizations show that the learned representation is discriminative and interpretable.