Multimodal Attention Transformer Framework for Parkinson’s Disease Detection Using Gait, Speech, Wearable Sensors, and Auxiliary Neuroimaging
Keywords:
Parkinson’s disease detection; multimodal artificial intelligence; medical imaging-supported AI; digital biomarkers; gait analysis; speech biomarkers; wearable sensors; Transformer model; attention-based fusion; explainable AI.Abstract
Early detection of Parkinson's disease (PD) is complicated by the insidious, heterogeneous nature of early symptoms that are incompletely measured with a single biomarker. This study introduces a multimodal artificial intelligence framework that combines gait, speech and wearable sensor data as primary digital biomarkers supported by an auxiliary validation branch for Parkinson's Progression Markers Initiative (PPMI) neuroimaging. These models included modality-specific and fusion-based models that were developed using publicly available datasets. Locomotor features were extracted from gait signals, acoustic biomarkers from speech data, movement variability metrics from wearable signals and imaging-derived representations from neuroimaging data. Several temporal regression models were explored: CNN, LSTM, CNN-LSTM, Transformer,CNN-Transformer and a multimodal attention-based transformer model. The gait-speech-wearable multimodal Transformer proposed here provided 93.1% accuracy and 0.965 AUC, which was superior to single-modality and conventional deep learning models with body-worn sensors alone. The highest performances were achieved by using the imaging-supported multimodal Transformer, with 94.3% accuracy and 0.973 AUC. Explainability of attention demonstrated that complementary diagnostic features from gait, speech and imaging were provided by wearable data. This approach may inform future decision-support systems in screening and monitoring for Parkinson’s disease, but clinical deployment would require validation against matched prospective multimodal datasets.





