Skip to main content
eScholarship
Open Access Publications from the University of California

Semantic-Aware Multimodal Fusion with Cross-Scale Features for Alzheimer's Disease Detection

Creative Commons 'BY' version 4.0 license
Abstract

With the increasing prevalence of Alzheimer's disease (AD), automatic detection has emerged as a practical auxiliary diagnostic tool. Communication-related signals, being non-invasive, have been widely used in unimodal and multimodal approaches for AD detection with promising results. However, existing methods face two limitations: (1) relying heavily on final-layer encoder representations prevents comprehensive characterization of modality-specific features, and (2) treating all modalities equally introduces noise and reduces discriminative capability. Therefore, we propose SCMF-Net, a multimodal fusion framework. Specifically, it incorporates a cross-scale feature alignment module to exploit complementary representations across encoder layers, producing comprehensive modality-specific representations. In addition, SCMF-Net integrates a semantic-aware attention module that uses textual semantics to guide cross-modal fusion, thereby promoting semantic alignment. Experiments on the ADReSS and ADReSSo benchmark datasets demonstrate the effectiveness of SCMF-Net, achieving accuracies of 91.67% and 90.14%, respectively, with ablation studies validating the contribution of each component.