Skip to main content
eScholarship
Open Access Publications from the University of California

Multi-Granularity EEG Decoding: Bridging Object-Background Semantics and Temporal Motion for Video Reconstruction

Creative Commons 'BY' version 4.0 license
Abstract

Reconstructing video from brain signals is a pivotal task in neural decoding. However, current frameworks struggle by directly aligning noisy EEG with coarse global captions and VAE latents, inevitably leading to semantic misalignment and temporal flickering. To bridge this gap, we propose M-STAR, a bio-inspired framework that mimics human cognitive decomposition. To our knowledge, M-STAR is the first framework to decompose dense captions into object- and environment-level semantics for robust alignment and employ an "Anchor-plus-Residual" strategy that anchors generation to a stable visual anchor and applies residual dynamics to maintain stability over time. Specifically, this is achieved via the Factorized Semantic Module (FSM) and the Anchor-Guided Dynamics Module (ADM), respectively. Extensive evaluations on the SEED-DV dataset demonstrate that M-STAR significantly outperforms state-of-the-art methods, generating high-fidelity videos with precise semantics and coherent motion. M-STAR establishes a new baseline for fine-grained visual decoding, marking a significant step towards high-fidelity visual brain-computer interfaces.