Skip to main content
eScholarship
Open Access Publications from the University of California

Coordinated Minorities Exploit Social Influence in Multi-Agent Debate Systems

Creative Commons 'BY' version 4.0 license
Abstract

Multi-agent debate (MAD) systems improve collective decision quality by aggregating diverse perspectives. However, this open interaction introduces security vulnerabilities. While existing research primarily focuses on single-agent attacks, threats from adversarial coalitions remain underexplored. Therefore, we propose the Dynamic Cognitive Manipulation Attack (DCMA), a framework to investigate how adversarial coalitions exploit sociocognitive dynamics to manipulate group consensus. DCMA consists of two components: a coordination mechanism that dynamically synchronizes coalition intentions through a perception-refinement-execution loop, and a strategic engine that selects appropriate persuasion strategies based on debate context. Empirical evaluations across four large language models validate the attack's effectiveness, with Attack Success Rate (ASR) reaching up to 22.9% and convergence delay nearly doubling compared to uncoordinated attacks. Our analysis also reveals three sociocognitive vulnerabilities: domain vulnerability asymmetry, latent process contamination, and scale-amplified vulnerability. Finally, we propose detection schemes and identify future research directions to address the limitations of existing defenses.