- Main
Multi-Agent Debate under Reasoning Chain Attacks: Exploiting Cognitive Vulnerabilities
Abstract
With the advancement in large language model (LLM) capabilities, multi-agent debate has emerged as a crucial approach for tackling complex problems. However, security research on this front remains largely focused on traditional adversarial attacks that manipulate surface-level semantics, neglecting the deep reasoning pathways behind model decisions. To bridge this gap, we propose a novel reasoning chain debate attack that targets cognitive vulnerabilities within LLMs' internal reasoning chains during debates. This framework dynamically identifies cognitive vulnerabilities in target agents by deploying virtual agents with real-time evolutionary capabilities. Subsequently, it executes precise cognitive attacks using a weighted fusion attack strategy. We also design an adaptive termination mechanism that automatically halts debates when attack effectiveness stabilizes, thereby reducing computational overhead. Experimental results validate the framework's effectiveness: across mainstream LLMs, the attacker achieves over 56% persuasion success rates while reducing token consumption by approximately 27%, providing novel foundational research for multi-agent security.