L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning

2026-07-10Artificial Intelligence

Artificial Intelligence
AI summary

The authors explore how having multiple expert agents debate helps in legal reasoning tasks, which require understanding complex and detailed information. They create a new system called Legal Multi-Agent Debate (L-MAD) that assigns different expert roles to each agent. This approach improves performance compared to using just one agent. They also find that adding more agents helps reduce mistakes, but having too many discussion rounds can cause agents to repeat errors. Their work helps define practical limits for using multi-agent debates in sensitive legal settings.

Multi-Agent DebateLegal Textual EntailmentExpert PersonasAgent AggregationOver-deliberation DriftCollaborative AILegal ReasoningInconsistencyAccuracyHigh-Stakes Environments
Authors
Tan-Minh Nguyen, Hoang-Trung Nguyen, Huu-Dong Nguyen, Dinh-Truong Do, Thi-Hai-Yen Vuong, Le-Minh Nguyen
Abstract
While multi-agent debate (MAD) frameworks have shown significant potential in general reasoning, their effectiveness in highly structured, knowledge-heavy legal domains remains under-explored. In this work, we introduce the Legal Multi-Agent Debate (L-MAD) framework to systematically evaluate different debate structures and aggregation methods within Legal Textual Entailment. By assigning distinct expert personas to multiple agents, L-MAD improves upon strong single-agent baselines by up to 8\%. Furthermore, analyzing how debate scales reveals a clear trade-off: increasing the agent population reduces inconsistency and improves accuracy, whereas extending discussion rounds induces a detrimental \textit{over-deliberation drift} where agents reinforce each other's mistakes. Ultimately, our findings outline the practical boundaries and safety margins of deploying collaborative multi-agent systems in high-stakes legal reasoning environments.