[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82192-en":3,"doc-seo-82192-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82192,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","L-MAD A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning","Multi-agent debate (MAD) frameworks show promise for general reasoning, yet their effectiveness in structured, knowledge-heavy legal settings remains insufficiently studied. This work proposes Legal Multi-Agent Debate (L-MAD) to systematically evaluate debate structures and aggregation methods for Legal Textual Entailment. Distinct expert-persona agents outperform strong single-agent baselines by up to 8%. Scaling analysis reveals trade-offs: more agents reduce inconsistency and improve accuracy, while more rounds cause over-deliberation drift, amplifying errors and harming performance.","L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures  \nin Legal Reasoning  \nTan-Minh Nguyen * 1 Hoang-Trung Nguyen * 2 Huu-Dong Nguyen * 2 Truong Dinh Do 1 Thi-Hai-Yen Vuong 2  \nLe-Minh Nguyen 1  \narXiv :2607 .09099v 1 [ cs .AI] 10 Jul 2026  \nAbstract  \nWhile multi-agent debate (MAD) frameworkshave shown significant potential in general reasoning, their effectiveness in highly structured, knowledge-heavy legal domains remains underexplored. In this work, we introduce the Legal Multi-Agent Debate (L-MAD) framework to systematically evaluate different debate structures and aggregation methods within Legal Textual Entailment. By assigning distinct expert personasto multiple agents, L-MAD improves upon strong single-agent baselines by up to 8% . Furthermore, analyzing how debate scales reveals a clear tradeoff: increasing the agent population reduces inconsistency and improves accuracy, whereas extending discussion rounds induces a detrimental over-deliberation drift where agents reinforce each other’s mistakes. Ultimately, our findings outline the practical boundaries and safety margins of deploying collaborative multi-agent systems in high-stakes legal reasoning environments.  \n1. Introduction  \nLarge language models (LLMs) show remarkable capabilities in general reasoning, but they frequently struggle in high-stakes, knowledge-heavy domains that demand strict logic. Legal Textual Entailment (LTE) (Aoki et al., 2022 ; Bilgin et al., 2024) is a challenging test for these models, requiring them to determine whether specific legal statutes apply to a given factual scenario (Aletras et al., 2016 ; Zhong et al., 2020) . While specialized models like LegalBERT (Chalkidis et al., 2020) and Lawformer (Xiao et al., 2021) outperform generalist baselines, standard single-model inference is still limited by a model’s basic reasoning capacity and its tendency to make factual or logical errors.  \n1Japan Advanced Institute of Science and Technology (JAIST), Ishikawa, Japan 2VNU University of Engineering and Technology, Hanoi, Vietnam. Correspondence to: Le-Minh Nguyen \u003C[nguyenml@jaist.ac.jp](nguyenml@jaist.ac.jp) >.  \nPreprint. July 13, 2026.  \nFigure 1. Illustration of decision-making protocols used in this study.  \nMulti-Agent Debate (MAD) frameworks offer a promising way to overcome these limitations by allowing models to spend more computing power during inference (Nie et al., 2020 ; Choi et al., 2026) . By assigning distinct expert personas to multiple LLM agents, these frameworks let models iterate and refine their arguments to arrive at a better final answer (Kaesberg et al., 2025b) . However, how these debate frameworks perform in highly structured, rule-bound fields like law remains underexplored. While open-ended debate typically helps in general tasks, it introduces unique challenges in legal reasoning. The trade-offs between forcing agents to agree versus letting them vote independently—and the risk of errors multiplying over several rounds—have not yet been systematically analyzed.  \nTo bridge this gap, we introduce Legal Multi-Agent Debate (L-MAD), a framework designed to analyze different multiagent debate structures under strict legal rules (Figure 1) . Unlike prior work that mainly focuses on connecting LLMs to external tools, L-MAD provides a controlled environment to evaluate core aggregation methods. Specifically, we compare forced consensus strategies (where agents debate  \nuntil they agree) against voting approaches (where agents maintain independent reasoning paths) .  \nOur empirical results show that L-MAD significantly improves performance, outperforming strong single-agent baselines by up to 8% in some settings. Crucially, we find that the best debate strategy depends heavily on the underlying model’s capability. Forcing a consensus works well with highly capable models (e.g., 30B+ parameters) by encouraging constructive debate, whereas independent voting protects smaller models (e.g., 8B parameters) from s","cbCaib3mLOeE2zcm","https://ap.wps.com/l/cbCaib3mLOeE2zcm","pdf",531111,1,13,"English","en",105,"# Introduction\n# Preliminaries\n## Legal Textual Entailment","[{\"question\":\"What does L-MAD aim to evaluate in legal reasoning tasks?\",\"answer\":\"L-MAD systematically evaluates different multi-agent debate structures and aggregation methods for Legal Textual Entailment, focusing on how debate protocols affect decision-making under legal rules.\"},{\"question\":\"How does L-MAD perform compared with single-agent baselines?\",\"answer\":\"L-MAD improves performance and can outperform strong single-agent baselines by up to 8% in certain settings.\"},{\"question\":\"What scaling trade-off does the document report for L-MAD?\",\"answer\":\"Increasing the number of agents reduces inconsistency and improves accuracy, but increasing the number of discussion rounds induces over-deliberation drift, where agents reinforce each other’s mistakes and performance degrades.\"}]",1784178717,33,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"l-mad-a-systematic-evaluation-of-multi-agent-debate-structures-in-legal-reasoning","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/l-mad-a-systematic-evaluation-of-multi-agent-debate-structures-in-legal-reasoning/82192/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What does L-MAD aim to evaluate in legal reasoning tasks?","Question",{"text":74,"@type":75},"L-MAD systematically evaluates different multi-agent debate structures and aggregation methods for Legal Textual Entailment, focusing on how debate protocols affect decision-making under legal rules.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does L-MAD perform compared with single-agent baselines?",{"text":79,"@type":75},"L-MAD improves performance and can outperform strong single-agent baselines by up to 8% in certain settings.",{"name":81,"@type":72,"acceptedAnswer":82},"What scaling trade-off does the document report for L-MAD?",{"text":83,"@type":75},"Increasing the number of agents reduces inconsistency and improves accuracy, but increasing the number of discussion rounds induces over-deliberation drift, where agents reinforce each other’s mistakes and performance degrades.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]