[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83291-en":3,"doc-seo-83291-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83291,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Collective Intelligence with Foundation Models","Coordinating multiple foundation models into cooperative, multi-agent reasoning systems supports safer and more globally reliable AI as model scale and diversity increase. The chapter proposes a framework with solver agents producing independent drafts, a dedicated critic agent performing structured critique and revision, and an aggregator agent synthesizing final consensus. A scoring module performs semantic, numerical, and procedural evaluation across agents. Ablation studies over calculus, physics, chemistry, biology, economics, optimization, statistics, and mathematics isolate architecture versus model diversity, showing heterogeneity drives substantial accuracy gains and improves intermediate step quality, strengthening explainability and auditability for global applied AI.","arXiv :2607 .07729v 1 [ cs .MA] 6 Jul 2026  \nCollective Intelligence with Foundation Models  \nJ. de Curt`o1,2 and [I. de](I. de) Zarz`a3  \n1 Department of Computer Applications in Science & Engineering, BARCELONA Supercomputing Center, 08034 Barcelona, Spain  \n2 Escuela T´ecnica Superior de Ingenier´ıa (ICAI), Universidad Pontificia Comillas, 28015 Madrid, Spain  \n3 Human Centered AI, Data & Software, LUXEMBOURG Institute of Science and Technology, L-4362  \nEsch-sur-Alzette, Luxembourg  \nAbstract. As foundation models grow in scale and diversity, coordinating multiple models into cooperative reasoning systems offers a promising path toward safer and more globally reliable AI. This chapter presents a multi-agent reasoning framework where multiple solver models generate independent drafts, each undergoes structured critique and revision by a dedicated critic agent, and a higher-level aggregator agent synthesizes a final consensus solution. A comprehensive scoring module provides semantic, numerical, and procedural evaluation across all agents. Through systematic ablation studies using a comprehensive benchmark spanning calculus, physics, chemistry, biology, economics, optimization, statistics, and mathematics, we isolate the distinct contributions of framework architecture versus model diversity. We compare four configurations: (1) Individual Baseline with no multi-agent framework,(2) Homogeneous Framework where all agents employ the same model, (3) Redundant Homogeneous Solvers using multiple instances of identical models, and (4) Heterogeneous Framework with diverse specialized models. Our results demonstrate that while framework structure and redundant sampling provide modest improvements, model heterogeneity emerges as the critical factor driving substantial performance gains. The heterogeneous configuration achieves superior step-wise accuracy (0.64 vs. 0.54 average for individual models and 2.3 × improvement over homogeneous configurations) with reduced variance across problem categories and difficulty levels. Notably, step-wise reasoning quality, which measures the correctness of intermediate reasoning steps rather than merely final answers, improves dramatically only with model diversity, indicating that heterogeneous agents provide complementary error detection and reasoning refinement capabilities essential for explainability and auditability. We discuss architectural principles, evaluation methodology, and implications for the future of Global Applied AI, highlighting how heterogeneous multi-agent coordination can support transparent, auditable, and high-confidence decision making across scientific and industrial domains.  \nKeywords: foundation models, multi-agent systems, reasoning frameworks, large language models, AI safety, consensus mechanisms  \n1 Introduction  \nThe rapid advancement of foundation models has fundamentally transformed artificial intelligence capabilities across diverse domains [3] . Large Language Models (LLMs) such as GPT-4 [15], Claude [1], and Llama [16] have demonstrated remarkable proficiency in natural language understanding, reasoning, and generation tasks. However, individual models exhibit inherent limitations including hallucinations, reasoning errors, biases, and inconsistent performance across problem domains [20,11] .  \nMulti-agent systems represent a promising paradigm shift in leveraging foundation model capabilities [19,10,4] . Rather than relying on a single model’s output, multi-agent architectures enable multiple models to collaborate, critique, and synthesize solutions collectively [8] . This approach offers several key advantages:  \n– Error Mitigation: Multiple independent reasoning paths reduce the probability of systematic errors propagating to final outputs.  \n– Diverse Perspectives: Different models trained on varied datasets bring complementary strengths and reasoning approaches.  \n– Consensus Building: Aggregation mechanisms can identify high-confidence solutions while f","cbCaicIIxD1CAMrz","https://ap.wps.com/l/cbCaicIIxD1CAMrz","pdf",1193887,3,1,12,"English","en",105,"# Introduction\n# Multi-Agent Reasoning Framework\n## Framework Components\n## Evaluation Setup\n# Contributions","[{\"question\":\"How does the proposed multi-agent framework produce a final consensus solution?\",\"answer\":\"Solver agents generate independent solution drafts, the critic agent performs structured critique and proposes revisions, and the aggregator agent synthesizes a consensus from multiple perspectives.\"},{\"question\":\"What role does the scoring module play in the framework?\",\"answer\":\"It provides comprehensive evaluation across agents, including semantic assessment, numerical correctness, and procedural validation.\"},{\"question\":\"Why do the results emphasize heterogeneous model diversity?\",\"answer\":\"Ablation comparisons show that while framework structure and redundant sampling yield modest gains, heterogeneous specialized models produce the largest improvements, including better step-wise reasoning quality and reduced variance across categories and difficulty levels.\"}]",1784186527,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"collective-intelligence-with-foundation-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/collective-intelligence-with-foundation-models/83291/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the proposed multi-agent framework produce a final consensus solution?","Question",{"text":75,"@type":76},"Solver agents generate independent solution drafts, the critic agent performs structured critique and proposes revisions, and the aggregator agent synthesizes a consensus from multiple perspectives.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What role does the scoring module play in the framework?",{"text":80,"@type":76},"It provides comprehensive evaluation across agents, including semantic assessment, numerical correctness, and procedural validation.",{"name":82,"@type":73,"acceptedAnswer":83},"Why do the results emphasize heterogeneous model diversity?",{"text":84,"@type":76},"Ablation comparisons show that while framework structure and redundant sampling yield modest gains, heterogeneous specialized models produce the largest improvements, including better step-wise reasoning quality and reduced variance across categories and difficulty levels.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]