[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85620-en":3,"doc-seo-85620-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85620,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","PerspectiveGap: A Benchmark for Multi-Agent Orchestration Prompting","Real-world LLM applications increasingly rely on orchestrated multi-agent systems rather than single-agent workflows, yet existing models often cannot determine what information each sub-agent should receive. PerspectiveGap introduces a benchmark to evaluate LLMs’ ability to generate orchestration prompts. It includes 110 scenarios, each assessed with two distractor-mixed formats—role-fragment assignment and free-form sub-agent prompt writing—organized into 10 topologies under the Prompt Economy principle. Experiments on 33 commercial models show GPT-5.5 leads, while all models still achieve a low average pass rate of 17.2% and high leakage behavior.","PerspectiveGap: A Benchmark for Multi-Agent Orchestration Prompting  \nYouran Sun 1 ,∗ Xingyu Ren2 ,∗ Kejia Zhang 1 Xinpeng Liu Jiaxuan Guo3 ,†  \narXiv :2606 .08878v2 [ cs .CL] 12 Jul 2026  \nAbstract  \nReal-world LLM applications are moving beyond single-agent workflows toward orchestrated multi-agent systems, yet current models still struggle to determine what each subagent needs to know. To measure this, we introduce PerspectiveGap, a benchmark for evaluating LLMs’ ability to compose orchestration prompts for multi-agent systems. PerspectiveGap contains 110 scenarios, each evaluated through two distractor-mixed task formats: role-fragment assignment and free-form prompt writing. These scenarios are organized into 10 topologies, which are distilled from the authors’ real-world engineering practice and framed by the Prompt Economy principle: building loop-centered orchestrations that maximize utility with minimal role and engineering overhead. In experiments with 33 commercial models from 10 companies, GPT- 5.5 substantially outperforms all competitors, whereas Opus 4.8 shows a notable weakness in orchestration prompting despite its strong coding performance. Nevertheless, PerspectiveGap remains challenging: the evaluated models achieve an average combined pass rate of only 17.2%(GPT-5.5 62.0%) and an average overall leakage rate of 217.9%(a per-scenario information leak-event count, not a proportion; GPT- 5.5 49.1%) . These findings suggest that multiagent orchestration prompting is a distinct and under-evaluated capability, and PerspectiveGap provides a foundation for measuring and improving it systematically.  \n1 Introduction  \nPrompt engineering has shifted from tuning a single prompt to designing multi-agent orchestras, in which a task is decomposed into specialized roles with distinct information boundaries and artifact handoffs (Li et al., 2023 ; Qian et al., 2023 ; Hong et al., 2023 ; Wu et al., 2023a ; Chen et al., 2023 ; Lu et al., 2024 ; Tran et al., 2025) .  \nConstructing these orchestras requires orchestration prompting: writing sub-agent instructions  \n1University of Maryland. 2The Chinese University of Hong Kong. 3 Stanford University. ∗Equal contribution. †Corresponding author. Emails: Youran Sun, [sun1245@umd.edu](sun1245@umd.edu); Jiaxuan Guo, [guojx@stanford.edu](guojx@stanford.edu).  \nthat specify each role’s task scope, context boundaries, and expected handoffs. Yet current LLMs struggle to determine what each sub-agent needs to know. The resulting failures are not cosmetic: main agents leak distractors, expose out-of-role information, drop shared context, confuse artifact ownership, and sometimes place instructions where the sub-agent cannot see them. These errors produce incomplete, contaminated, or self-defeating sub-agent prompts. Section 6 analyzes these failure modes in detail.  \nExisting evaluations do not target this orchestration-prompting ability. Theory-of-mind (ToM) benchmarks typically score question answering or belief tracking (Le et al., 2019 ; Kim et al., 2023 ; Sclar et al., 2024), multiple-choice action prediction (Gu et al., 2024), dialogue acts (YS et al., 2026), or functional behavior labels (Riemer et al., 2025) . Agent benchmarks instead score downstream task success, tool use, environment-level performance, or instruction following (Liu et al., 2023 ; Qin et al., 2023 ; Orogat et al., 2026 ; Qi et al., 2025) . Neither line directly evaluates whether a main agent can write sub-agent prompts that respect asymmetric context and role-specific information needs. Section 2 discusses this boundary.  \nWe introduce PerspectiveGap, to our knowledge the first benchmark for multi-agent orchestration prompt writing. PerspectiveGap contains 110 scenarios, each consisting of a role list, a shuffled set of information fragments f1 ,..., fN , and a reference answer specifying which fragments each role needs. Each scenario includes two tasks: rolefragment assignment and free-form prompt writing. The ","cbCaic5TzSecstM2","https://ap.wps.com/l/cbCaic5TzSecstM2","pdf",5652030,4,1,20,"English","en",105,"# Abstract\n# Introduction\n## Orchestration prompting and failure modes\n## Related work and evaluation gap\n## PerspectiveGap benchmark design\n## Scenario topologies and Prompt Economy framing\n## Experimental evaluation results","[{\"question\":\"What capability does PerspectiveGap evaluate?\",\"answer\":\"PerspectiveGap evaluates whether an LLM can write orchestration prompts for multi-agent systems that respect each sub-agent’s task scope, context boundaries, and information handoffs.\"},{\"question\":\"How are the 110 scenarios in PerspectiveGap structured?\",\"answer\":\"Each scenario includes a role list, shuffled information fragments, and a reference answer mapping which fragments each role needs, plus an injected distractor fragment that may seem helpful but is not.\"},{\"question\":\"What are the two task formats used for each scenario?\",\"answer\":\"PerspectiveGap uses role-fragment assignment and free-form prompt writing. Both tasks share the same fragments, enabling analysis of models that identify role needs but fail to apply them in the written prompts.\"}]",1784204972,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"perspectivegap-a-benchmark-for-multi-agent-orchestration-prompting","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/perspectivegap-a-benchmark-for-multi-agent-orchestration-prompting/85620/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What capability does PerspectiveGap evaluate?","Question",{"text":75,"@type":76},"PerspectiveGap evaluates whether an LLM can write orchestration prompts for multi-agent systems that respect each sub-agent’s task scope, context boundaries, and information handoffs.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are the 110 scenarios in PerspectiveGap structured?",{"text":80,"@type":76},"Each scenario includes a role list, shuffled information fragments, and a reference answer mapping which fragments each role needs, plus an injected distractor fragment that may seem helpful but is not.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the two task formats used for each scenario?",{"text":84,"@type":76},"PerspectiveGap uses role-fragment assignment and free-form prompt writing. Both tasks share the same fragments, enabling analysis of models that identify role needs but fail to apply them in the written prompts.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]