[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86421-en":3,"doc-seo-86421-105":30,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86421,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","Multi-Agent Routing as Set-Valued Prediction: A WildChat Benchmark and Cost-Aware Evaluation","Tool and agent routing from natural-language prompts is treated as a set-valued prediction problem, where one query may need multiple agents and over-selection increases execution cost. The work introduces a WildChat-derived benchmark with 3,000 prompts over a fixed 12-agent catalog, using AI-assisted heuristic labels, controlled rebalancing, and multi-label set evaluation. The protocol combines Precision, Recall, F1, Jaccard, Exact Match, latency, simulation-based capability coverage, and constrained weighted routing across ordinal cost tiers. Results compare KNN, linear multi-label, dependency-aware baselines, fine-tuned encoders, WAR post-scoring, and zero-shot LLM routing, showing stronger accuracy and cost-aware utility.","Multi-Agent Routing as Set-Valued Prediction: A WildChat Benchmark and Cost-Aware Evaluation  \nAnanto Nayan Bala  \nAhsanullah University of Science and Technology Dhaka, Bangladesh [nayan.ananto@gmail.com](nayan.ananto@gmail.com)  \nFaisal Muhammad Shah  \nAhsanullah University of Science and Technology Dhaka, Bangladesh [faisal.cse@aust.edu](faisal.cse@aust.edu)  \narXiv :2606 .28925v2 [ cs .LG] 12 Jul 2026  \nAbstract  \nTool and agent routing from natural-language prompts is naturally a set-valued prediction problem: a single query may require multiple agents, while over-selection increases execution cost. The benchmark introduced here is derived from WildChat and contains 3,000 prompts over a fixed 12-agent catalog, with AI-assisted heuristic labels under a fixed schema and controlled rebalancing for multi-label evaluation. The evaluation protocol combines setlevel metrics (Precision, Recall, F1, Jaccard, and Exact Match), latency, an execution-oriented capability-coverage simulation, anda constrained weighted-routing setting based on ordinal agentcost tiers. Compared methods include nearest-neighbor matching, linear multilabel classification, dependency-aware baselines, a finetuned encoder, deterministic weighted post-scoring via Weighted Agent Routing (WAR), and a zero-shot LLM baseline. Results show that supervised routers substantially outperform nearest-neighbor and zero-shot LLM routing. The fine-tuned encoder achieves the strongest unconstrained set accuracy, while the linear multilabel model provides the strongest practical baseline. In the constrained setting, the weighted routing layer improves utility when applied on top of strong supervised scorers, with the largest gain observed for Encoder+WAR. Overall, the benchmark and evaluation protocol support reproducible study of accuracy–cost trade-offs in fixed-catalog multi-agent routing.  \n1 Introduction  \nModern AI systems increasingly rely on catalogs of tools or agents, where the system must select one or more agents to fulfill a user request (e.g., query a database, fetch an API, run statistical analysis, or generate a plot). This setting maps naturally to set-valued routing: given a query, the system predicts a small set of relevant agents that jointly satisfy the request. Unlike top-1 routing, this formulation captures real multi-step workflows and enables explicit trade-offs between coverage and execution cost.  \nDespite growing interest in tool-augmented assistants, thereis limited work that treats routing as a multi-label set prediction problem with set-level evaluation. Prior routing pipelines often select a single agent or rank tools without a principled decision rule for multi-agent execution. We address this gap by treating agent routing explicitly as set-valued prediction over a fixed inventory and evaluating it with standard set metrics and cost-aware utility. We build a WildChat-derived benchmark with controlled agent coverage and set-size distribution. Starting from real user prompts, we assign AI-assisted heuristic labels under a fixed 12-agent catalog and rebalance the pool for stable multi-label evaluation, then split it into train/dev/test partitions. Because prompt-to-agent routing  \ncan admit more than one defensible routed set depending on redundancy tolerance, cost sensitivity, and user preference, these labels are best interpreted as protocol-defined reference sets for comparative evaluation. We evaluate three families of methods: (i) content-based nearest neighbor retrieval,(ii) supervised multi-label classification, and (iii) a fine-tuned encoder that provides stronger semantic matching. We also study a cost-aware selection policy that trades off prediction quality and execution cost. Our contributions are:  \n• A set-valued prediction formulation of agent routing that makes multi-agent selection and cost-aware evaluation explicit.  \n• A WildChat-derived benchmark with real prompts, heuristic labels under a fixed 12-agent catalog, and controll","cbCaieEjfrduICLP","https://ap.wps.com/l/cbCaieEjfrduICLP","pdf",1383996,4,1,9,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n## Tool/Agent Routing in LLM Systems\n## Bundle/Slate and Set Selection","[{\"question\":\"为什么将多智能体路由建模为“集合值预测”而不是Top-1路由？\",\"answer\":\"因为一次用户查询可能需要多个智能体共同完成任务；集合化表述能显式刻画“覆盖率”与“执行成本”之间的权衡，而不是只选择单个最相关智能体。\"},{\"question\":\"该研究构建了怎样的WildChat-derived基准与标注方式？\",\"answer\":\"基准来自WildChat，包含3,000条提示，覆盖固定的12个智能体目录；使用在固定schema下的AI辅助启发式标注，并通过受控重平衡来稳定多标签评测，再划分训练/开发/测试集。\"},{\"question\":\"评测方法如何同时衡量准确性与成本？\",\"answer\":\"评测同时使用集合级指标（Precision、Recall、F1、Jaccard、Exact Match），并结合延迟与执行导向的能力覆盖模拟；在约束场景下还引入基于序数成本分档的加权路由（WAR）加权选择策略。\"},{\"question\":\"主要方法的对比结果表明了什么结论？\",\"answer\":\"在无约束场景下，微调编码器在集合准确性上表现最好，而线性多标签模型提供最强的实践基线；在约束加权路由场景中，将加权路由层叠加到强监督打分器上可提升效用，Encoder+WAR带来最大收益。\"}]",1784211654,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":28},"multi-agent-routing-as-set-valued-prediction-a-wildchat-benchmark-and-cost-aware-evaluation","",{"@graph":36,"@context":89},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/multi-agent-routing-as-set-valued-prediction-a-wildchat-benchmark-and-cost-aware-evaluation/86421/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"为什么将多智能体路由建模为“集合值预测”而不是Top-1路由？","Question",{"text":75,"@type":76},"因为一次用户查询可能需要多个智能体共同完成任务；集合化表述能显式刻画“覆盖率”与“执行成本”之间的权衡，而不是只选择单个最相关智能体。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"该研究构建了怎样的WildChat-derived基准与标注方式？",{"text":80,"@type":76},"基准来自WildChat，包含3,000条提示，覆盖固定的12个智能体目录；使用在固定schema下的AI辅助启发式标注，并通过受控重平衡来稳定多标签评测，再划分训练/开发/测试集。",{"name":82,"@type":73,"acceptedAnswer":83},"评测方法如何同时衡量准确性与成本？",{"text":84,"@type":76},"评测同时使用集合级指标（Precision、Recall、F1、Jaccard、Exact Match），并结合延迟与执行导向的能力覆盖模拟；在约束场景下还引入基于序数成本分档的加权路由（WAR）加权选择策略。",{"name":86,"@type":73,"acceptedAnswer":87},"主要方法的对比结果表明了什么结论？",{"text":88,"@type":76},"在无约束场景下，微调编码器在集合准确性上表现最好，而线性多标签模型提供最强的实践基线；在约束加权路由场景中，将加权路由层叠加到强监督打分器上可提升效用，Encoder+WAR带来最大收益。","https://schema.org",{"og:url":52,"og:type":91,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":93,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":96},[97,101,105,109,114,119,124,127,131,134,138],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},5,"Comic",60,"comic",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},6,"Technology",50,"technology",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":125,"slug":126},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":129,"slug":130},"Religion & Spirituality",20,"religion-spirituality",{"id":129,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":129,"slug":133},"World Cup","world-cup",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":135,"slug":137},10,"Lifestyle","lifestyle",{"id":139,"doc_module":4,"doc_module_name":46,"category_name":140,"show_sort_weight":110,"slug":141},19,"General","general"]