[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86331-en":3,"doc-seo-86331-105":30,"detail-sidebar-cat-0-en-105":96},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86331,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","MM-ToolSandBox：用于评估视觉工具调用智能体的统一框架","MM-ToolSandBox is a benchmark and evaluation framework for visually grounded tool-calling agents, designed around a stateful environment with 500+ tools across 16 application domains. It supports multi-image, multi-turn tasks where agents must progressively ground arriving visual inputs into executable tool calls while handling realistic conversation dynamics such as goal revisions, error corrections, and state mutations. An automated scenario generation pipeline creates diverse visual scenarios, and evaluations on 12 state-of-the-art models show success rates below 50%, with visual precision as a key bottleneck.","arXiv :2607 . 11818v1 [ cs .CV] 13 Jul 2026  \nMM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents  \nKaixin Ma∗, Di Feng∗, Alexander Metz, Jiarui Lu, Eshan Verma, Afshin Dehghan Apple  \n∗ Equal Contribution  \nWe introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a stateful execution environment spanning 500+ tools across 16 application domains, supporting multi-image, multi-turn tasks where agents must ground progressively arriving visual inputs into executable tool calls while handling realistic conversational phenomena (goal revisions, error corrections, state mutations) . An automated scenario generation pipeline produces diverse, visually grounded scenarios through information-flow-guided planning and multi-stage quality filtering, yielding 258 human-verified nominal scenarios and 50 variants targeting interactive UI applications. Evaluating 12 state-of-the-art models, from 4B open-weight to frontier proprietary systems, shows that current models still lack robust visual tool-calling capability: even the best model achieves below 50% success rate. Our failure analysis further reveals that visual precision, not only planning, is a primary bottleneck for capable models: 53% of failures stem from incorrect information extraction from images despite otherwise correct task workflows. A planning-to-precision crossover emerges with scale: smaller models fail at deciding what to do, while larger models fail at perceiving what they see, suggesting fundamentally different research directions for improving models at different capability levels. The framework and the benchmark are publicly available at [https://github.com/apple/ml-mmtoolsandbox](https://github.com/apple/ml-mmtoolsandbox).  \nDate: July 14, 2026  \n1 Introduction  \nEvaluating LLM-based agents on tool-augmented tasks has become an active area of research, with benchmarks targeting function calling (Li et al. , 2023 ; Patil et al. , 2025), stateful environment interaction (Lu et al. , 2025 ; Trivedi et al. , 2024), and multi-turn conversational dynamics (Froger et al. , 2026 ; Barres et al. , 2025 ; Xiu et al. , 2026) . However, most of these frameworks remain text-centric, and cannot evaluate many realistic assistant tasks with visual inputs, for example, a user may share a screenshot to diagnose an issue or provide a photo of an event poster to schedule an event.  \nSolving visual tool-calling tasks requires more than visual recognition. The agent must extract task-relevant information from images, map visual evidence to the correct tool or code action, execute that action against astateful environment, and continue the interaction when the user’s intent evolves. This makes visual tool calling a distinct agentic capability: the model must not only understand what is shown, but also decide how to act on it through tools.  \nAmong the few benchmarks that incorporate visual inputs, images are typically provided as a static prefix in the initial query (Wang et al. , 2024a), or the tasks are limited to single-round user-agent interactions (Xie et al. , 2024 ; Rawles et al. , 2025) . Moreover, many existing benchmarks operate with relatively small tool spaces, often containing only 15–100 tools (Froger et al. , 2026 ; Wang et al. , 2024a ; Kong et al. , 2025 ; Lu et al. , 2025) . As a result, there is no unified framework and benchmark for systematically studying the visual tool-calling capability of multimodal agents under dynamic, multi-turn, and large-tool-space settings.  \nTo close this gap, we introduce MM-ToolSandBox, a diverse simulation framework for evaluating visual tool-calling agents. MM-ToolSandBox extends the stateful tool-use environments from ToolSandBox (Lu et al. , 2025), with explicit image handling, visual-specific tools, and support for both structured tool-use  \nEnvironment  \n\n| Entity Database\u003Cbr>Images Documents Contacts Mails Transactions\u003Cbr>Service-pro","cbCaitENFfDOv7YG","https://ap.wps.com/l/cbCaitENFfDOv7YG","pdf",9961773,7,1,32,"English","en",105,"# Introduction\n## Motivation and Gap in Existing Benchmarks\n## Visual Tool-Calling as a Distinct Capability\n## MM-ToolSandBox Overview","[{\"question\":\"MM-ToolSandBox评估的核心能力是什么？\",\"answer\":\"它评估具备视觉扎根能力的工具调用智能体：从图像中提取与任务相关的信息，将视觉证据映射到正确的工具或代码动作，并在有状态环境中执行后继续多轮互动。\"},{\"question\":\"该框架如何处理动态的多轮对话与真实偏差？\",\"answer\":\"MM-ToolSandBox支持多图片、多轮任务，并显式建模目标修订、错误纠正和状态突变等对话现象，使评测更贴近真实使用过程。\"},{\"question\":\"实验结果显示当前模型存在什么主要瓶颈？\",\"answer\":\"失败分析表明瓶颈不仅在规划，还在视觉精度：约53%的失败来自图像信息提取错误，而工作流本身往往是正确的。\"},{\"question\":\"基准数据与场景是如何构建和校验的？\",\"answer\":\"通过信息流引导的规划与多阶段质量过滤生成多样化的视觉场景，得到经人工验证的名义场景，并使用专家进行修正与校验，以保证视觉扎根、可执行任务规格和完成标准可靠。\"}]",1784210522,81,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":91,"head_meta":93,"extra_data":95,"updated_unix":28},"mm-toolsandbox-a-unified-framework-for-evaluating-visual-tool-calling-agents","",{"@graph":36,"@context":90},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/mm-toolsandbox-a-unified-framework-for-evaluating-visual-tool-calling-agents/86331/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82,86],{"name":73,"@type":74,"acceptedAnswer":75},"MM-ToolSandBox评估的核心能力是什么？","Question",{"text":76,"@type":77},"它评估具备视觉扎根能力的工具调用智能体：从图像中提取与任务相关的信息，将视觉证据映射到正确的工具或代码动作，并在有状态环境中执行后继续多轮互动。","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"该框架如何处理动态的多轮对话与真实偏差？",{"text":81,"@type":77},"MM-ToolSandBox支持多图片、多轮任务，并显式建模目标修订、错误纠正和状态突变等对话现象，使评测更贴近真实使用过程。",{"name":83,"@type":74,"acceptedAnswer":84},"实验结果显示当前模型存在什么主要瓶颈？",{"text":85,"@type":77},"失败分析表明瓶颈不仅在规划，还在视觉精度：约53%的失败来自图像信息提取错误，而工作流本身往往是正确的。",{"name":87,"@type":74,"acceptedAnswer":88},"基准数据与场景是如何构建和校验的？",{"text":89,"@type":77},"通过信息流引导的规划与多阶段质量过滤生成多样化的视觉场景，得到经人工验证的名义场景，并使用专家进行修正与校验，以保证视觉扎根、可执行任务规格和完成标准可靠。","https://schema.org",{"og:url":52,"og:type":92,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":94,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":97},[98,102,106,110,115,120,124,127,132,135,139],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Exam",70,"exam",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},5,"Comic",60,"comic",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":125,"slug":126},30,"research-report",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":130,"slug":131},9,"Religion & Spirituality",20,"religion-spirituality",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":130,"slug":134},"World Cup","world-cup",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":136,"slug":138},10,"Lifestyle","lifestyle",{"id":140,"doc_module":4,"doc_module_name":46,"category_name":141,"show_sort_weight":111,"slug":142},19,"General","general"]