[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86146-en":3,"doc-seo-86146-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86146,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","VIA: Visual Interface Agent for Robot Control","Robot manipulation demands visual understanding, physical reasoning, planning, and reliable closed-loop control. Foundation models show strong vision and reasoning, yet standard vision-language-action fine-tuning often yields much smaller models that struggle with long-horizon physical skills. VIA reframes robot control as an agentic visual tool-use task: an off-the-shelf model-powered agent operates a browser-based 3D interface using screenshots, intuitive commands, observation, and iterative re-planning, without robot-specific tuning or privileged state access, achieving strong zero-shot performance on multiple tabletop benchmarks.","arXiv :2607 . 11119v1 [ cs .RO] 13 Jul 2026  \nVIA: VISUAL INTERFACE AGENT FOR ROBOT CONTROL  \nHengyuan Hu, Priya Sundaresan, Jensen Gao, Dorsa Sadigh  \nStanford University  \nABSTRACT  \nRobot manipulation is a complex task that requires visual understanding, physical reasoning, planning, and closed-loop control. General-purpose foundation models (FMs) have grown remarkably capable of some of these, especially vision and reasoning. To leverage this for generalist robot policies, current methods typically involve converting existing FMs into vision-language-action (VLA) models by fine-tuning on robot data to output low-level actions. However, VLAs are often orders of magnitude smaller than frontier FMs given the limited data and compute available for fine-tuning, which in turn limits their general capability. Inspired by the growing ability of FMs to operate software through visual interfaces, we ask whether that same competence suffices to control a robot. We present VIA (Visual Interface Agent for robot control), a framework that recasts robot control as anagentic task: an off-the-shelf FM-powered agent drives a manipulator through a browser-based 3D interface by taking screenshots, issuing intuitive commands, observing the outcome, and adjusting. The agent receives no robot-specific finetuning and no access to privileged state information: it perceives visual input and acts through a small set of general tools. VIA inherits the agent’s general reasoning, closed-loop error recovery, and ability to plan and re-plan from what it observes. It solves a diverse suite of tabletop manipulation tasks zero-shot with both Claude Code and Codex. With the strongest model (Fable 5) it achieves 96. 7% success on three LIBERO-Goal tasks and 100% on a long-horizon rainbow assembly task. Performance improves with the scale and strength of the underlying model. These results suggest that frontier agents already possess skills that transfer directly to robot control given the right interface: your coding or computer-use agent is, in a sense, secretly a robot-control agent.  \n1 INTRODUCTION  \nGeneral robot manipulation is a long-standing research goal at the intersection of robotics and learning. It requires policies that are capable of complex visual understanding, physical reasoning, longterm planning, and precise closed-loop control. Foundation models (FMs) (Radford et al., 2019 ; Bommasani et al., 2021 ; OpenAI, 2023) trained on internet-scale data have become remarkably capable of several of these, such as visual perception and reasoning, which has made leveraging them for robotic manipulation highly attractive.  \nOne approach for this is to convert existing vision-language models into vision-language-action (VLA) models, by fine-tuning them on robot data to produce low-level actions (Zitkovich et al., 2023 ; Black et al., 2025) . However, due to a lack of both compute and robot-specific data, VLAs are often orders of magnitude smaller than frontier FMs. This gap has a profound impact on capabilities, particularly the physical reasoning and long-horizon planning skills that current robot policies lack.  \nAnother paradigm is Code-as-Policies (CaP) (Liang et al., 2023 ; Singh et al., 2023), where FMs are prompted to generate programs that call perception APIs and hand-crafted skill primitives, e.g., mask = segment_text_prompt(\"green cube\"), grasp_pose = sample_grasp_pose(\"red cube\") (Fu et al., 2026) . Although this is a natural way to tap into the strong coding capabilities ofFMs, the FM usually does not perceive the scene directly, and the system is often bottlenecked by human-crafted abstractions rather than by the FM. Fu et al. (2026) find that CaP, even when built on state-of-the-art FMs, remains dependent on these abstractions, as success rates degrade sharply when the high-level primitives are replaced by low-level ones.  \n[Correspondence to hengyuan.hhu@gmail.com](Correspondence to hengyuan.hhu@gmail.com)  \nWeb UI —3D scene reconstructe","cbCairuvRv11daLQ","https://ap.wps.com/l/cbCairuvRv11daLQ","pdf",2729878,3,1,20,"English","en",105,"# Abstract\n# Introduction\n## Motivation and Background\n## Limitations of VLA and Code-as-Policies\n## VIA Framework Overview","[{\"question\":\"What problem does VIA address in robot manipulation?\",\"answer\":\"VIA addresses the difficulty of transferring foundation-model capabilities to general robot manipulation, especially long-horizon planning and physical reasoning under limited fine-tuning data and compute.\"},{\"question\":\"How does VIA control a robot without robot-specific fine-tuning or privileged state access?\",\"answer\":\"VIA uses an off-the-shelf agent that observes through a browser-based 3D interface (screenshots and camera feeds) and acts via a small set of human-legible tools, then re-plans in a multi-round observe-act loop.\"},{\"question\":\"What are VIA’s reported results on tabletop tasks?\",\"answer\":\"With the strongest model mentioned, VIA reaches 96.7% success on three LIBERO-Goal tasks and 100% success on a long-horizon rainbow assembly task, with performance improving as the underlying model size/strength increases.\"}]",1784208899,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"via-visual-interface-agent-for-robot-control","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/via-visual-interface-agent-for-robot-control/86146/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does VIA address in robot manipulation?","Question",{"text":75,"@type":76},"VIA addresses the difficulty of transferring foundation-model capabilities to general robot manipulation, especially long-horizon planning and physical reasoning under limited fine-tuning data and compute.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does VIA control a robot without robot-specific fine-tuning or privileged state access?",{"text":80,"@type":76},"VIA uses an off-the-shelf agent that observes through a browser-based 3D interface (screenshots and camera feeds) and acts via a small set of human-legible tools, then re-plans in a multi-round observe-act loop.",{"name":82,"@type":73,"acceptedAnswer":83},"What are VIA’s reported results on tabletop tasks?",{"text":84,"@type":76},"With the strongest model mentioned, VIA reaches 96.7% success on three LIBERO-Goal tasks and 100% success on a long-horizon rainbow assembly task, with performance improving as the underlying model size/strength increases.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]