[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86225-en":3,"doc-seo-86225-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86225,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure","Recent advances in large language models (LLMs) and vision-language models (VLMs) enable capable digital agents for computer use, yet long-horizon tasks often accumulate evolving contexts with observations, edits, failures, and partial executions. Existing agents rely on raw interaction history, hindering interpretation, verification, and recovery. This paper proposes StructAgent, a state-centered framework using a unified causal representation for verifiable task progress, verifier-backed state transitions, checkpointing, evidence-driven completion, targeted failure recovery, and tool-supported execution.","arXiv :2607 . 1 1388v 1 [ cs .AI] 13 Jul 2026  \nStructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure  \nWenyi Wu 1 ,2 , § , Sibo Zhu 1 ,2 , § , Kun Zhou2 ,†, Aayush Salvi 1 , Zixuan Song 1 , Biwei Huang 1 ,2  \n1University of California, San Diego 2Aether AI Lab  \n§ Work done during an internship at Aether AI.  \n†Project Lead & Corresponding Author: [franciskunzhou@gmail.com](franciskunzhou@gmail.com).  \nAbstract  \nRecent advances in large language models (LLMs) and vision-language models (VLMs) have enabled increasingly capable digital agents for computer use.  \nHowever, real-world tasks are often long-horizon and involve evolving contexts containing accumulated observations, intermediate edits, failed attempts, and partially completed executions. Existing agents typically operate over raw interaction history, making task progress difficult to interpret, verify, and recover, which ultimately limits reliable long-horizon execution. In this paper, we argue that addressing this challenge requires explicitly structuring both the agent’s state and workflow around a unified causal representation of task progress. We present StructAgent, a state-centered framework that introduces a unified state for maintaining compact, verifiable task progress and a structured workflow that regulates progress through verifier-backed state transitions. Building on this design, StructAgent further enables explicit progress checkpointing, evidence-driven task completion, targeted failure recovery, and tool-supported execution, while ensuring that all progress updates remain grounded in verification. Extensive experiments demonstrate that StructAgent consistently improves a wide range of LLM and VLM backbones on long-horizon computer-use tasks. On OSWorld-Verified, it improves Qwen3.5-9B from 27.0% to 46.9% success rate and Qwen3.5-27B from 31.6% to 62.2%, while achieving a new open-source state of the art of 78.9% with MiniMaxM3 . Moreover, the same framework generalizes beyond desktop environments to Minecraft, demonstrating the generality of our design.  \n Project Page  Code  \n“The purpose of abstracting is not to be vague, but to create a new semantic level in which one can be  \nabsolutely precise.” —Edsger W. Dijkstra  \n1 Introduction  \nDriven by scaling laws, large language models (LLMs) and vision-language models (VLMs) have become increasingly capable, demonstrating stronger abilities in perception, reasoning, and planning [1, 2] . Building on these advances, digital agents have emerged as a promising paradigm for automating user tasks in computer environments [3, 4] . It completes user requests by observing the digital environment through visual screenshots or accessibility information, taking actions such as clicking, typing, and invoking system tools when necessary [5–8] .  \nPreprint.  \nFigure 1: Results on OSWorld-Verified, Mind2Web, and Minecraft. StructAgent improves matched open backbones and reaches 78.9% with MiniMax-M3, the state-of-the-art open-source model result.  \nIn real-world applications, however, digital agents must execute long-horizon, complex workflows involving information retrieval, document editing, cross-application coordination, while preserving intermediate results across many steps. This setting is particularly challenging because the agent must continually manage an evolving context that accumulates large amounts of factual information, intermediate edits, failed attempts, and partially completed executions. As this context grows, critical information can become buried and the current task state increasingly ambiguous. Consequently, the agent struggles to maintain a clear causal understanding of task progress, making it difficult to promptly detect abnormal states, identify intermediate errors, and perform accurate planning and execution [8, 9] .  \nTo solve it, the agent’s working process must be explicitly organized to support tracing the effects of individual decisions, identifying the key ","cbCaip5x1WRfQfu0","https://ap.wps.com/l/cbCaip5x1WRfQfu0","pdf",7630871,6,1,32,"English","en",105,"# Abstract\n# Introduction\n## Problem: long-horizon execution and ambiguous context\n## Solution: unified state and structured workflow\n## StructAgent contributions\n## Experimental results","[{\"question\":\"为什么长程数字代理在真实任务中难以可靠执行？\",\"answer\":\"真实任务包含不断演化的上下文，积累了观测、编辑、失败尝试和未完成执行，导致关键信息被淹没、任务状态变得模糊，从而难以保持清晰的因果理解与正确规划执行。\"},{\"question\":\"StructAgent 如何通过统一因果结构改善任务进展的可解释性和可验证性？\",\"answer\":\"StructAgent引入统一状态来维护可共享且可验证的任务进展表征，并通过结构化工作流建立固定执行循环，由规划器、执行器和验证器驱动状态更新，从而保留透明因果链。\"},{\"question\":\"StructAgent 在实验中取得了哪些主要效果？\",\"answer\":\"在OSWorld-Verified上，StructAgent将Qwen3.5-9B的成功率从27.0%提升到46.9%，将Qwen3.5-27B从31.6%提升到62.2%，并在MiniMax-M3上达到78.9%的开源状态下的最新水平；同时在Minecraft等环境中也体现了框架泛化能力。\"}]",1784209621,81,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"structagent-harness-long-horizon-digital-agents-with-unified-causal-structure","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/structagent-harness-long-horizon-digital-agents-with-unified-causal-structure/86225/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"为什么长程数字代理在真实任务中难以可靠执行？","Question",{"text":76,"@type":77},"真实任务包含不断演化的上下文，积累了观测、编辑、失败尝试和未完成执行，导致关键信息被淹没、任务状态变得模糊，从而难以保持清晰的因果理解与正确规划执行。","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"StructAgent 如何通过统一因果结构改善任务进展的可解释性和可验证性？",{"text":81,"@type":77},"StructAgent引入统一状态来维护可共享且可验证的任务进展表征，并通过结构化工作流建立固定执行循环，由规划器、执行器和验证器驱动状态更新，从而保留透明因果链。",{"name":83,"@type":74,"acceptedAnswer":84},"StructAgent 在实验中取得了哪些主要效果？",{"text":85,"@type":77},"在OSWorld-Verified上，StructAgent将Qwen3.5-9B的成功率从27.0%提升到46.9%，将Qwen3.5-27B从31.6%提升到62.2%，并在MiniMax-M3上达到78.9%的开源状态下的最新水平；同时在Minecraft等环境中也体现了框架泛化能力。","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]