[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-seo-148435-105":3,"detail-sidebar-cat-0-en-105":81,"doc-detail-148435-en":131},{"code":4,"msg":5,"data":6},0,"ok",{"site_id":7,"language":8,"slug":9,"title":10,"keywords":11,"description":12,"schema_data":13,"social_meta":74,"head_meta":76,"extra_data":78,"updated_unix":80},105,"en","web-agent-reasoning-enhancement-webcot-targeted-reasoning-skills-improvements","Web agent reasoning enhancement - WEBCOT - targeted reasoning skills improvements","","Web agents powered by large language models can struggle in uncertain, dynamic web settings due to limited reasoning. This paper defines core web-agent reasoning abilities—reflection & lookahead, branching, and rollback—and builds trajectory data by reconstructing the agent’s inference-time reasoning into chain-of-thought rationales. Experiments on the OpenWebVoyager self-improving benchmark show that simple fine-tuning can distill these reasoning patterns into the backbone LLM, improving performance on WebVoyager, Mind2web-live, and SimpleQA.",{"@graph":14,"@context":73},[15,34,56],{"@type":16,"itemListElement":17},"BreadcrumbList",[18,23,27,31],{"item":19,"name":20,"@type":21,"position":22},"https://docshare.wps.com","Home","ListItem",1,{"item":24,"name":25,"@type":21,"position":26},"https://docshare.wps.com/document/","Document",2,{"item":28,"name":29,"@type":21,"position":30},"https://docshare.wps.com/document/research-report/","Research & Report",3,{"item":32,"name":10,"@type":21,"position":33},"https://docshare.wps.com/document/web-agent-reasoning-enhancement-webcot-targeted-reasoning-skills-improvements/148435/",4,{"url":32,"name":10,"@type":35,"image":36,"author":41,"headline":10,"publisher":44,"fileFormat":47,"inLanguage":8,"description":12,"dateModified":48,"datePublished":49,"encodingFormat":47,"isAccessibleForFree":50,"interactionStatistic":51},"DigitalDocument",{"url":37,"@type":38,"width":39,"height":40},"https://docshare.wps.com/thumbnails/web-agent-reasoning-enhancement-webcot-targeted-reasoning-skills-improvements/148435.png","ImageObject",300,407,{"name":42,"@type":43},"Liam","Person",{"url":19,"name":45,"@type":46},"DocShare","Organization","application/pdf","2026-10-07","2026-08-26",true,{"@type":52,"interactionType":53,"userInteractionCount":55},"InteractionCounter",{"@type":54},"ViewAction",14,{"@type":57,"mainEntity":58},"FAQPage",[59,65,69],{"name":60,"@type":61,"acceptedAnswer":62},"What problem does WEBCOT address in web agents?","Question",{"text":63,"@type":64},"WEBCOT targets the limited reasoning ability of current language models when operating in uncertain, dynamic web environments, which reduces robustness in real web workflows.","Answer",{"name":66,"@type":61,"acceptedAnswer":67},"Which three reasoning skills are central to the paper’s approach?",{"text":68,"@type":64},"The paper focuses on reflection & lookahead, branching, and rollback, each implemented by corresponding representative modules and used to guide trajectory generation.",{"name":70,"@type":61,"acceptedAnswer":71},"How does WEBCOT improve web-agent performance and where is it evaluated?",{"text":72,"@type":64},"It reconstructs inference-time reasoning into chain-of-thought rationales, then distills salient reasoning patterns into the backbone LLM via simple fine-tuning. Results on OpenWebVoyager and additional benchmarks show consistent gains across WebVoyager, Mind2web-live, and SimpleQA.","https://schema.org",{"og:url":32,"og:type":75,"og:title":10,"og:site_name":45,"og:description":12},"article",{"robots":77,"canonical":32},"index,follow",{"doc_id":79,"site_id":7},148435,1787779780,{"code":4,"msg":82,"data":83},"success",[84,88,92,96,101,106,111,115,120,123,127],{"id":22,"doc_module":4,"doc_module_name":25,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":26,"doc_module":4,"doc_module_name":25,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":33,"doc_module":4,"doc_module_name":25,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":97,"doc_module":4,"doc_module_name":25,"category_name":98,"show_sort_weight":99,"slug":100},5,"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":25,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":25,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":112,"doc_module":4,"doc_module_name":25,"category_name":29,"show_sort_weight":113,"slug":114},8,30,"research-report",{"id":116,"doc_module":4,"doc_module_name":25,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":25,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":25,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":25,"category_name":129,"show_sort_weight":97,"slug":130},19,"General","general",{"code":4,"msg":82,"data":132},{"doc_id":79,"user_id":133,"nickname":42,"user_avatar":134,"doc_module":4,"category_id":112,"category_name":29,"doc_title":10,"doc_description":12,"doc_content":135,"file_id":136,"file_url":137,"file_type":138,"file_size":139,"view_count":55,"is_deleted":4,"is_public":22,"is_downloadable":22,"audit_status":22,"page_count":128,"language":140,"language_code":8,"site_id":7,"html_lang":8,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":12,"update_tm":80,"read_time":144},8796095461564,"https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d","WEBCOT: Enhancing Web Agent Reasoning by Reconstructing Chain-of-Thought in Reflection, Branching, and Rollback  \nMinda Hu♣♠ * , Tianqing Fang♠∗ , Jianshu Zhang♥ , Junyu Ma♠ , Zhisong Zhang♠ , Jingyan Zhou♣ , Hongming Zhang♠ , Haitao Mi♠ , Dong Yu♠ , Irwin King♣♣ Chinese University of Hong Kong, ♠Tencent AI Lab, ♥Wuhan University  \n{mindahu21, [king}@cse.cuhk.edu.hk](king}@cse.cuhk.edu.hk), [tianqfang@tencent.com](tianqfang@tencent.com)  \narXiv :2505 .20013v2 [ cs .CL] 18 Sep 2025  \nAbstract  \nWeb agents powered by Large Language Models (LLMs) show promise for next-generation AI, but their limited reasoning in uncertain, dynamic web environments hinders robust deployment. In this paper, we identify key reasoning skills essential for effective web agents, i.e., reflection & lookahead, branching, and rollback, and curate trajectory data that exemplifies these abilities by reconstructing the agent’s (inference-time) reasoning algorithms into chain-of-thought rationales. We conduct experiments in the agent self-improving benchmark, OpenWebVoyager, and demonstrate that distilling salient reasoning patterns into the backbone LLM via simple fine-tuning can substantially enhance its performance. Our approach yields significant improvements across multiple benchmarks, including WebVoyager, Mind2web-live, and SimpleQA (web search), highlighting the potential of targeted reasoning skill enhancement for web agents.  \n1 Introduction  \nThe rise of large language models (LLMs) has sparked significant interest in developing intelligent agents capable of interacting with the web through a browser, commonly referred to as web agents (Yao et al., 2023 ; Monica.Im, 2025 ; Lianget al., 2025) . However, despite recent advancements, even the best-performing web agents still lag far behind human performance—even when compared to users unfamiliar with a website’s structure or functionality (Zhang et al., 2024b ; Mialonet al., 2024 ; Song et al., 2025) . This performance gap is primarily attributed to the limited reasoning abilities of current language models when applied to web agent workflows.  \nDespite recent advances in Large Reasoning Models (LRM, e.g., DeepSeek-R1; DeepSeek-AI  \n*Equal Contribution  \nAgent Inference-time Reasoning Algorithm  \nReconstructed Chain-of-thought  \nFigure 1: Overview of our framework. We leverage a language model to translate inference-time processes, i.e., reflection and look-ahead, branching, and rollback, into natural language chain-of-thoughts, which are then used to train the agent language model.  \net al., 2025, QwQ; Qwen, 2025), these models primarily focus on arithmetic reasoning and are prone to overthinking and generating unnecessarily complex solutions for agent tasks (Cuadron et al., 2025 ; Kumar et al., 2025 ; Su et al., 2025) . While directly applying Reinforcement Learning (RL) in agentic environments (Qi et al., 2025 ; Li et al., 2025 ; Liu et al., 2025 ; Singh et al., 2025 ; Wei et al., 2025) is a viable alternative, the resulting reasoning abilities are often unpredictable and lack structured priors. Moreover, these approaches incur prohibitively high costs (Xu et al., 2025 ; Dang and Ngo, 2025) when conducting real-world rollouts. Additionally, they often focus on static and deterministic environments such as WebArena (Zhou et al., 2024), whereas applying them to stochastic real-world open-domain web environments can be problematic due to the randomness inherent in  \nrollouts. In contrast, distilling specific reasoning patterns into agents (Chen et al., 2024 ; Zhao et al., 2024 ; Hu et al., 2025) combines the adaptability of learned policies with the interpretability and taskaware heuristics of curated reasoning, mitigating both the overthinking problem and the exploration burden of pure RL.  \nIn this paper, we carefully examine and design the specific reasoning abilities required for effective web agents, and sample corresponding agent trajectories to conduct Supervised Fine-Tuning (SFT) on LLMs. In ","cbCaii6MsOm3FKu7","https://ap.wps.com/l/cbCaii6MsOm3FKu7","pdf",781678,"English","# Introduction\n## Motivation and limitations of current web agents\n## Large Reasoning Models vs. distillation approaches\n## Proposed reasoning abilities: reflection & lookahead, branching, rollback\n## Method overview and trajectory construction\n## Experimental setup and evaluation benchmarks","[{\"question\":\"What problem does WEBCOT address in web agents?\",\"answer\":\"WEBCOT targets the limited reasoning ability of current language models when operating in uncertain, dynamic web environments, which reduces robustness in real web workflows.\"},{\"question\":\"Which three reasoning skills are central to the paper’s approach?\",\"answer\":\"The paper focuses on reflection \\u0026 lookahead, branching, and rollback, each implemented by corresponding representative modules and used to guide trajectory generation.\"},{\"question\":\"How does WEBCOT improve web-agent performance and where is it evaluated?\",\"answer\":\"It reconstructs inference-time reasoning into chain-of-thought rationales, then distills salient reasoning patterns into the backbone LLM via simple fine-tuning. Results on OpenWebVoyager and additional benchmarks show consistent gains across WebVoyager, Mind2web-live, and SimpleQA.\"}]","Web agent reasoning enhancement - WEBCOT - targeted reasoning skills improvements | PDF",48]