[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86464-en":3,"doc-seo-86464-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86464,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation","Digital Adoption Platforms (DAPs) embed web overlays that help users complete operations by highlighting elements and providing short instructions. Real task completion, however, requires multi-step action across changing page states, while prior work largely separates web agent actions from guide writing and relies on DOM/accessibility text instead of rendered screenshots. MAG unifies execution and stepwise guide generation with two screenshot grounding schemes and provides an end-to-end harness for LLM-assisted annotation, human verification, live training and evaluation, and joint action-guide metrics. Results show \u003C40% completion for top models and guide-quality improvements via GRPO training.","MAG: A Web-Agent Benchmark and Harness for Multimodal Action and  \nGuide Generation  \nChengguang Gan 1 , Hanjun Wei2 , Yunhao Liang2 , Zhixi Cai3 , Qinghao Zhang4 , Shiwen Ni5  \n1Independent Researcher, 2University of Chinese Academy of Sciences, 3Monash University, 4Pusan National University,  \n5 Shenzhen University of Advanced Technology  \n[chengguangg1024@gmail.com](chengguangg1024@gmail.com) , {weihanjun23, [liangyunhao22}@mails.ucas.ac.cn](liangyunhao22}@mails.ucas.ac.cn)  \n[zhixi.cai@monash.edu](zhixi.cai@monash.edu), [zhangqinghao@pusan.ac.kr](zhangqinghao@pusan.ac.kr), [nishiwen@suat-sz.edu.cn](nishiwen@suat-sz.edu.cn)  \narXiv :2607 . 10079v 1 [ cs .AI] 11 Jul 2026  \nAbstract  \nDigital Adoption Platforms (DAPs) are embedded overlays widely used on web systems to guide users through operations inside a page, helping them get started with unfamiliar interfaces quickly. Completing a real task, however, rarely means clicking a few buttons on a single page: it takes a sequence of actions that unfolds across changing page states. Prior studies have also treated automated web agent actions and guide text generation as two separate problems, and most of them feed models textual page representations such as the DOM or accessibility trees rather than the rendered screens that humans actually operate on. In this work we introduce MAG, the first benchmark that unifies task execution and guide writing into a single Multimodal Action and Guide task, with two grounding schemes over screenshots:  \nSet-of-Mark element selection and raw pixel coordinates. We further build a complete harness for this compound task, covering annotation with LLM assistance and human verification, training, evaluation in live environments, and joint metrics for actions and guides. With this harness we evaluate frontier API models and open multimodal models, and report detailed analyses. Finally, we design a GRPO training method augmented with expert trajectories, which nearly doubles the success rate of a supervised 9B agent (from 6.9% to 13.2%) and improves guide quality at the same time. Even the strongest model completes fewer than 40% of the tasks, leaving ample room for future research.  \n1 Introduction  \nCommercial Digital Adoption Platforms (DAPs) overlay guidance on web systems to help new users operate complex, unfamiliar interfaces. When a user faces an unfamiliar page, the platform highlights the element that matters and shows a short instruction next to it (Figure 1(a)) . This convenience rests on manual labor: vendors locate the  \n(a) Digital Adoption Platform (DAP): human written guides overlaid on live web UIs  \ntoday: authored and updated by hand, page by page  \n(b) MAG: one agent acts and writes the guide, step by step  \nrepeated until the agent outputs finish(answer)  \n(c) one run, two artifacts  \nauthored once by the agent, reused by every future user  \nFigure 1: Overview of MAG. (a) A Digital Adoption Platform overlays human written guides on live webpages; today these guides are authored and updated by hand. (b) The MAG task asks a single agent to complete the task and to write a guide at every step, under two grounding schemes: Set-of-Mark elementselection (Step 1) and raw pixel coordinates (Step 2) . (c) One run yields two artifacts: a verifiable task outcome and a guide that future users can reuse as a DAP overlay.  \ntarget element of every step and write its guide text, and every redesign forces both to be redone. Because one goal often spans several pages anda chain of dependent operations, guidance for a whole system is expensive to author and maintain. Web agents are a natural way to remove this labor, but the two relevant research lines have developed separately. Agent benchmarks score task completion alone (Zhou et al., 2024 ; Deng et al.,  \n2023 ; Yao et al., 2022a ; Koh et al., 2024a): none asks the agent to produce, or scores, the instruction a human would need at each step. Guide generation has been studied in the opposite ","cbCaicCuorEHwUFW","https://ap.wps.com/l/cbCaicCuorEHwUFW","pdf",2483892,4,1,21,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does MAG address compared with prior web-agent and guide-generation work?\",\"answer\":\"MAG unifies task execution with step-by-step guide writing, instead of treating automated web actions and instruction generation as separate problems. It also grounds decisions on rendered screenshots rather than only textual page representations like the DOM.\"},{\"question\":\"How does MAG ground the agent’s actions on screenshots?\",\"answer\":\"MAG provides two grounding schemes: Set-of-Mark element selection over numbered elements, and raw pixel coordinates. Each step uses the chosen grounding to connect agent actions with the corresponding guide text.\"},{\"question\":\"What does the MAG harness include for building and evaluating the benchmark?\",\"answer\":\"The harness covers LLM-assisted annotation paired with human correction and verification, supervised and reinforcement training pipelines, live evaluation using functional checkers and an LLM judge, and joint metrics for action success and guide quality.\"}]",1784211882,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mag-a-web-agent-benchmark-and-harness-for-multimodal-action-and-guide-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/mag-a-web-agent-benchmark-and-harness-for-multimodal-action-and-guide-generation/86464/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does MAG address compared with prior web-agent and guide-generation work?","Question",{"text":75,"@type":76},"MAG unifies task execution with step-by-step guide writing, instead of treating automated web actions and instruction generation as separate problems. It also grounds decisions on rendered screenshots rather than only textual page representations like the DOM.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does MAG ground the agent’s actions on screenshots?",{"text":80,"@type":76},"MAG provides two grounding schemes: Set-of-Mark element selection over numbered elements, and raw pixel coordinates. Each step uses the chosen grounding to connect agent actions with the corresponding guide text.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the MAG harness include for building and evaluating the benchmark?",{"text":84,"@type":76},"The harness covers LLM-assisted annotation paired with human correction and verification, supervised and reinforcement training pipelines, live evaluation using functional checkers and an LLM judge, and joint metrics for action success and guide quality.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]