[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84723-en":3,"doc-seo-84723-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84723,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure","Multimodal GUI agents interpret an interface through two redundant channels: screenshot pixels and a serialized representation such as a DOM or accessibility tree. Benchmarks evaluate task success and robustness but rarely test whether the agent’s state belief is actually sourced from pixels. The study formalizes visual state reliance, attributes state beliefs to pixels, structure, or priors, and measures it via paired single-channel interventions on 310 real GUI probes using deterministic scoring. A key metric, Perception–Fusion Gap, reveals that textual state beliefs defer to structure under conflict while image-only accuracy remains near ceiling.","Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels  \nversus Structure  \nGuijia Zhang 1 Harry Yang2  \n1 Shenzhen University  \n2The Hong Kong University of Science and Technology  \narXiv :2607 .04334v 1 [ cs .AI ] 5 Jul 2026  \nAbstract  \nMultimodal GUI agents read an interface through two redundant channels: the rendered pixels of a screenshot and a serialized structure such as a DOM or accessibility tree. Before acting, an agent forms a belief about the current interface state, but existing benchmarks score task success, element grounding, or attack resistance and do not ask whether that belief is drawn from the pixels. We formalize visual state reliance, the attribution of a state belief to pixels, structure, or priors, and measure it with paired single-channel interventions over 310 real web, mobile, and desktop probes. Every probe is scored by deterministic forced choice, with no modelgenerated item and no model judge. Our central metric is the Perception-Fusion Gap pfg, the fraction of probes a model perceives correctly yet resolves toward structure under conflict. Across five models from three vendors, textual state beliefs defer to structure while image-only accuracy stays near ceiling, and pfg is positive for every model; non-text identity, by contrast, stays largely pixel-bound. The substitution is specific to the serialized-text and indexed-action channel, and coordinate-action agents are largely immune. For textual conflicts a white-box ablation traces the effect to a single copied structural value, and in two live environments the conflict drives wrong actions and real task failure. Visual state reliance therefore gives a measurable diagnostic of whether agent state beliefs are visually grounded, and the errors it exposes propagate to actions.  \nIntroduction  \nMultimodal GUI agents now operate computers, browsers, and phones by consuming two redundant views of the same interface: the rendered pixels of a screenshot and a serialized structure such as a DOM or accessibility tree (Deng et al. 2023; Koh et al. 2024) . Before an agent can plan or act, it must form a belief about the current state of the interface: whether a toggle is on, which tab is selected, or what text afield contains. This belief precedes every action, and a wrong belief propagates into the decisions that follow even when the planner is otherwise competent. Existing benchmarks score the end of that loop, namely task completion, element grounding, or robustness to injected attacks (Koh et al. 2024; Cheng et al. 2024; Cao et al. 2025), and so leave the origin of the belief unmeasured.  \nThe two channels are usually consistent, yet they diverge often enough to matter: accessibility trees go stale, con-  \npfg: perceive correctly yet defer  \nFigure 1: The core scenario, on a real probe from Multimodal-Mind2Web (Deng et al. 2023) . The same interface state reaches the agent through two channels: the rendered pixels P, a screenshot whose highlighted control reads Reservations, and the serialized structure S, a DOM node we edit to read Budget Truck. Read from the pixels alone the agent answers Reservations correctly, yet its fused belief follows the edited structure. We report the fraction of such perceive-correctly-yet-defer probes as the Perception– Fusion Gap pfg, our central metric.  \ntent renders after the structure is captured, and localized strings disagree with displayed labels. Such divergence is documented at scale: human CLAY (Li et al. 2022) annotations mark 10.6% of RICO structure nodes as having no valid visual representation, and such a ghost node appears on 37.4% of screens. A model that reports the correct state by copying a structural string it never visually verified matches  \nthe screen only as long as the structure happens to agree, and it fails without warning once structure and reality diverge. This divergence motivates the question but does not by itself reveal which channel a belief follows, and provenance i","cbCaiutqLhYeohut","https://ap.wps.com/l/cbCaiutqLhYeohut","pdf",1034619,1,15,"English","en",105,"# Abstract\n# Introduction\n# Findings","[{\"question\":\"What two channels do multimodal GUI agents use to understand an interface state?\",\"answer\":\"They consume rendered screenshot pixels and a serialized structure such as a DOM or accessibility tree. The agent forms its interface-state belief before planning or acting based on these channels.\"},{\"question\":\"Why do existing GUI-agent benchmarks not fully answer whether beliefs come from pixels?\",\"answer\":\"They typically score end results like task completion, element grounding, or robustness to attacks, without isolating where the belief underlying decisions originates. As a result, belief provenance is left unmeasured.\"},{\"question\":\"What is the Perception–Fusion Gap and what does it diagnose?\",\"answer\":\"Perception–Fusion Gap is the fraction of probes where a model perceives the correct visible state yet resolves toward structure when the channels conflict. It provides a measurable diagnostic of whether state beliefs are visually grounded and how those grounding errors propagate to actions.\"}]",1784197857,38,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"do-gui-agents-believe-their-eyes-diagnosing-state-belief-reliance-on-pixels-versus-structure","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/do-gui-agents-believe-their-eyes-diagnosing-state-belief-reliance-on-pixels-versus-structure/84723/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What two channels do multimodal GUI agents use to understand an interface state?","Question",{"text":75,"@type":76},"They consume rendered screenshot pixels and a serialized structure such as a DOM or accessibility tree. The agent forms its interface-state belief before planning or acting based on these channels.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why do existing GUI-agent benchmarks not fully answer whether beliefs come from pixels?",{"text":80,"@type":76},"They typically score end results like task completion, element grounding, or robustness to attacks, without isolating where the belief underlying decisions originates. As a result, belief provenance is left unmeasured.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the Perception–Fusion Gap and what does it diagnose?",{"text":84,"@type":76},"Perception–Fusion Gap is the fraction of probes where a model perceives the correct visible state yet resolves toward structure when the channels conflict. It provides a measurable diagnostic of whether state beliefs are visually grounded and how those grounding errors propagate to actions.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]