[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83472-en":3,"doc-seo-83472-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83472,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","AI Sees What You Don’t: Exploiting New Attack Surfaces in Third-Party Mobile Agents","Third-party mobile agents powered by vision-language models act as high-privilege decision makers, using screenshot-based perception and VLM reasoning to control other apps and the OS. This shift creates new exploitable attack surfaces that can turn otherwise benign interfaces into practical entry points. The work contrasts agents with general apps, analyzes their security posture, and identifies two agent-specific surfaces: Screen Perception Attacks and Misused Channel Attacks, supported by seven concrete attacks and evaluations across five popular frameworks.","(A)I Sees What You Don’t: Exploiting New Attack Surfaces in Third-Party Mobile Agents  \nZidong Zhang  \nSimon Fraser University [zza323@sfu.ca](zza323@sfu.ca)  \nZhentao Xie  \nShandong University [fxizenta@mail.sdu.edu.cn](fxizenta@mail.sdu.edu.cn)  \nWenrui Diao  \nShandong University [diaowenrui@link.cuhk.edu.hk](diaowenrui@link.cuhk.edu.hk)  \nJianliang Wu  \nSimon Fraser University [wujl@sfu.ca](wujl@sfu.ca)  \narXiv :2607 .00333v 1 [ cs .CR] 1 Jul 2026  \nAbstract—Third-party mobile agents powered by VisionLanguage Models (VLMs) have emerged as a promising paradigm for automating smartphone interactions. These agents act as high-privilege decision-makers, perceiving device states through screenshots and executing actions via VLM reasoning, transforming how an agent app interacts with the environment (i.e., other apps or the OS). Correspondingly, this transformation introduces new attack surfaces or transforms benign/harmless interfaces into exploitable ones for mobile devices.  \nIn this paper, we summarize key differences between third-party mobile agent apps and general apps when interacting with the environment, analyze the security posture of agents, and identify two unique attack surfaces compared to general mobile apps: the Screen Perception Attack Surface, which exploits the gap between human and machine vision, and the Misused Channel Attack Surface, which intercepts or manipulates the agent’s execution pipeline. We design and implement seven concrete attacks, from subliminal text injection and invisible pixel zone exploitation to screenshot tampering and host PC command injection. Our evaluation of five popular mobile agent frameworks demonstrates that a malicious app can hijack agent actions and achieve arbitrary command execution even without any privilege permissions, while remaining visually indistinguishable to users. These findings reveal a fundamental trust mismatch in autonomous agent design and highlight the urgent need for perception-aware security modelson multi-tenant platforms.  \nI. INTRODUCTION  \nMobile agents powered by large language models (LLMs) and vision-language models (VLMs) are rapidly transforming how users interact with their smartphones. Intelligent assistants (such as AppAgent [47], Mobile-Agent [46], and Open-AutoGLM [35]) can now understand natural-language instructions, perceive screen content through visual analysis, and autonomously execute complex multi-step tasks ranging from messaging and shopping to financial transactions. Unlike traditional automation scripts [27] that require explicit programming for each scenario, modern mobile agents leverage the reasoning capabilities of foundation models to generalize across diverse applications and contexts, promising a future where smartphones truly become intelligent personal assistants. However, this paradigm differs from conventional mobile apps as it introduces new sources of input and communication channels. On the one hand, while conventional apps barely take screen perception as their input, third-party agents heavily rely on these inputs to interpret UI states. On the other hand, certain communication channels (e.g., ADB [1]), which are rarely used (if not unused at all) by conventional apps, are  \nnecessary for the agents. Accordingly, this paradigm opens new exploitable attack surfaces that are either harmless or nonexistent for conventional apps.  \nIn this paper, we analyze the security posture of thirdparty mobile agents and identify two unique attack surfaces inherent to their design: (1) the Screen Perception Attack Surface, rooted in the new inputs of screen perception, and (2) the Misused Channel Attack Surface, stemming from the misuse of debug interfaces and unauthenticated system broadcasts. We uncover fundamental vulnerabilities where agents implicitly trust visual artifacts that humans cannot see, and communication channels that lack authentication, creating opportunities for stealthy exploitation.  \nOur investigation reveal","cbCaiee95zbvlj6H","https://ap.wps.com/l/cbCaiee95zbvlj6H","pdf",1815452,1,16,"English","en",105,"# Introduction\n## Threat model and attack surfaces\n## Screen Perception Attacks\n## Misused Channel Attacks\n## Evaluation across mobile agent frameworks\n## Countermeasures and contributions","[{\"question\":\"为什么第三方移动智能体会带来新的攻击面？\",\"answer\":\"移动智能体依赖截图感知与VLM推理来理解界面状态，并通过额外通信/调试通道与系统交互，这会把原本对传统App较少使用或未被利用的接口与输入转换为可被滥用的入口。\"},{\"question\":\"文中提出的两种独特攻击面是什么？\",\"answer\":\"两种攻击面分别是Screen Perception Attack Surface（利用人类与机器视觉差距）和Misused Channel Attack Surface（拦截或操纵智能体的执行流水线）。\"},{\"question\":\"实验结果表明攻击者可以实现什么效果？\",\"answer\":\"在对五个开源移动智能体框架的评估中，恶意应用可劫持智能体动作，并在缺少权限的情况下实现任意命令执行，同时对用户保持视觉不可区分。\"}]",1784188212,40,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"ai-sees-what-you-dont-exploiting-new-attack-surfaces-in-third-party-mobile-agents","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/ai-sees-what-you-dont-exploiting-new-attack-surfaces-in-third-party-mobile-agents/83472/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么第三方移动智能体会带来新的攻击面？","Question",{"text":75,"@type":76},"移动智能体依赖截图感知与VLM推理来理解界面状态，并通过额外通信/调试通道与系统交互，这会把原本对传统App较少使用或未被利用的接口与输入转换为可被滥用的入口。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"文中提出的两种独特攻击面是什么？",{"text":80,"@type":76},"两种攻击面分别是Screen Perception Attack Surface（利用人类与机器视觉差距）和Misused Channel Attack Surface（拦截或操纵智能体的执行流水线）。",{"name":82,"@type":73,"acceptedAnswer":83},"实验结果表明攻击者可以实现什么效果？",{"text":84,"@type":76},"在对五个开源移动智能体框架的评估中，恶意应用可劫持智能体动作，并在缺少权限的情况下实现任意命令执行，同时对用户保持视觉不可区分。","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":28,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]