[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84097-en":3,"doc-seo-84097-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84097,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","UI2App Benchmarking Visual Interaction Inference in Executable Web Application Generation","Large language models can generate webpages, yet text-driven methods depend on complex prompts that limit control over layout and cross-page visual coherence. Image-driven approaches using UI screenshots better match development workflows, but existing benchmarks emphasize visual fidelity and do not systematically evaluate whether generated artifacts support real interaction behavior. UI2App is introduced as the first benchmark for interaction inference from screenshots alone. It contains 327 screenshots in 45 state-coherent runnable multi-route web application sets and evaluates executability, navigation reachability, visual fidelity, and interaction inference.","arXiv :2607 .06306v 1 [ cs . SE] 7 Jul 2026  \nUI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation  \nGrace Man Chen 1 , Litao Guo 1 , Yifan Wu 1 , Yiyu Chen 1 , Yenchi Tseng 1 , Sicheng Liu 1 ,  \nYuyu Luo 1 ,2 ,†, Ying-Cong Chen 1 ,2 ,†  \n1The Hong Kong University of Science and Technology (Guangzhou)  \n2The Hong Kong University of Science and Technology  \n†Corresponding author  \n􀂀 Project § Code  Dataset  \nAbstract  \nLarge language models (LLMs) have demonstrated growing competence in webpage generation. However, existing text-driven approaches rely on complex prompts that impose substantial demands on users and offer limited expressivity for page layout and cross-page visual coherence. Image-driven paradigms, which take UI screenshots as input, align more closely with real development workflows. However, current benchmarks focus primarily on visual fidelity and lack a systematic evaluation of the interaction capabilities in generated artifacts. To address this gap, we introduce UI2App, the first benchmark targeting interaction inference, the ability to recover application behavior from screenshots alone, without any textual or behavioral guidance. UI2App comprises 327 screenshots grouped into 45 state-coherent screenshot sets for runnable multi-route web applications.  \nWe design an end-to-end pipeline that evaluates each artifact along four dimensions: executability, navigation reachability, visual fidelity, and interaction inference. The interaction metric (IIS) assesses inferred interactions by functional correctness and state-management complexity, crediting any valid implementation rather than matching a single reference. Experiments on six frontier vision-language models reveal a marked capability mismatch between visual reconstruction and interaction realization: the visual-fidelity leader scores only 7.5 on IIS, ranking fourth and trailing the IIS leader by 5.2 × . High-complexity interactions such as cross-page state remain a pervasive bottleneck, with half of the evaluated models scoring exactly zero on this dimension. Overall, the results indicate that inferring complete interaction behavior from static screenshots remains a key challenge for models.  \n1 Introduction  \nLarge language models (LLMs) have made notable progress in code generation, task planning, and reasoning [Roziere et al., 2023, Guo et al., 2024, 2025] . Building on these advances, recent work has demonstrated the feasibility of end-to-end software development from natural-language instructions [Jimenez et al., 2024, Hong et al., 2024, Qian et al., 2024] . However, existing text-driven paradigms largely rely on carefully designed or highly detailed prompts as input [Lu et al., 2025, Zhang et al., 2025a], which introduces practical limitations. First, textual descriptions are limited in precisely specifying detailed visual requirements, making it difficult to accurately define page layoutsand maintain visual coherence across multiple pages. Second, functional requirements involving cross-page interactions, such as state management and data synchronisation, are difficult to specify precisely and consistently in natural language.  \nTo alleviate these limitations, image-to-webpage generation has emerged as an alternative paradigm. Instead of relying on textual specifications, this approach takes UI designs as input and reconstructs webpage structure and appearance directly from visual signals. This paradigm better aligns with  \nPreprint.  \nInput: static UI screenshots alone (no behavior spec) → generate a runnable app  \nBuild & Render (EXEC) Navigation (NRS) Visual fidelity (VFS) Interaction inference-IIS (ours)  \n| Internal state\u003Cbr>(simplified) |  |\n| --- | --- |\n\nBuild & Render (EXEC)  Navigation (NRS)  Visual fidelity (VFS)  Interaction inference-IIS (ours)    \nClick  \n'Add to Cart'(size M)  \ncart = [] // no items, no schema  \n// CartContext.add() not implemented  \nFigure 1: UI2App task overview: visual fidelit","cbCaic3C9WRi6svm","https://ap.wps.com/l/cbCaic3C9WRi6svm","pdf",13199754,3,1,35,"English","en",105,"# Abstract\n# Introduction\n## Limitations of text-driven webpage generation\n## Image-to-webpage generation and visual benchmarks\n## Interaction inference as a missing evaluation dimension","[{\"question\":\"What problem does UI2App address in executable web application generation?\",\"answer\":\"UI2App addresses the lack of benchmarks that evaluate interaction inference from UI screenshots alone, instead of focusing only on visual fidelity.\"},{\"question\":\"How does UI2App evaluate generated artifacts?\",\"answer\":\"It evaluates each artifact along four dimensions: executability, navigation reachability, visual fidelity, and interaction inference using the IIS metric.\"},{\"question\":\"What do the experiments on vision-language models reveal?\",\"answer\":\"They show a capability mismatch: models with high visual fidelity can still perform poorly on interaction realization, and high-complexity cross-page state interactions often yield zero scores for half of the models.\"}]",1784192769,88,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"ui2app-benchmarking-visual-interaction-inference-in-executable-web-application-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/ui2app-benchmarking-visual-interaction-inference-in-executable-web-application-generation/84097/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does UI2App address in executable web application generation?","Question",{"text":75,"@type":76},"UI2App addresses the lack of benchmarks that evaluate interaction inference from UI screenshots alone, instead of focusing only on visual fidelity.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does UI2App evaluate generated artifacts?",{"text":80,"@type":76},"It evaluates each artifact along four dimensions: executability, navigation reachability, visual fidelity, and interaction inference using the IIS metric.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the experiments on vision-language models reveal?",{"text":84,"@type":76},"They show a capability mismatch: models with high visual fidelity can still perform poorly on interaction realization, and high-complexity cross-page state interactions often yield zero scores for half of the models.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]