[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85967-en":3,"doc-seo-85967-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85967,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Unibrowse A Data-to-Agent Framework for Multimodal BrowseComp","Multimodal BrowseComp agents must integrate perception, tool use, and long-horizon reasoning over evolving web content, coping with compositional structure, open-world uncertainty, and cross-modal grounding. Real-world browsing includes three information-flow patterns: text-only, image-to-text, and text-to-image, yet prior pipelines cover only the first two, leaving text-to-image insufficiently trained. UNIBROWSE delivers a unified data pipeline spanning all patterns, adds live web retrieval to improve fidelity, and uses an exploration-degree metric to filter low-signal instances for efficient RL. A 35B agent trained with SFT and exploration-aware RL reaches state-of-the-art 54.4 average accuracy.","UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp  \nXiyu Wei 1 ,2 * , Qingwei Zong 1 ,3 * ,Zhuocheng Yu 1 ,3 * , Sujian Li 1 ,3 †  \n1 Key Laboratory of Computational Linguistics, MOE, Peking University  \n2 School of Software and Microelectronics, Peking University  \n3 School of Computer Science, Peking University  \n{wxylemon, [shiinasama}@stu.pku.edu.cn](shiinasama}@stu.pku.edu.cn) [lisujian@pku.edu.cn](lisujian@pku.edu.cn)  \narXiv :2607 . 10557v 1 [ cs .CL] 12 Jul 2026  \nAbstract  \nMultimodal BrowseComp tasks require agents to combine perception, tool use, and longhorizon reasoning over dynamic web content, challenging their ability to handle compositional structure, open-world uncertainty, and multimodal integration across extended interactions. Crucially, real-world multimodal browsing involves three distinct information-flow patterns: text-only, image-to-text, and textto-image, yet existing data construction methods cover only the text-only and image-to-text patterns, leaving text-to-image largely unaddressed and limiting agent generality and robustness. We introduce UNIBROWSE, a unified data pipeline that for the first time simultaneously generates training data covering all three patterns, augments curated knowledge graphs with live web retrieval for improved fidelity, and introduces a novel metric of exploration degree to filter low-signal instances for efficient reinforcement learning. Through this pipeline, we produce high-quality cold-start tool-use trajectories and explorationrich QA pairs, and train a 35B-scale agent via supervised fine-tuning and explorationaware RL. The resulting UNIBROWSE agent achieves state-of-the-art performance on multimodal BrowseComp benchmarks, attaining an average accuracy of 54.4 across five diverse benchmarks—an improvement of 10.5 points over its base model Qwen3.5-35B-A3B—and surpassing serveral closed-source agent workflows such as GPT-5 (42.9), Gemini-2.5 Pro (44 . 8), and Gemini-2 .5 Flash (41 .3) .  \n1 Introduction  \nBrowseComp-style tasks evaluate whether web browsing agents can solve questions whose answers are difficult to locate and must be composed from multiple pieces of web evidence (Wei et al., 2025) . Unlike standard fact-seeking tasks (Nakano  \n*Equal contribution.  \n†Corresponding authors.  \net al., 2021), these problems cannot be solved by issuing a single query or extracting a single webpage snippet. Instead, an agent must plan searches, follow partial clues, assess evidence reliability, and combine information across sources to arrive at a uniquely determined answer. Real-world browsing further extends this challenge beyond text: users often provide image clues, ask agents to identify visually grounded entities, or require visual verification after textual search has narrowed down the target. This gives rise to the more challenging setting of multimodal BrowseComp, where agents must jointly perform textual retrieval, visual grounding, tool use, and long-horizon reasoning over openweb evidence. This emerging setting has therefore attracted increasing attention from the research  \ncommunity (Jiang et al., 2024 ; Li et al., 2025b) .  \nDespite this growing interest, constructing effective training data for multimodal BrowseComp remains difficult. Early steps toward injecting visual information into the data generation process include WebWatcher (Geng et al., 2025), which relies on curated knowledge graphs to obtain structured relations, as well as Vision-DeepResearch (Zenget al., 2026) and Skywork R1V4 (Zhang et al., 2025), which employ web random walks to mimic authentic browsing environments. Although these pioneering methods succeed in producing largescale multimodal training instances, their further application is hindered by a lack of task diversity and insufficient training efficiency, which manifests in three key aspects. First, the existing dataset exclusively covers the image-to-text pattern where an initial image clue is resolved into textual evid","cbCaid8sSHrAX1At","https://ap.wps.com/l/cbCaid8sSHrAX1At","pdf",9169098,4,1,17,"English","en",105,"# Introduction\n## Multimodal BrowseComp challenges\n## Limitations of existing dataset construction\n## UNIBROWSE contributions and results","[{\"question\":\"What makes multimodal BrowseComp tasks harder than standard web question answering?\",\"answer\":\"They require composing answers from multiple web evidence pieces, while planning searches, assessing evidence reliability, and combining sources. Multimodal settings further add visual grounding and visual verification after textual retrieval.\"},{\"question\":\"Which information-flow patterns does UNIBROWSE address that earlier methods miss?\",\"answer\":\"UNIBROWSE covers text-only, image-to-text, and text-to-image patterns together. Earlier data construction methods mainly cover text-only and image-to-text, leaving text-to-image under-addressed.\"},{\"question\":\"How does UNIBROWSE improve data quality and training efficiency?\",\"answer\":\"It unifies a data pipeline that generates training data for all patterns and augments curated knowledge graphs with live web retrieval for higher fidelity. It also introduces an exploration degree metric to filter low-signal instances, enabling more efficient reinforcement learning.\"}]",1784207452,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"unibrowse-a-data-to-agent-framework-for-multimodal-browsecomp","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/unibrowse-a-data-to-agent-framework-for-multimodal-browsecomp/85967/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What makes multimodal BrowseComp tasks harder than standard web question answering?","Question",{"text":75,"@type":76},"They require composing answers from multiple web evidence pieces, while planning searches, assessing evidence reliability, and combining sources. Multimodal settings further add visual grounding and visual verification after textual retrieval.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which information-flow patterns does UNIBROWSE address that earlier methods miss?",{"text":80,"@type":76},"UNIBROWSE covers text-only, image-to-text, and text-to-image patterns together. Earlier data construction methods mainly cover text-only and image-to-text, leaving text-to-image under-addressed.",{"name":82,"@type":73,"acceptedAnswer":83},"How does UNIBROWSE improve data quality and training efficiency?",{"text":84,"@type":76},"It unifies a data pipeline that generates training data for all patterns and augments curated knowledge graphs with live web retrieval for higher fidelity. It also introduces an exploration degree metric to filter low-signal instances, enabling more efficient reinforcement learning.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]