[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85339-en":3,"doc-seo-85339-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85339,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos","Continuous egocentric video enables assistance that is proactive rather than merely reactive, but prior systems either wait for user queries or respond to detected events without judging whether interruption is appropriate. The work frames proactive help as a context-dependent decision problem requiring temporal reasoning over accumulated history. It introduces Vinci2, a proactive on-device assistant advancing Vinci, and EgoServe, the first large-scale benchmark covering multiple service categories and memory horizons. It also proposes EgoMemo, a training-free memory-augmented agent for retrieval-augmented, context-grounded responses.","arXiv :2607 . 11523v1 [ cs .CV] 13 Jul 2026  \nVinci2: Providing Proactive Assistance in Continuous Egocentric Videos  \nSitong Gong 1 ,2 ,3, Tianyu Yan⋆1 ,2 ,3, Caixin Kang⋆2 ,3, Bo Zheng2 , Xiang Ruan 1, Huchuan Lu 1, Kaipeng Zhang2, Yoichi Sato3, and  \nYifei Huang†2 ,3  \n1 Dalian University of Technology 2 Alaya Lab 3 The University of Tokyo [stgong@mail.dlut.edu.cn](stgong@mail.dlut.edu.cn) , {tianyu.yan, caixin.kang, bo.zheng, [kaipeng.zhang}@shanda.com](kaipeng.zhang}@shanda.com) , {ruanxiang, [lhchuan}@dlut.edu.cn](lhchuan}@dlut.edu.cn) ,  \n[ysato@iis.u-tokyo.ac.jp](ysato@iis.u-tokyo.ac.jp) , [hyf015@gmail.com](hyf015@gmail.com)  \nAbstract. When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a new form of assistance: one that is proactive rather than merely reactive. Yet existing approaches either wait passively for user queries or treat every detected event as requiring a response, without considering the user’s history, current activity, or whether assistance would actually be welcome. We reframe proactive assistance as a context-dependent decision problem: the agent must not only perceive what is happening, but reason over accumulated temporal context to determine when and whether to intervene. To this end, we present Vinci2, a proactive egocentric assistance system that advances the on-device assistant Vinci from reactive response toward proactivity. On the evaluation side, we present EgoServe, the first large-scale benchmark for proactive assistance in continuous egocentric video. EgoServe comprises over 3,000 service instances organized along 4 temporal memory horizons, ranging from immediate safety alerts to long-term habit coaching, across 10 service categories. On the modeling side, we propose EgoMemo, a training-free, memory-augmented agent that maintains three complementary memory representations: multi-scale temporal summaries, a semantic knowledge graph, and visual embedding archives. At each timestep, EgoMemo performs retrieval-augmented reasoning to determine whether assistance is warranted and, if so, produces contextually grounded responses. Experiments demonstrate that EgoMemo establishes strong baselines on EgoServe while remaining competitive on existing egocentric benchmarks. Our benchmark and code are publicly available at Vinci2 .  \nKeywords: Egocentric Vision · Proactive VLM · LLM Agent  \n⋆ Equal contribution.† Corresponding author.  \n2 Sitong Gong et al.  \n(a) Reactive Paradigm (b) Semi-proactive Paradigm  \nReaching that high may shake the phone. Please use the side volume button for a steadier shot.  \nYou left a file transfer running on your laptop downstairs a few minutes ago. Since you're free now, do you want to check its status?  \nYou’re looking for salt. You put it on the kitchen table after buying it the day before yesterday.  \nMemory Construction  \nMultimodal Context Knowledge Graph  \nEpisodic Messages  \n(c) Proactive Paradigm (Ours)  \nEgocentric Video  \nStream   \nFig. 1: Three paradigms of egocentric assistants. (a) Reactive: responds only to explicit user queries. (b) Semi-proactive: monitors the video stream for predefined taskrelevant events. (c) Proactive (ours): autonomously decides when and how to intervene without any user prompt.  \n1 Introduction  \nThe promise of proactive egocentric intelligence is an assistant that sees what you see, understands your context as it evolves, and offers help at the right moment without being asked. Recent progress in Video-LLMs [30, 33 , 51 , 53 , 54], streaming visual perception [27, 46 , 70], and egocentric foundation models [22, 64 , 69] has brought this vision within reach, enabling continuous comprehension and reasoning over first-person video. Yet while the perception capabilities are maturing rapidly, the question of how and when an assistant should proactively intervene remains largely unaddressed.  \nExisting egocentric assistants operate under two limitin","cbCaijZC5VJdkjW9","https://ap.wps.com/l/cbCaijZC5VJdkjW9","pdf",7790032,1,36,"English","en",105,"# Abstract\n# Keywords\n# Introduction\n## Motivation and problem framing\n## Limitations of existing paradigms\n## Proposed approach: Vinci2\n## Evaluation gap and EgoServe\n## Modeling approach: EgoMemo","[{\"question\":\"What problem does Vinci2 address in continuous egocentric videos?\",\"answer\":\"It tackles when an intelligent assistant should speak up without being asked, deciding whether and how to intervene based on evolving context and user history.\"},{\"question\":\"How does EgoServe support evaluation of proactive assistance?\",\"answer\":\"EgoServe is a large-scale benchmark with service instances organized across multiple temporal memory horizons and service categories, including needs ranging from safety alerts to habit coaching.\"},{\"question\":\"What is EgoMemo and how does it determine whether assistance is warranted?\",\"answer\":\"EgoMemo is a training-free, memory-augmented agent that maintains multi-scale temporal summaries, a semantic knowledge graph, and visual embedding archives, then performs retrieval-augmented reasoning at each timestep to decide and generate grounded responses.\"}]",1784202610,91,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"vinci2-providing-proactive-assistance-in-continuous-egocentric-videos","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/vinci2-providing-proactive-assistance-in-continuous-egocentric-videos/85339/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Vinci2 address in continuous egocentric videos?","Question",{"text":75,"@type":76},"It tackles when an intelligent assistant should speak up without being asked, deciding whether and how to intervene based on evolving context and user history.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does EgoServe support evaluation of proactive assistance?",{"text":80,"@type":76},"EgoServe is a large-scale benchmark with service instances organized across multiple temporal memory horizons and service categories, including needs ranging from safety alerts to habit coaching.",{"name":82,"@type":73,"acceptedAnswer":83},"What is EgoMemo and how does it determine whether assistance is warranted?",{"text":84,"@type":76},"EgoMemo is a training-free, memory-augmented agent that maintains multi-scale temporal summaries, a semantic knowledge graph, and visual embedding archives, then performs retrieval-augmented reasoning at each timestep to decide and generate grounded responses.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]