[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81589-en":3,"doc-seo-81589-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81589,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Act, Sense, Act: Learning Active Perception from Large-Scale Egocentric Human Data","Achieving generalizable robotic manipulation in unconstrained settings requires proactive handling of information uncertainty through active perception. Existing methods often restrict sensing to a narrow set of behaviors, limiting performance in complex scenes. Active perception is formalized as a history-dependent perception-action loop driven by information-seeking actions and decision branching, with visual paradigms categorized. A cognitive memory-aware vision-language-action framework, CoMe-VLA, uses large-scale egocentric human data to learn exploration and manipulation priors.","arXiv :2602 .04600v2 [ cs .RO] 10 Jul 2026  \nAct, Sense, Act: Learning Active Perception from Large-Scale Egocentric Human Data  \nJialiang Li*, Yi Qiao*, Yunhan Guo, Changwen Chen, Wenzhao Lian  \n*Equal Contribution  \nSchool of Artificial Intelligence, Shanghai Jiao Tong University  \nFig. 1. By unifying large-scale egocentric human data and robot data within a shared action space, a wheel-based humanoid develops robust and adaptive active perception capabilities through memory-driven temporal and cognitive learning mechanisms. This establishes a tight \"Act, Sense, Act\" coupling between perception and action, which effectively addresses diverse active perception tasks.  \nAbstract. Achieving generalizable manipulation in unconstrained environments requires the robot to proactively resolve information uncertainty, i.e., the capability of active perception. However, existing methods are often confined in limited types of sensing behaviors, restricting their applicability to complex environments. In this work, we formalize active perception as a history-dependent perception-action loop driven by information-seeking action and decision branching, providing a structured categorization of visual active perception paradigms.  \nBuilding on this perspective, we introduce CoMe-VLA, a cognitive and memory-aware vision-language-action (VLA) framework that leverages large-scale human egocentric data to learn versatile exploration and manipulation priors. Our framework integrates a cognitive auxiliary head for autonomous sub-task transitions and a dual-track memory system to maintain consistent self and environmental awareness by fusing proprioceptive and visual temporal contexts. By aligning human and robot hand-eye coordination behaviors in a unified egocentric action space, we train the model progressively in three stages. Extensive experiments on a wheel-based humanoid have demonstrated strong robustness and adaptability of our proposed method across diverse long-horizon tasks spanning multiple active perception scenarios.View our project in [https:](https:)//[jern-li.github.io/asa/](jern-li.github.io/asa/)  \nKeywords: Active Perception · Egocentric Human Data · Humanoid.  \n2 Li et al.  \n1 Introduction  \nThe significant progress in robotic manipulation in recent years has been largely driven by the rapid advancement of imitation learning (IL) [41,7] and foundation models [17,4] . These approaches have demonstrated promising performance on deterministic linear tasks in structured environments, where the observation-toaction mapping along the task execution is stationary. However, to deploy robots in complex unstructured environments, intentional gazing and exploration based on previous interaction history, rather than just reactive movement, are required. For example, to retrieve a wrench buried in a cluttered toolbox, a robot may need to explore different viewpoints or manipulate surrounding objects to reveal occluded regions. The robot then continuously adapts its execution based on newly observed information, such as reaching towards the left if the wrench appears on the left, or towards the right if revealed on the right. Such scenarios require a cognitive framework that resolves ambiguity with a coupled act-senseact loop, where perception provides evolving sensory information that triggers adaptive decisions, and decisions in turn guide actions to resolve perceptual ambiguity or adopt alternative manipulation strategies, to progress towards task success or generate new observations informing further decisions. This closedloop process illustrates the core principle of Active Perception.  \nA few efforts [31,28,9,40,39,34] have been attempted to tackle the above mentioned perceptual passivity issue. However, these approaches treat active perception control (such as head/eye movement) merely as an extra action dimension, and optimize for immediate task completion. Consequently, these models are commonly confined to head movement-based vie","cbCaiauYX7bzwxHa","https://ap.wps.com/l/cbCaiauYX7bzwxHa","pdf",5240194,2,1,27,"English","en",105,"# Introduction\n## Problem: Perceptual passivity in active perception methods\n## Proposed framing: History-dependent act-sense-act loop\n## Method: CoMe-VLA and training strategy","[{\"question\":\"Why does generalizable manipulation in unstructured environments require active perception?\",\"answer\":\"Because robots must proactively resolve information uncertainty by actively seeking information rather than only reacting to current observations.\"},{\"question\":\"How is active perception formalized in the document?\",\"answer\":\"As a history-dependent perception-action loop driven by information-seeking actions and decision branching, enabling adaptive coupling between perception and action.\"},{\"question\":\"What is CoMe-VLA and how does it improve active perception?\",\"answer\":\"CoMe-VLA is a cognitive and memory-aware vision-language-action framework that leverages large-scale egocentric human data, includes a cognitive auxiliary head for sub-task transitions, and uses a dual-track memory system to maintain consistent self and environmental awareness.\"}]",1784174551,68,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"act-sense-act-learning-active-perception-from-large-scale-egocentric-human-data","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/act-sense-act-learning-active-perception-from-large-scale-egocentric-human-data/81589/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does generalizable manipulation in unstructured environments require active perception?","Question",{"text":75,"@type":76},"Because robots must proactively resolve information uncertainty by actively seeking information rather than only reacting to current observations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is active perception formalized in the document?",{"text":80,"@type":76},"As a history-dependent perception-action loop driven by information-seeking actions and decision branching, enabling adaptive coupling between perception and action.",{"name":82,"@type":73,"acceptedAnswer":83},"What is CoMe-VLA and how does it improve active perception?",{"text":84,"@type":76},"CoMe-VLA is a cognitive and memory-aware vision-language-action framework that leverages large-scale egocentric human data, includes a cognitive auxiliary head for sub-task transitions, and uses a dual-track memory system to maintain consistent self and environmental awareness.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]