[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85530-en":3,"doc-seo-85530-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85530,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","InCoM Intent-Driven Perception and Structured Coordination for Mobile Manipulation","Mobile manipulation enables general-purpose robotic agents, but it requires coordinated control of both the mobile base and manipulator and robust perception as viewpoints change. Existing methods struggle with strong coupling between base and arm action decisions and with perceptual attention being poorly allocated during motion. InCoM introduces intent-driven perception that infers latent motion intent to reweight multi-scale features for stage-adaptive attention. It also uses geometric-semantic alignment for cross-modal robustness and a decoupled coordinated flow matching decoder to reduce optimization difficulties.","arXiv :2602 .23024v5 [ cs .RO] 13 Jul 2026  \nInCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation  \nJiahao Liu 1 ,2 , Wenbo Cui 1 , Zhongpu Xia3 , Yongliang Wang 1 , Haoran Li 1 ,∗ , Dongbin Zhao 1 ,∗  \n1Institute of Automation, Chinese Academy of Sciences  \n2 School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences  \n3Anyverse Dynamics  \n∗ Corresponding author  \n[liujiahao2077@gmail.com](liujiahao2077@gmail.com), [xiazhongpu5@163.com](xiazhongpu5@163.com)  \n{cuiwenbo2023, lihaoran2015, [yongliangwang1997@gmail.com](yongliangwang1997@gmail.com) , [dongbin.zhao](dongbin.zhao}@ia.ac.cn)[}](dongbin.zhao}@ia.ac.cn)[@ia.ac.cn](dongbin.zhao}@ia.ac.cn),  \nAbstract: Mobile manipulation is a fundamental capability for general-purpose  \nrobotic agents, requiring both coordinated control of the mobile base and manip  \nulator and robust perception under dynamically changing viewpoints. However,  \nexisting approaches face two key challenges: strong coupling between base and  \narm actions complicates control optimization, and perceptual attention is often  \npoorly allocated as viewpoints shift during mobile manipulation. We propose  \nInCoM, an intent-driven perception and structured coordination framework for  \nmobile manipulation. InCoM infers latent motion intent to dynamically reweight  \nmulti-scale perceptual features, enabling stage-adaptive allocation of perceptual  \nattention. To support robust cross-modal perception, InCoM further incorporates  \na geometric-semantic structured alignment mechanism that enhances multimodal  \ncorrespondence. On the control side, we design a decoupled coordinated flow  \nmatching action decoder that explicitly models coordinated base-arm action genera  \ntion, alleviating optimization difficulties caused by control coupling. Experimental  \nresults demonstrate that InCoM significantly outperforms state-of-the-art methods,  \nachieving success rate gains of 28.2%, 26.1%, and 23.6% across three ManiSkill  \nHAB scenarios without privileged information. Furthermore, its effectiveness  \nis consistently validated in real-world mobile manipulation tasks, where InCoM  \nmaintains a superior success rate over existing baselines. The project website is  \navailable at [https://incom-anonymous.github.io/InCoM.github.io/](https://incom-anonymous.github.io/InCoM.github.io/) .  \nKeywords: Mobile Manipulation, Imitation Learning, Robotics  \n1 Introduction  \nMobile manipulation is a fundamental capability for realizing general-purpose robotic policies in realworld environments and has attracted increasing attention from both academia and industry [1, 2, 3, 4] . Compared with tabletop manipulators or mobile bases, mobile manipulation requires simultaneous control of the robotic arm and the mobile base, as well as continuous decision-making and control in complex, dynamic, and partially observable environments [5, 6] . A key challenge is generating stable coordinated actions between the mobile base and the manipulator. Due to inevitable execution errors in real robots, small deviations in base motion can be amplified at the manipulator end-effector under decoupled or weakly coordinated control [3], leading to degraded manipulation accuracy or task failure. Introducing conditional dependencies between the mobile base and the manipulator atthe action decoding stage [2, 3] can partially alleviate this issue. However, such joint action modeling typically represents coordination as a one-way dependency, limiting the model’s ability to capture the bidirectional coupling and mutual compensation inherent in real-world tasks.  \nAnother important yet often overlooked challenge is the allocation of perceptual information under dynamically changing viewpoints, which we refer to as the dynamic perceptual attention problem. Unlike single-mode settings such as tabletop manipulation or navigation, mobile manipulation combines and switches among multiple operational modes with dist","cbCaijSUkhFqZGFh","https://ap.wps.com/l/cbCaijSUkhFqZGFh","pdf",10158549,1,21,"English","en",105,"# Abstract\n# Introduction\n## Coordinated base–arm action generation challenges\n## Dynamic perceptual attention under changing viewpoints\n## Overview of the InCoM framework and contributions","[{\"question\":\"What are the two main challenges InCoM targets in mobile manipulation?\",\"answer\":\"InCoM targets strong coupling between mobile-base and arm actions that complicates control optimization, and the dynamic perceptual attention problem where attention is not well reallocated as viewpoints shift during execution.\"},{\"question\":\"How does InCoM handle dynamic perceptual attention during changing viewpoints?\",\"answer\":\"InCoM infers latent motion intent and dynamically reweights multi-scale perceptual features so perceptual attention adapts to the current stage of the task.\"},{\"question\":\"What mechanism improves cross-modal perception in InCoM?\",\"answer\":\"It incorporates a geometric-semantic structured alignment mechanism that enhances correspondence across modalities, improving multimodal fusion.\"}]",1784204204,53,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"incom-intent-driven-perception-and-structured-coordination-for-mobile-manipulation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/incom-intent-driven-perception-and-structured-coordination-for-mobile-manipulation/85530/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What are the two main challenges InCoM targets in mobile manipulation?","Question",{"text":75,"@type":76},"InCoM targets strong coupling between mobile-base and arm actions that complicates control optimization, and the dynamic perceptual attention problem where attention is not well reallocated as viewpoints shift during execution.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does InCoM handle dynamic perceptual attention during changing viewpoints?",{"text":80,"@type":76},"InCoM infers latent motion intent and dynamically reweights multi-scale perceptual features so perceptual attention adapts to the current stage of the task.",{"name":82,"@type":73,"acceptedAnswer":83},"What mechanism improves cross-modal perception in InCoM?",{"text":84,"@type":76},"It incorporates a geometric-semantic structured alignment mechanism that enhances correspondence across modalities, improving multimodal fusion.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]