[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83806-en":3,"doc-seo-83806-105":28,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":13,"seo_description":14,"update_tm":26,"read_time":27},83806,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","SurgAM Surgical Affordance Map Prediction with Multimodal Feature Fusion for Robot Autonomy","SurgAM presents surgical affordance map prediction to bridge visual scene understanding and autonomous action planning for surgical robot autonomy. The work proposes an adaptive multimodal feature fusion framework that combines a self-supervised vision transformer encoder for semantic understanding with a large-scale generative model encoder for spatial awareness. A hierarchical prompt learning strategy adapts to diverse procedural contexts, and a scene-guided attention decoder targets critical surgical regions while suppressing background noise. A new dataset covers three actions—aspiration, clipping, and retraction—and extensive experiments show state-of-the-art performance, validated on lung and prostate phantoms enabling autonomous actions.","SurgAM: Surgical Affordance Map Prediction with Multimodal Feature Fusion for Robot Autonomy  \nLei Song, Yonghao Long, Mengya Xu, Jiayi Geng, Xiuyuan Chen†, Qi Dou†  \narXiv :2607 .04378v 1 [ cs .RO] 5 Jul 2026  \nAbstract—Surgical automation is being increasingly studied, yet bridging visual scene understanding with autonomous action planning remains a fundamental challenge. While much research effort has been made on scene perception (e.g., tool recognition and scene segmentation), understanding and predicting actionable possibilities for surgical automation is still underexplored. In this paper, we introduce surgical affordance prediction, which identifies actionable regions for fundamental surgical actions from visual data. Specifically, a novel adaptive feature fusion framework is proposed that leverages the complementary strengths of a self-supervised vision transformer encoder for its superior semantic understanding and a largescale generative model encoder for its spatially-aware capability. Furthermore, we introduce a hierarchical prompt learning mechanism to adapt to varying procedural contexts. Finally, a scene-guided attention decoder is proposed to focus on critical surgical areas while suppressing background distractions. To validate the effectiveness, we established a new dataset, derived from publicly available surgical datasets with affordance annotations for three basic surgical actions: aspiration, clipping, and retraction. Extensive experiments demonstrate that our approach achieves state-of-the-art performance. Moreover, we validate our framework’s applicability for downstream automation on a realistic lung and prostate phantom, and results show that the predicted affordance maps successfully enable autonomous surgical actions.  \nI. INTRODUCTION  \nRobotic surgery has been increasingly adopted in modern healthcare, demonstrating promising clinical benefits. Building on this success, there is a growing trend toward surgical automation [1], which promises to further reduce surgeon workload while improving procedural efficiency toward higher-level autonomy and human-robot collaboration [2] . To this end, autonomous robotic surgical systems must not only recognize what is in the surgical scene (e.g., surgical instruments or anatomical structure), but also interpret the scene in a way that directly supports manipulation, understanding what can be done where, and how to do it safely and effectively [3] . Typical perception tasks, such as key point extraction, instrument detection, and scene segmentation [4],  \nL. Song, Y. Long, M. Xu, and Q. Dou are with the Department of Computer Science and Engineering, The Chinese University of Hong Kong.  \nJ. Geng and X. Chen are with the Department of Thoracic Surgery, Peking University People’s Hospital, Beijing, China. J. Geng and X. Chen are with the Thoracic Oncology Institute, Peking University People’s Hospital, Beijing, China. J. Geng and X. Chen are with the Research Unit of Intelligence Diagnosis and Treatment in Early Non-small Cell Lung Cancer, Chinese Academy of Medical Sciences, 2021RU002, Peking University People’s Hospital, Beijing, China. J. Geng and X. Chen are with the Institute of Advanced Clinical Medicine, Peking University, Beijing, China. J. Geng and X. Chen are with Beijing Key Laboratory of Innovative Application of Big Data in Lung Cancer, Peking University People’s Hospital, Beijing, China. Corresponding authors: Qi Dou ([qidou@cuhk.edu.hk](qidou@cuhk.edu.hk)), Xiuyuan Chen (dr [chenxy@pku.edu.cn](chenxy@pku.edu.cn)).  \nFig. 1. Overall concept for surgical scene affordance map prediction: Model generates affordance map to identify optimal manipulation regions for specific surgical tasks, enabling downstream applications in robotic surgery and surgical planning.  \nhave been widely studied, which provided valuable scene information. However, they just describe what is in the field of view, not where the robot should act [5] . The bridge between visual ","cbCaivcE0txjSvlj","https://ap.wps.com/l/cbCaivcE0txjSvlj","pdf",18686183,1,"English","en",105,"# Introduction\n## Bridging scene understanding and action planning\n## Limitations of existing surgical automation methods\n## Affordance prediction as an interpretable bridge\n## Motivation for surgical affordance prediction","[{\"question\":\"What does SurgAM predict for surgical robot autonomy?\",\"answer\":\"SurgAM predicts surgical affordance maps that identify actionable regions for fundamental surgical actions directly from visual data.\"},{\"question\":\"How does the proposed framework fuse multimodal features?\",\"answer\":\"It adaptively fuses features from a self-supervised vision transformer encoder for semantic understanding and a large-scale generative model encoder for spatially aware representations.\"},{\"question\":\"What dataset and surgical actions are used to evaluate the method?\",\"answer\":\"A new dataset is built from publicly available surgical datasets with affordance annotations for three actions: aspiration, clipping, and retraction.\"}]",1784190525,20,{"code":4,"msg":29,"data":30},"ok",{"site_id":23,"language":22,"slug":31,"title":13,"keywords":32,"description":14,"schema_data":33,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":26},"surgam-surgical-affordance-map-prediction-with-multimodal-feature-fusion-for-robot-autonomy","",{"@graph":34,"@context":84},[35,52,67],{"@type":36,"itemListElement":37},"BreadcrumbList",[38,42,46,49],{"item":39,"name":40,"@type":41,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":43,"name":44,"@type":41,"position":45},"https://docshare.wps.com/document/","Document",2,{"item":47,"name":12,"@type":41,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":41,"position":51},"https://docshare.wps.com/document/surgam-surgical-affordance-map-prediction-with-multimodal-feature-fusion-for-robot-autonomy/83806/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":39,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What does SurgAM predict for surgical robot autonomy?","Question",{"text":74,"@type":75},"SurgAM predicts surgical affordance maps that identify actionable regions for fundamental surgical actions directly from visual data.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does the proposed framework fuse multimodal features?",{"text":79,"@type":75},"It adaptively fuses features from a self-supervised vision transformer encoder for semantic understanding and a large-scale generative model encoder for spatially aware representations.",{"name":81,"@type":72,"acceptedAnswer":82},"What dataset and surgical actions are used to evaluate the method?",{"text":83,"@type":75},"A new dataset is built from publicly available surgical datasets with affordance annotations for three actions: aspiration, clipping, and retraction.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":44,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":45,"doc_module":4,"doc_module_name":44,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":44,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":44,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":44,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":44,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":44,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":44,"category_name":124,"show_sort_weight":27,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":27,"doc_module":4,"doc_module_name":44,"category_name":127,"show_sort_weight":27,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":44,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":44,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]