[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83303-en":3,"doc-seo-83303-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83303,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Monocular Vision Based Control Framework for Grasping","Grasping items in unstructured settings demands handling objects with widely varying mechanical properties, from soft deformable materials to rigid everyday items. A unified monocular vision-based framework is proposed to control a position-controlled gripper using only RGB input. It integrates open-vocabulary object detection, image segmentation, boundary-aware point assignment, real-time point tracking, and monocular depth estimation to recover object motion and geometry. A language-driven stiffness estimation model infers expected compliance from semantic descriptions, enabling object-level control mode selection before contact. Procrustes-based keypoint dissimilarity guides deformation adaptation, while rigid-object gripper width is set by scaling tracked point distances. Real-world pick-and-place experiments validate stable grasping across lettuce, mozzarella cheese, croissants, paper towels, and hard plastic bottles via vision-only feedback for generalizable, sensor-efficient household manipulation.","Monocular Vision Based Control Framework for Grasping  \nShail Jadav 1 and Dongheui Lee 1 ,2  \narXiv :2607 .07897v 1 [ cs .RO] 8 Jul 2026  \nAbstract—Grasping in unstructured environments requires handling objects with widely different mechanical properties, from soft and deformable items to rigid everyday objects. Most existing approaches address these categories separately and often rely on tactile sensing, object-specific models, or specialized grippers. In this paper, we present a unified monocular visionbased grasping framework that targets both soft and rigid objects within a single control pipeline, using only RGB input and a position-controlled gripper. The proposed system combines open-vocabulary object detection, image segmentation, boundary-aware point assignment, real-time point tracking, and monocular depth estimation to recover object motion and geometry from visual observations. A key component of the framework is a language-based stiffness estimation model that infers an object’s expected compliance from its semantic description and provides an object-level prior for selecting the grasping strategy before contact. For deformable objects, grasp adaptation is governed by a Procrustes-based dissimilarity measure computed from tracked keypoints, which acts as a visual proxy for deformation. For rigid objects, the gripper width is regulated through the scaling of tracked point distances. We validate the proposed method in real-world pickand-place experiments on a Franka Emika Research 3 arm using objects with substantially different mechanical properties, including lettuce, fresh mozzarella cheese, croissants, paper towels, and hard plastic bottles. Results demonstrate that the framework achieves stable grasping across both soft and rigid objects using visual feedback alone, highlighting a practical, sensor-efficient, and generalizable approach for food handling and household manipulation.  \nI. INTRODUCTION  \nAs robots increasingly integrate into our daily lives, their ability to manipulate various objects, particularly deformable items, becomes crucial [1] . These objects include a wide range of items, from food products to flexible materials like cloths and rigid objects. The successful grasping of deformable objects offers significant potential for many industries, especially in food processing and home automation [2] . However, the manipulation of deformable objects presents a unique challenge due to their high degrees of freedom and variable material properties [2] . Traditional force-based grasping methods are often inadequate for manipulating deformable objects, as they do not consider object deformation [2], [3] . Consequently, advanced sensing and control strategies are necessary to overcome these challenges.  \n1 Shail Jadav and Dongheui Lee are with Autonomous Systems, Technische Universitt Wien (TU Wien), Vienna, Austria (e-mail: [shail.jadav@tuwien.ac.at](shail.jadav@tuwien.ac.at) , [dongheui.lee@tuwien.ac.at](dongheui.lee@tuwien.ac.at)).  \n2Dongheui Lee is also with the Institute of Robotics and Mechatronics (DLR), German Aerospace Center, Wessling, Germany.  \nThis work was supported by the Vienna Science and Technology Fund (WWTF) under the project SafeDiffusion (ICT25068) and by the European Union project INVERSE (No. 101136067) .  \nDeformation Metric  \nScaling Factor  \n3D Optical Flow  \nGeometric State Estimation  \nSegmentation  \nPoint Tracking  \nDepth Estimation  \nVisual Perception  \nLanguage-Driven Material Prior  \n“Text”  StiffNET   \nEstimated  \nStiffness  \nControl Mode Selection  \nFig. 1: Overview of the proposed framework. A unified monocular vision-based grasping system enables a standard position controlled gripper to handle both compliant and rigid objects using RGB input alone. By combining semantic priors from language with visual feedback during manipulation, the framework selects an appropriate grasping behavior before contact and adapts online to maintain stable grasps across diverse every","cbCaieqqOvyqGsnc","https://ap.wps.com/l/cbCaieqqOvyqGsnc","pdf",3855844,3,1,"English","en",105,"# Introduction\n## Challenges in deformable object grasping\n## Limitations of vision-based tactile sensing\n## Mechanics-based modeling and semantic prior approaches","[{\"question\":\"What problem does the proposed framework address for robot grasping?\",\"answer\":\"It targets stable grasping in unstructured environments where objects can vary widely in mechanical properties, including both soft deformable and rigid items.\"},{\"question\":\"What sensing and inputs does the framework require?\",\"answer\":\"The system uses only RGB input and a position-controlled gripper, combining object detection, segmentation, point tracking, and monocular depth estimation for visual state recovery.\"},{\"question\":\"How does the framework adapt grasping for deformable versus rigid objects?\",\"answer\":\"For deformable objects, it adapts using a Procrustes-based dissimilarity measure computed from tracked keypoints as a deformation proxy; for rigid objects, it regulates gripper width by scaling tracked point distances.\"}]",1784186630,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"monocular-vision-based-control-framework-for-grasping","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":20},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/monocular-vision-based-control-framework-for-grasping/83303/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does the proposed framework address for robot grasping?","Question",{"text":74,"@type":75},"It targets stable grasping in unstructured environments where objects can vary widely in mechanical properties, including both soft deformable and rigid items.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What sensing and inputs does the framework require?",{"text":79,"@type":75},"The system uses only RGB input and a position-controlled gripper, combining object detection, segmentation, point tracking, and monocular depth estimation for visual state recovery.",{"name":81,"@type":72,"acceptedAnswer":82},"How does the framework adapt grasping for deformable versus rigid objects?",{"text":83,"@type":75},"For deformable objects, it adapts using a Procrustes-based dissimilarity measure computed from tracked keypoints as a deformation proxy; for rigid objects, it regulates gripper width by scaling tracked point distances.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]