[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84283-en":3,"doc-seo-84283-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84283,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Time-to-Collision Based Dynamic Obstacle Avoidance Using Pretrained Vision Models for Robots in Unstructured Environments","Dynamic obstacle avoidance in unstructured outdoor settings remains difficult for autonomous mobile robots, especially when collecting large robot-specific datasets or building simulation-trained policies is impractical. This work proposes a data-efficient, interpretable vision pipeline that uses real-world data only, avoiding the sim-to-real transfer gap. UniDepth generates dense monocular depth from RGB video, while SuperPoint/SuperGlue track keypoints across long sequences. Bundle adjustment with camera intrinsics yields 3D keypoints and per-keypoint TTC. A ground-plane motion primitive steers the robot away from the minimum-TTC point, achieving strong precision/recall on M3ED and detecting TTC \u003C 1s for most test obstacles with minimal hyperparameter tuning data.","Time-to-Collision Based Dynamic Obstacle Avoidance Using Pretrained Vision Models for Robots in Unstructured Environments  \nErik Jagnandan 1,* , Mulugeta Haile2,†, Gregory Barber2 Pratik Chaudhari 1  \n1 GRASP Laboratory, University of Pennsylvania, PA  \n2U.S. Army Research Laboratory, Aberdeen Proving Ground, MD  \narXiv :2607 .07885v 1 [ cs .RO] 8 Jul 2026  \nAbstract—Dynamic obstacle avoidance in unstructured outdoor environments remains a critical challenge for autonomous mobile robots, particularly when large-scale robot-specific training data and simulation-based policies are impractical. We present a data-efficient, interpretable method for vision-based dynamic obstacle avoidance that operates entirely on real-world data, avoiding the sim-to-real transfer problem inherent in simulation-trained policies. Our approach leverages UniDepth, a large pretrained monocular depth estimation model, to produce dense depth maps from RGB video without requiring stereo cameras or LiDAR at inference time. Dynamic obstacle avoidance is achieved by extending the SuperPoint and SuperGlue feature correspondence pipeline to track keypoints across long frame sequences, projecting their 2D pixel-space positions into 3D using camera intrinsics and predicted depth, running bundle adjustment initialized from these 3D keypoints, and computing per-keypoint time-to-collision (TTC). A 2D motion primitive in the ground plane is then selected to move the robot away from the closest point of approach of the minimum-TTC keypoint. Evaluated on real-world data from the M3ED dataset, our pipeline achieves a precision of 0.49 and a recall of 0.38 in identifying frames with a ground truth TTC below 1 second, and correctly generates the evasive motion direction in 84% of true positive detections. Crucially, it detects at least one frame with TTC less than 1 second for 20 out of 22 unique physical obstacles present in our test sequences. Unlike end-to-end learned methods that demand thousands of hours of robot-specific training data, our approach eliminates model training entirely, requiring only 74 seconds of data for hyperparameter tuning. This demonstrates exceptional data efficiency while preserving interpretable and generalizable behavior across diverse obstacle types.  \nI. INTRODUCTION  \nTransformers have emerged as powerful architectures for robotic perception and control, achieving state-of-the-art performance and strong generalization across tasks. Prior work in robotic manipulation and visual navigation shows that transformer-based models can outperform traditional convolutional neural network (CNN) and recurrent neural network (RNN) based approaches when trained on sufficient data [1]  \n[2] [3] [4] [5] . However, collecting the large-scale robotspecific datasets required to train such models from scratch is prohibitively expensive. For example, Robot Transformer [1] required 17 months of continuous data collection from 13 robots, creating a major bottleneck to scalability.  \n†Corresponding [author: mulugeta.a.haile.civ@army.mil](author: mulugeta.a.haile.civ@army.mil)  \n*[erik.jagnandan@gmail.com](erik.jagnandan@gmail.com)  \nA common alternative is to train reinforcement learning (RL) policies in simulation and deploy them in the real world. While simulation removes data constraints, it introduces a“sim-to-real” gap, where performance does not fully transfer due to discrepancies in visual realism, contact dynamics, and environmental diversity between the simulated and real environments. Although techniques such as domain randomization [6] and SimOpt [7] mitigate this issue, simulationbased training necessarily presents two other key limitations. First, policies learn only from simulated experience, making their behavior unpredictable in real-world scenarios, especially in complex, unstructured environments with unforeseen obstacles. Second, RL policies are typically black-box models with limited interpretability, making it difficult to understand or gua","cbCais4C7mvipCIW","https://ap.wps.com/l/cbCais4C7mvipCIW","pdf",8493162,3,1,9,"English","en",105,"# Abstract\n# I. Introduction\n## Motivation and limitations of data collection and sim-to-real RL\n## Proposed approach using pretrained vision foundation models","[{\"question\":\"What problem does the method address for mobile robots?\",\"answer\":\"It tackles dynamic obstacle avoidance in unstructured outdoor environments where large-scale robot-specific training data and simulation-based policies are impractical.\"},{\"question\":\"How does the approach compute dynamic avoidance decisions?\",\"answer\":\"It generates dense monocular depth with UniDepth, tracks keypoints across frames using SuperPoint/SuperGlue, reconstructs 3D keypoints via bundle adjustment, computes per-keypoint time-to-collision (TTC), and selects a ground-plane motion primitive to move away from the minimum-TTC keypoint.\"},{\"question\":\"What results does the method report on real-world data?\",\"answer\":\"On M3ED, it reports precision 0.49 and recall 0.38 for detecting frames with ground-truth TTC below 1 second, produces the correct evasive direction in 84% of true positives, and finds TTC \\u003c 1 second for 20 out of 22 physical obstacles.\"}]",1784194580,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"time-to-collision-based-dynamic-obstacle-avoidance-using-pretrained-vision-models-for-robots-in-unstructured-environments","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/time-to-collision-based-dynamic-obstacle-avoidance-using-pretrained-vision-models-for-robots-in-unstructured-environments/84283/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the method address for mobile robots?","Question",{"text":75,"@type":76},"It tackles dynamic obstacle avoidance in unstructured outdoor environments where large-scale robot-specific training data and simulation-based policies are impractical.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the approach compute dynamic avoidance decisions?",{"text":80,"@type":76},"It generates dense monocular depth with UniDepth, tracks keypoints across frames using SuperPoint/SuperGlue, reconstructs 3D keypoints via bundle adjustment, computes per-keypoint time-to-collision (TTC), and selects a ground-plane motion primitive to move away from the minimum-TTC keypoint.",{"name":82,"@type":73,"acceptedAnswer":83},"What results does the method report on real-world data?",{"text":84,"@type":76},"On M3ED, it reports precision 0.49 and recall 0.38 for detecting frames with ground-truth TTC below 1 second, produces the correct evasive direction in 84% of true positives, and finds TTC \u003C 1 second for 20 out of 22 physical obstacles.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]