[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86166-en":3,"doc-seo-86166-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86166,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","When Depth Is Better Told Than Shown Depth-Ordinal Prompting for Vision-Language Spatial Reasoning","Vision-language models are expected to reason about physical space, yet relative depth and 3D arrangement judgments remain weak even when depth maps are available. Providing depth as an image can worsen performance because frozen models struggle to access and regulate the rendered pseudodepth input. Depth-Ordinal Prompting (DOP) converts monocular depth into a targeted ordinal text cue for queried objects, without adding a depth image, training modules, feature injection, or labels. Across benchmarks, DOP improves spatial reasoning when pseudo-depth yields reliable object-level ordering and stays competitive with training-free depth prompting while remaining simpler.","When Depth Is Better Told Than Shown: Depth-Ordinal Prompting for  \nVision-Language Spatial Reasoning  \nQuynh Vo Phuc Dao Cong-Duy Nguyen Thong Nguyen*  \nNational University of Singapore Center of AI Research, VinUniversity  \n[thong.nguyen@u.nus.edu](thong.nguyen@u.nus.edu)  \narXiv :2607 . 11173v1 [ cs .CV] 13 Jul 2026  \nAbstract  \nVision-language models (VLMs) are expected to reason about physical space—which object is closer, what lies behind what, and how objects are arranged in 3D—yet they still struggle with such spatial judgments. A natural remedy is to show the model a depth map, but we find that this can make performance worse. We show that depth is not absent: it reaches the language model, but becomes difficult to access for downstream reasoning, while rendered pseudodepth maps act as noisy auxiliary images that frozen VLMs cannot easily regulate. We propose Depth-Ordinal Prompting (DOP), a training-free method that converts monocular depth into a single question-targeted ordinal text cue at the queried objects, without adding a depth image, training a module, injecting features, or using labels. Our key finding is form dependence: the same depth signal can hurt when shown as an image but help when told as text.Across benchmarks, models, and depth estimators, DOP improves spatial reasoning when pseudo-depth provides reliable object-level ordering and remains largely neutral in strong originalimage regimes. It is also competitive with the strongest training-free depth-prompting alternative while being simpler and more targeted.  \n1. Introduction  \nVision-language models (VLMs) are increasingly expected to reason about physical space: which object is closer, what lies behind what, and how objects are arranged in 3D. Such capabilities are central to robotics, embodied AI, and assistive perception [2, 3, 5] . Yet spatial reasoning remains a persistent weakness. Recent benchmarks show that even strong VLMs struggle with relative depth, direction, and objectlevel spatial relations [6, 8, 19, 22, 30] . A natural remedy is to provide geometry explicitly, for example by appending a monocular depth map to the original color image or by training models with depth-and 3D-aware supervision [2–  \n*Corresponding author  \nFigure 1 . Tell, don’t show. The same monocular depth signal can hurt when shown as an additional image but help when converted into a targeted ordinal text cue. DOP keeps the original color image and communicates only the queried-object depth relation, allowing the frozen VLM to weigh the cue against its original-image evidence.  \n5, 20, 37] . Surprisingly, we find that the simplest version of this remedy often makes performance worse: when pseudodepth is rendered as an additional image, frozen VLMs can degrade substantially on point-and object-level spatial reasoning. Fig. 1 previews this form contrast: rendering depth as another image can mislead the frozen VLM, whereas expressing the queried-object relation as a short text cue can make the same geometric information easier to use. This raises a basic question: if depth is useful, why does showing depth hurt?  \nThis contrast suggests that the issue is not simply the absence of depth, but the interface through which depth is made usable by a frozen VLM. We therefore ask whether the model fails because it never encodes depth, or because depth becomes difficult to use later in the computation. To answer this, we conduct layer-wise probing of the model’s hidden states, as summarized in Fig. 2. We find that early visual features contain a clear depth signal, and that this signal reaches the first language-model layers with little loss. However, as computation proceeds, the same signal becomes hard to read with a simple linear probe, even though a non-linear probe can still recover it. In other words, depth is not erased; it becomes entangled in a form that  \nis less accessible to the model’s downstream reasoning. A causal intervention further supports this view: when","cbCaiij1nom0RBmB","https://ap.wps.com/l/cbCaiij1nom0RBmB","pdf",4888189,3,1,10,"English","en",105,"# Abstract\n# Introduction\n## Spatial reasoning limitations in VLMs\n## Why showing depth can hurt\n## Probing and causal evidence\n## Depth interface diagnosis\n## Proposed method: Depth-Ordinal Prompting (DOP)","[{\"question\":\"Why can adding a depth map as an image make vision-language spatial reasoning worse?\",\"answer\":\"When rendered as pseudodepth, the extra visual input is hard for frozen VLMs to align, interpret, and selectively ignore, leading to degraded object-level spatial judgments.\"},{\"question\":\"What is Depth-Ordinal Prompting (DOP)?\",\"answer\":\"DOP is a training-free interface that reads monocular pseudo-depth at the queried object regions and converts it into a single ordinal text cue (e.g., which object appears closer), appending it to the original prompt without adding a depth image or labels.\"},{\"question\":\"What is the key finding behind DOP’s effectiveness?\",\"answer\":\"The same depth signal can hurt when shown as an image but help when expressed as a targeted text cue, suggesting a form-dependent accessibility gap for downstream reasoning.\"}]",1784209041,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"when-depth-is-better-told-than-shown-depth-ordinal-prompting-for-vision-language-spatial-reasoning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/when-depth-is-better-told-than-shown-depth-ordinal-prompting-for-vision-language-spatial-reasoning/86166/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why can adding a depth map as an image make vision-language spatial reasoning worse?","Question",{"text":75,"@type":76},"When rendered as pseudodepth, the extra visual input is hard for frozen VLMs to align, interpret, and selectively ignore, leading to degraded object-level spatial judgments.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is Depth-Ordinal Prompting (DOP)?",{"text":80,"@type":76},"DOP is a training-free interface that reads monocular pseudo-depth at the queried object regions and converts it into a single ordinal text cue (e.g., which object appears closer), appending it to the original prompt without adding a depth image or labels.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the key finding behind DOP’s effectiveness?",{"text":84,"@type":76},"The same depth signal can hurt when shown as an image but help when expressed as a targeted text cue, suggesting a form-dependent accessibility gap for downstream reasoning.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]