[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83210-en":3,"doc-seo-83210-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83210,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models","Vision-language models require spatial reasoning grounded in language and image context, especially when resolving spatial deictic expressions such as “this” and “that”. This work develops a multilingual benchmark to assess how VLMs use these context-dependent references across four languages, including linguistic distinctions tied to spatial distance and ambiguity. Experiments show tested models select demonstratives differently from humans, notably failing to shift choices according to object distance, reflecting weaker human-like spatial deixis understanding.","Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models  \nKaito Watanabe1,2 , Taisei Yamamoto1,2 , Tomoki Doi1,2 , Hitomi Yanaka1,2,3  \n1 The University of Tokyo, 2 Riken, 3 Tohoku University  \n{nglhdf,yamamo96,doi-tomoki701,[hyanaka}@is.s.u-tokyo.ac.jp](hyanaka}@is.s.u-tokyo.ac.jp)  \narXiv :2607 .0725 1v 1 [ cs .CL] 8 Jul 2026  \nAbstract  \nOne of the expected abilities of vision-language models (VLMs) is spatial reasoning ability based on a given text and image. To evaluate the spatial reasoning abilities of VLMs, we focus on the use of spatial deictic expressions, which are defined as spatial expressions whose referent is determined by their situational context, such as “this” and “that”. To handle spatial deictic expressions, VLMs must jointly reason over language and visual space, grounding context-dependent references in the image’s spatial structure. In addition, selecting appropriate spatial deictic expressions across languages requires VLMs to understand the language-specific spatial distinctions encoded by these expressions. In this paper, we develop a benchmark1 to evaluate the multilingual ability of VLMs to use spatial deictic expressionsin four languages. Our experiments using this benchmark reveal that the tested models use demonstratives in a manner different from that of humans, particularly in selecting the appropriate demonstratives based on the distance to the object.  \n1 Introduction  \nIn recent years, large language models (LLMs) and vision-language models (VLMs) have achieved remarkable progress due to their scalability. Furthermore, a key feature of LLMs and VLMs is their ability to handle a wide range of tasks not only in English but also in multiple languages. The development prompted vigorous attempts to evaluate their reasoning ability. Previous studies have investigated the abilities of VLMs to capture spatial relations and spatial expressions, such as frames of reference2 (Zhang et al., 2025b ; Khemlani et al., 2025) .  \n1Our benchmark is available in [https://github.com/](https://github.com/)[ ](https://github.com/)ynklab/multilingual-demonstratives-eval  \n2Frames of reference(FoR) are frameworks which are used to express the relative position of an object from the perspective of the other object.  \nDespite efforts, the ability of VLMs to utilize spatial deictic expressions, an important type of spatial expression, has not been explicitly studied. Deixis is the usage of expressions whose referent is dependent on the situation of utterance. Spatial deictic expressions are expressions that depend on the space in which the utterance occurred, such as“here” or “that”. For example, suppose there is a pen in front of the speaker. If the pen is far away, especially at an unreachable point, the speaker tends to say “that pen”, while if it is near the speaker, the speaker may indicate it by “this pen”. As this example shows, deictic functions are fundamental to human language because they connect linguistic expressions with the physical environment. Therefore, the evaluation of a VLM’s proficiency in using spatial deictic expressions has linguistic significance for benchmarking its capacity for spatial reasoning.  \nWe argue that the use of spatial deixis poses two major challenges for VLMs: cross-linguistic variation and inherent ambiguity. First, to use spatial deictic expressions appropriately, multilingual VLMs need to understand their semantic differences across languages. For example, while in English we use two demonstratives, proximal 3 “this” and distal “that”, in Japanese we use three demonstratives, proximal kono, distal ano, and medial sono. As such, across languages, the number of kinds of spatial deictic expressions varies. In addition, as in another example, although Spanish and Japanese both have three demonstratives: proximal, distal, and medial, the meaning of each word is not the same. According to Diessel (1999), Spanish medial ese signifies a rela","cbCaihHoGQfwsuNd","https://ap.wps.com/l/cbCaihHoGQfwsuNd","pdf",2017696,4,1,9,"English","en",105,"# Abstract\n# Introduction\n## Spatial deixis and challenges\n## Cross-linguistic variation and ambiguity\n# Background\n## Analysis of spatial reasoning ability in VLMs","[{\"question\":\"What ability does this paper evaluate in vision-language models?\",\"answer\":\"It evaluates spatial reasoning through the use of spatial deictic expressions, where the reference depends on situational context (e.g., “this” and “that”).\"},{\"question\":\"How does the proposed benchmark test multilingual capability?\",\"answer\":\"The benchmark is designed to evaluate how VLMs select appropriate spatial demonstratives across four languages, focusing on language-specific spatial distinctions and how absolute distance affects choices.\"},{\"question\":\"What do the experiments reveal about how models differ from humans?\",\"answer\":\"The tested models use demonstratives differently from humans, especially by not shifting demonstrative selection with object distance the way people do.\"}]",1784185965,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"evaluation-of-multilingual-ability-to-use-spatial-deictic-expressions-in-vision-language-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/evaluation-of-multilingual-ability-to-use-spatial-deictic-expressions-in-vision-language-models/83210/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What ability does this paper evaluate in vision-language models?","Question",{"text":75,"@type":76},"It evaluates spatial reasoning through the use of spatial deictic expressions, where the reference depends on situational context (e.g., “this” and “that”).","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed benchmark test multilingual capability?",{"text":80,"@type":76},"The benchmark is designed to evaluate how VLMs select appropriate spatial demonstratives across four languages, focusing on language-specific spatial distinctions and how absolute distance affects choices.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the experiments reveal about how models differ from humans?",{"text":84,"@type":76},"The tested models use demonstratives differently from humans, especially by not shifting demonstrative selection with object distance the way people do.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]