[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85065-en":3,"doc-seo-85065-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85065,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","UniRef-UAV: A Multimodal Benchmark for Universal Referring in UAV Imagery","Unmanned aerial vehicles increasingly require visual grounding that links diverse instructions to task-relevant targets in complex aerial scenes. Traditional referring expression comprehension benchmarks mostly support text-only queries and single-object outputs, limiting UAV applicability when reference images, multimodal instructions, absent targets, or multiple valid instances are involved. UniRef-UAV introduces Universal Referring and a multimodal benchmark supporting text, image, and text+image queries with modality-dependent target cardinality, plus in-domain and cross-domain evaluation protocols. It also presents UAV-URNet, a set-prediction baseline with stable, reproducible performance and analyses showing multimodal queries reduce ambiguity and unify query–target alignment.","This work has been submitted to the IEEE Transactions on Multimedia for possible publication.  \nCopyright may be transferred without notice, after which this version may no longer be accessible.  \nIEEE TRANSACTIONS ON MULTIMEDIA 1  \nUniRef-UAV: A Multimodal Benchmark for Universal Referring in  \nUAV Imagery  \nHaibin Tian, Huichao Xie, Xuelin Qian, Member, IEEE, Ruitao Lu, Junwei Han, Fellow, IEEE, and  \nDingwen Zhang, Member, IEEE  \narXiv :2607 .08267v 1 [ cs .CV] 9 Jul 2026  \nAbstract—Unmanned aerial vehicles (UAVs) increasingly rely on visual grounding capabilities to localize task-relevant targets from diverse instructions in complex aerial scenes. Existing referring expression comprehension (REC) benchmarks and methods, however, are largely built around text-only queries and single-object outputs, which limits their applicability to practical UAV scenarios involving reference images, multimodal instructions, absent targets, and multiple valid target instances. To address this gap, we introduce Universal Referring, a generalized UAV referring task that jointly expands the query modality and the output cardinality. We construct UniRef-UAV, amultimodal benchmark that supports text-only, image-only, and text+image queries with modality-dependent target cardinality, where text-only and text+image queries admit no-target, singletarget, and multi-target grounding while image-only queries focus on existence-aware single-instance grounding. It also provides indomain and cross-domain evaluation protocols for visual-query generalization. We further present UAV-URNet, a detection-style baseline that maps heterogeneous queries into a shared query space and predicts variable-size target sets through set prediction. Extensive experiments show that UAV-URNet provides a stable and reproducible baseline with more consistent no-target discrimination and a more lightweight, reproducible implementation than large general-purpose MLLMs. Additional domain analysis, query-representation analysis, and ablation studies demonstrate that multimodal queries help reduce visual-query ambiguity and promote a more unified query–target alignment space. The annotations, visual query crops/images, train/validation/test splits, evaluation scripts, and baseline code will be made publicly available to facilitate reproducible research.  \nIndex Terms—UAV vision, multimodal referring, universal referring, visual grounding, variable-cardinality detection.  \nI. INTRODUCTION  \nUNMANNED aerial vehicles (UAVs) have become an im  \nportant mobile sensing platform for security patrols [1], emergency search and rescue [2], [3], traffic monitoring [4], and industrial inspection [5] . Related visual systems have also expanded toward large-scale scene reconstruction and open-vocabulary scene querying [6]–[8], but practical UAV operations still require interfaces that ground task instructions to concrete visual targets. Beyond autonomous flight and navigation, these applications increasingly require UAVs  \nHaibin Tian and Huichao Xie contributed equally to this work.  \nHaibin Tian, Huichao Xie, Xuelin Qian, and Dingwen Zhang are with the School of Automation, Northwestern Polytechnical University, Xi’an 710072, China (e-mail: [haibintian@foxmail.com](haibintian@foxmail.com); [xiehuichao@mail.nwpu.edu.cn](xiehuichao@mail.nwpu.edu.cn); [xlqian@nwpu.edu.cn](xlqian@nwpu.edu.cn); [zdw2006yyy@nwpu.edu.cn](zdw2006yyy@nwpu.edu.cn)). Corresponding author: Dingwen Zhang.  \nRuitao Lu is with the College of Missile Engineering, Rocket Force University of Engineering, Xi’an 710038, China (e-mail: [lrt19880220@163.com](lrt19880220@163.com)).  \nJunwei Han is with the School of Artificial Intelligence, Chongqing University of Posts and Telecommunications, Chongqing 400065, China (email: [jhan@nwpu.edu.cn](jhan@nwpu.edu.cn)).  \nto interpret task-level instructions and associate them with specific visual targets in complex aerial scenes. For example, a rescue UAV must localize the indicated vic","cbCairQJ7nzsGa9G","https://ap.wps.com/l/cbCairQJ7nzsGa9G","pdf",3286328,2,1,14,"English","en",105,"# Introduction\n## Motivation and problem setting\n## Limitations of existing REC in UAV scenarios\n## Universal Referring and multimodal inputs\n## Proposed benchmark and baseline (overview)","[{\"question\":\"What problem does UniRef-UAV address in UAV referring tasks?\",\"answer\":\"It addresses the mismatch between existing REC benchmarks (typically text-only, single-target outputs) and practical UAV needs involving multimodal instructions, reference images, absent targets, and multiple valid instances in aerial scenes.\"},{\"question\":\"What query modalities and output cardinalities does the UniRef-UAV benchmark support?\",\"answer\":\"UniRef-UAV supports text-only, image-only, and text+image queries. Target cardinality depends on modality: text-only and text+image support no-target, single-target, and multi-target grounding, while image-only focuses on existence-aware single-instance grounding.\"},{\"question\":\"What is UAV-URNet and how does it perform query understanding?\",\"answer\":\"UAV-URNet is a detection-style baseline that maps heterogeneous queries into a shared query space and uses set prediction to output variable-size target sets, providing a stable and reproducible baseline with improved no-target discrimination.\"}]",1784200768,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"uniref-uav-a-multimodal-benchmark-for-universal-referring-in-uav-imagery","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/uniref-uav-a-multimodal-benchmark-for-universal-referring-in-uav-imagery/85065/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does UniRef-UAV address in UAV referring tasks?","Question",{"text":75,"@type":76},"It addresses the mismatch between existing REC benchmarks (typically text-only, single-target outputs) and practical UAV needs involving multimodal instructions, reference images, absent targets, and multiple valid instances in aerial scenes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What query modalities and output cardinalities does the UniRef-UAV benchmark support?",{"text":80,"@type":76},"UniRef-UAV supports text-only, image-only, and text+image queries. Target cardinality depends on modality: text-only and text+image support no-target, single-target, and multi-target grounding, while image-only focuses on existence-aware single-instance grounding.",{"name":82,"@type":73,"acceptedAnswer":83},"What is UAV-URNet and how does it perform query understanding?",{"text":84,"@type":76},"UAV-URNet is a detection-style baseline that maps heterogeneous queries into a shared query space and uses set prediction to output variable-size target sets, providing a stable and reproducible baseline with improved no-target discrimination.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]