[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84483-en":3,"doc-seo-84483-105":28,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":13,"seo_description":14,"update_tm":26,"read_time":27},84483,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","AeroRAG Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning","AeroRAG targets reliable visual question answering in aerial imagery where critical evidence is often carried by small objects, explicit quantities, coarse locations, and inter-object relations. Conventional dense visual-token representations lack alignment with these structured semantics. The method introduces a scene-graph-guided multimodal retrieval-augmented generation framework: it converts images into structured visual knowledge and retrieves query-relevant semantic chunks to build compact prompts. Experiments on AUG and VG-150 show consistent gains over strong MLLM baselines, with larger improvements in dense and relation-sensitive scenes.","AeroRAG: Structured Multimodal Retrieval-Augmented LLM for Fine-Grained Aerial Visual Reasoning  \nJunxiao Xue 1†, Quan Deng2 , 1†, Tingqi Hu3 , Meicong Si3 , Xinyi Yin3 , Yunyun Shi4 , and Xuecheng Wu4∗‡  \narXiv :2604 . 17889v2 [ cs .CV] 12 Jul 2026  \nAbstract—Recent progress in multimodal large language models (MLLMs) has yet to resolve the challenge of reliable visual question answering in aerial imagery. In such scenes, task-critical evidence is often carried by small objects, explicit quantities, coarse locations, and inter-object relations, whereas conventional dense visual-token representations are not well aligned with these structured semantics. To address this interface mismatch, we propose AeroRAG, a scene-graphguided multimodal retrieval-augmented generation framework for visual question answering. The framework first convertsan input image into structured visual knowledge, including object categories, quantities, spatial locations, and semantic relations, and then retrieves query-relevant semantic chunks to construct compact prompts for a text-based large language model. Rather than relying on direct reasoning over dense visual tokens, our method introduces a more explicit intermediate interface between perception and language reasoning. Experiments on the AUG aerial dataset and the generaldomain VG-150 benchmark show consistent improvements over six strong MLLM baselines, with the largest gains observed in dense aerial scenes and relation-sensitive reasoning. We further evaluate the framework on VQAv2 to demonstrate that the proposed interface remains compatible with standard visual reasoning settings. These results suggest that structured retrieval is a practical design direction for deployment-oriented and grounded visual reasoning systems.  \nI. INTRODUCTION  \nMultimodal large language models (MLLMs) have become essential frameworks for merging visual and linguistic understanding, fueled by various advanced models [1]–[3] . These approaches connect pre-trained vision encoders with large language models (LLMs), achieving outstanding performance in general visual question answering (VQA) [4],[5] and image captioning [6] . At the same time, there is increasing interest in deploying such models in mission critical visual analysis scenarios, especially in aerospace  \n*This work was supported by the Central Government Guiding Local Science and Technology Development Fund under Grant No. 2026ZY04001 .  \n1Research Center for Space Computing System, Zhejiang Lab, Hangzhou 311100, China (E-mail: [xuejx@zhejianglab.cn](xuejx@zhejianglab.cn));  \n2Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences, Hangzhou 311000, China (E-mail: [dengquan23@mails.ucas.ac.cn](dengquan23@mails.ucas.ac.cn));  \n3 School of Cyber Science and Engineering, Zhengzhou University, Zhengzhou 450002, China (E-mail: {htqiserendipity, smc ggjxb, [yinxinyi](yinxinyi}@stu.zzu.edu.cn)[}](yinxinyi}@stu.zzu.edu.cn)[@stu.zzu.edu.cn](yinxinyi}@stu.zzu.edu.cn)).  \n4 School of Computer Science and Technology, Xi’an Jiaotong University, Xi’an 710049, China (E-mail: [yunyunshi@stu.xjtu.edu.cn](yunyunshi@stu.xjtu.edu.cn); [wuxc3@ieee.org](wuxc3@ieee.org));  \nWork done during Quan Deng’s research internship at Zhejiang Lab.  \n∗ Corresponding author: Xuecheng Wu.  \n†Equal contributions. ‡Project lead.  \nFig. 1. The qualitative comparison on challenging aerial VQA examples. Existing MLLMs often miss small objects or confuse counts and spatial relations, whereas our proposed multimodal RAG framework provides more accurate and visually grounded answers.  \nand aerial imagery applications such as remote sensing, environmental monitoring, urban planning, and autonomous inspection. In these settings, the requirement is not only strong perception, but also reliable and grounded reasoning over complex scenes.  \nHowever, directly applying existing MLLMs to aerial VQA remains challenging. As illustrated in Fig. 1, even advanced models such as GPT-4o","cbCairqfubEkB4xe","https://ap.wps.com/l/cbCairqfubEkB4xe","pdf",515936,1,"English","en",105,"# I. Introduction\n## Motivation: limitations of existing MLLMs in aerial VQA\n## Problem framing: interface mismatch in perception-to-reasoning\n## Proposed approach: AeroRAG framework","[{\"question\":\"What problem does AeroRAG address in aerial visual question answering?\",\"answer\":\"AeroRAG addresses unreliable visual QA caused by small objects, counting, coarse localization, and relation reasoning that conventional MLLMs struggle with in aerial scenes.\"},{\"question\":\"How does AeroRAG differ from directly prompting multimodal LLMs with dense visual tokens?\",\"answer\":\"AeroRAG converts the image into scene-graph-based structured knowledge (categories, quantities, locations, relations) and then performs query-conditioned retrieval to build compact text prompts.\"},{\"question\":\"Which datasets and evaluation setups are used to demonstrate the effectiveness of AeroRAG?\",\"answer\":\"Experiments are conducted on the AUG aerial dataset and the VG-150 benchmark, and compatibility with standard visual reasoning settings is evaluated using VQAv2.\"}]",1784195960,20,{"code":4,"msg":29,"data":30},"ok",{"site_id":23,"language":22,"slug":31,"title":13,"keywords":32,"description":14,"schema_data":33,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":26},"aerorag-structured-multimodal-retrieval-augmented-llm-for-fine-grained-aerial-visual-reasoning","",{"@graph":34,"@context":84},[35,52,67],{"@type":36,"itemListElement":37},"BreadcrumbList",[38,42,46,49],{"item":39,"name":40,"@type":41,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":43,"name":44,"@type":41,"position":45},"https://docshare.wps.com/document/","Document",2,{"item":47,"name":12,"@type":41,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":41,"position":51},"https://docshare.wps.com/document/aerorag-structured-multimodal-retrieval-augmented-llm-for-fine-grained-aerial-visual-reasoning/84483/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":39,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does AeroRAG address in aerial visual question answering?","Question",{"text":74,"@type":75},"AeroRAG addresses unreliable visual QA caused by small objects, counting, coarse localization, and relation reasoning that conventional MLLMs struggle with in aerial scenes.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does AeroRAG differ from directly prompting multimodal LLMs with dense visual tokens?",{"text":79,"@type":75},"AeroRAG converts the image into scene-graph-based structured knowledge (categories, quantities, locations, relations) and then performs query-conditioned retrieval to build compact text prompts.",{"name":81,"@type":72,"acceptedAnswer":82},"Which datasets and evaluation setups are used to demonstrate the effectiveness of AeroRAG?",{"text":83,"@type":75},"Experiments are conducted on the AUG aerial dataset and the VG-150 benchmark, and compatibility with standard visual reasoning settings is evaluated using VQAv2.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":44,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":45,"doc_module":4,"doc_module_name":44,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":44,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":44,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":44,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":44,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":44,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":44,"category_name":124,"show_sort_weight":27,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":27,"doc_module":4,"doc_module_name":44,"category_name":127,"show_sort_weight":27,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":44,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":44,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]