[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84828-en":3,"doc-seo-84828-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84828,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Modality Relevance is not Modality Utility Post-hoc Selective Modality Escalation for Cost-Aware Multimodal RAG","Multimodal retrieval-augmented generation (RAG) grounds generation in evidence from heterogeneous modalities such as text, tables, and images, yet real deployments face a cost asymmetry between cheap text/table processing and expensive vision–language model (VLM) image understanding. Prior adaptive routing decides modality or fidelity pre-retrieval, but modality relevance poorly predicts true answer utility. Oracle analysis on MultiModalQA shows many image-supported questions are solvable without images. The proposed post-hoc selective escalation answers first with text+tables, verifies and localizes missing evidence by modality, and invokes VLM only when value justifies visual cost, recovering always-on accuracy with fewer visual calls.","arXiv :2607 .05438v 1 [ cs .IR] 3 Jul 2026  \nModality Relevance is not Modality Utility: Post-hoc Selective Modality Escalation for Cost-Aware Multimodal  \nRAG  \nXue Li, Yiming Gai  \nHangzhou International Innovation Institute, Beihang University  \nHangzhou, China  \nAbstract  \nMultimodal retrieval-augmented generation (RAG) grounds a generator in evidence drawn from heterogeneous modalities—text, tables, and images. The dominant deployment choice is binary and made before the model has tried to answer: either run a cheap text(+table) pipeline, or pay for an expensive vision–language model (VLM) over every image. Recent adaptive systems improve on this by selecting the modality or fidelity pre-retrieval, from a question-conditioned predictor of which modality will be needed. We show that this is the wrong decision point. Through an oracle headroom analysis on MultiModalQA, we find that the relevance of a modality to a question is a weak predictor of whether that modality is actually needed to answer correctly: a large fraction of questions whose gold support includes an image are nonetheless answerable from text and tables alone, and a pre-retrieval router that escalates on apparent visual relevance over-escalates substantially relative to an oracle. We propose post-hoc selective modality escalation: answer cheaply from text and tables, run a verifier on the (query, draft answer, evidence) tuple that localizes which modality is missing, and pay for VLM evidence only there. A calibrated value-of-escalation router then decides whether the expected accuracy gain justifies the visual cost. On MultiModalQA, our router recovers the accuracy of an always-on VLM pipeline while issuing far fewer visual calls, and closes most of the gap to the oracle escalation rate. The result extends a routing-signal hierarchy established for retrieval depth and reasoning hops to a third axis—modality—under a single cost-aware selective-escalation view.  \n1 Introduction  \nRetrieval-augmented generation (RAG) grounds large language models in external evidence by retrieving content and conditioning generation on it. In the multimodal setting, that evidence is heterogeneous: a question may be answerable from a passage, from a table cell, from the content of an image, or from a combination. Production multimodal RAG systems face a basic cost asymmetry. Reading text and tables is cheap, but turning an image into something a language model can use requires a vision–language model (VLM), which is markedly more expensive in latency and compute. The default deployment pattern resolves this asymmetry rigidly: either restrict the pipeline to text(+table) and accept failures on visually grounded questions, or invoke a VLM over every candidate image and pay that cost on every query—including the majority that never needed it.  \nRecent adaptive multimodal RAG aims to spend the visual budget selectively. A prominent line decides which modality or fidelity to retrieve from a question-conditioned success predictor, before any evidence is read or any answer is drafted [Anonymous, 2026]. Agentic pipelines go further, planning modality-specific sub-queries and aggregating across retrieval branches [Anonymous, 2025a,d], but at the cost of many model calls per question. The first family is attractive because it is cheap, but it commits to the visual decision at the  \nearliest and least-informed point in the pipeline: before the system has observed any evidence or attempted an answer.  \nThis paper argues, and shows empirically, that the modality decision should be made after a cheap attempt, not before. Our central observation is a modality analogue of a relevance–utility gap that has been documented for retrieval scores in single-hop RAG: the relevance of a modality to a question is not the same as its utility for answering. A question can be about an entity that has an associated image and still be answerable entirely from the surrounding text and table evidence; conversel","cbCaief5yQcpYuyH","https://ap.wps.com/l/cbCaief5yQcpYuyH","pdf",197320,1,9,"English","en",105,"# Introduction\n## Cost asymmetry in multimodal RAG\n## Relevance–utility gap for modality\n## Post-hoc selective modality escalation","[{\"question\":\"Why does modality relevance fail to predict modality utility in multimodal RAG?\",\"answer\":\"The paper shows that a modality can be relevant to the question while still being unnecessary for correct answering. Oracle headroom analysis on MultiModalQA finds many questions whose gold support includes an image are answered correctly using only text and tables.\"},{\"question\":\"What is the proposed post-hoc selective modality escalation method?\",\"answer\":\"The system first generates a draft answer from cheap text+table evidence, then runs a verifier on the (query, draft answer, evidence) tuple to localize which modality’s information is missing. It escalates to VLM evidence only when the missing gap is localized to the image modality, then regenerates using compact textual sidecars.\"},{\"question\":\"How does the system decide whether to pay the VLM cost?\",\"answer\":\"A calibrated value-of-escalation router compares the expected accuracy gain from escalation against the visual cost. This yields a tunable operating point on the accuracy–cost frontier and reduces unnecessary visual calls.\"}]",1784198582,23,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"modality-relevance-is-not-modality-utility-post-hoc-selective-modality-escalation-for-cost-aware-multimodal-rag","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/modality-relevance-is-not-modality-utility-post-hoc-selective-modality-escalation-for-cost-aware-multimodal-rag/84828/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does modality relevance fail to predict modality utility in multimodal RAG?","Question",{"text":75,"@type":76},"The paper shows that a modality can be relevant to the question while still being unnecessary for correct answering. Oracle headroom analysis on MultiModalQA finds many questions whose gold support includes an image are answered correctly using only text and tables.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the proposed post-hoc selective modality escalation method?",{"text":80,"@type":76},"The system first generates a draft answer from cheap text+table evidence, then runs a verifier on the (query, draft answer, evidence) tuple to localize which modality’s information is missing. It escalates to VLM evidence only when the missing gap is localized to the image modality, then regenerates using compact textual sidecars.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the system decide whether to pay the VLM cost?",{"text":84,"@type":76},"A calibrated value-of-escalation router compares the expected accuracy gain from escalation against the visual cost. This yields a tunable operating point on the accuracy–cost frontier and reduces unnecessary visual calls.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]