[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86038-en":3,"doc-seo-86038-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86038,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Trust Before Fusion QIMG-7 and Source-Aware Resolution for Polluted Multimodal RAG","Multimodal retrieval-augmented generation (RAG) is evaluated with clean evidence, but real retrieval often returns topically relevant yet unreliable material, including false text and misleading images caused by corrupted metadata, entity swaps, typographic overlays, semantic edits, adversarial patches, or style transfer. The paper presents QIMG-7, a controlled benchmark for multimodal retrieval pollution in multi-sentence factual QA, spanning four datasets, seven image-attack families, and 16 paired clean/polluted regimes. Across four generator/gate stacks, naive fusion proves brittle; source-aware trust resolution (SATR) selects among Parametric, Text-only, and Full-MM candidates using source reliability, achieving a best balanced score of 0.816.","Trust Before Fusion: QIMG-7 and Source-Aware Resolution for Polluted Multimodal RAG  \nSaadeldine Eletter Owais Aijaz Preslav Nakov  \nMohamed bin Zayed University of Artificial Intelligence {saadeldine.eletter,owais.aijaz,[preslav.nakov}@mbzuai.ac.ae](preslav.nakov}@mbzuai.ac.ae)  \narXiv :2607 . 10798v 1 [ cs .CL] 12 Jul 2026  \nAbstract  \nMultimodal retrieval-augmented generation (RAG) is often evaluated with clean evidence, yet real retrieval can return topically relevant but unreliable content: false text and misleading images from corrupted metadata, entity swaps, typographic overlays, semantic edits, adversarial patches, blends, or style transfer. We introduce QIMG-7, a controlled benchmark for multimodal retrieval pollution in multisentence factual QA, spanning four datasets, seven image-attack families, and 16 paired clean/polluted regimes, for 1,760 evaluation rows per method. Across four generator/ -gate stacks, naive multimodal fusion is brittle: in the main gpt-4o-mini stack, Full-MM support drops from 0.908 with clean text to 0.490 with polluted text, often making Parametric fallback safer than retrieval. We propose source-aware trust resolution (SATR), a training-free approach that compares Parametric, Text-only, and Full-MM candidate answers and selects among candidate answers or falls back based on source reliability. The Field-Selector variant achieves the best balanced score, 0.816, improving over Full-MM by 11.7 points and over the Cascaded Router by 2.7 points. Ablations show that, in this text-first setting, explicit text-reliability modeling is the dominant driver of these gains. Overall, in text-first factual QA with multimodal retrieval conflict, our results support selective trust rather than unconditional fusion.  \nArtifacts are available at [https://github.com/](https://github.com/)[ ](https://github.com/)SaadElDine/Trust_Before_Fusion.  \n1 Introduction  \nRetrieval-augmented generation (RAG) is widely used to ground large language models in external evidence, improving factuality for knowledgeintensive QA and generation (Lewis et al., 2020 ; Izacard and Grave, 2021) . However, relevant retrieved content is not necessarily reliable: misinformation in retrieval corpora can steer models toward  \nRetrieved evidence is topical but mixed  \nRetriever  \nFigure 1: Motivating failure case for multimodal retrieval pollution. The retrieved evidence is topically relevant but mixed: clean text supports the correct Tokyo Disneyland answer, while polluted text and image evidence suggest the false claim that it opened in 1979 near Paris. A naïve multimodal RAG system may fuse these conflicting signals and produce a confident wrong answer, motivating source-aware trust resolution.  \nconfident but false answers (Pan et al., 2023 ; Zenget al., 2025) . In long-form settings, this problem is amplified because an early factual mistake can propagate through a multi-paragraph answer.  \nThe multimodal setting makes the problem harder. Modern systems retrieve images alongside text, and multimodal models can use both. Images can help by providing grounding, but they can also mislead via false captions, typographic overlays, out-of-context crops, swapped entities, or visually edited evidence. While robustness research has focused mostly on text-only retrieval, and multimodal RAG benchmarks usually assume clean evidence, the combined problem of multimodal polluted retrieval remains underexplored (Yu et al., 2025 ; Hu et al., 2025 ; Mortaheb et al., 2025) .  \nFigure 1 illustrates the core failure mode: the retrieved multimodal evidence can be relevant but unreliable, and naïve fusion can amplify polluted evidence into a confident hallucination.  \nIn this paper, we study multimodal retrieval pollution for multi-sentence factual QA, making the following four contributions:  \n• Benchmark. We construct QIMG-7, a paired clean/polluted multimodal RAG benchmark for multi-sentence factual QA, spanning four datasets, seven image-attack famili","cbCaifI47i7UghPb","https://ap.wps.com/l/cbCaifI47i7UghPb","pdf",7007059,5,1,23,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What problem does the paper address in multimodal RAG systems?\",\"answer\":\"Retrieved multimodal evidence can be topically relevant but unreliable due to corrupted metadata, misleading images, and edited or adversarial content. Naive multimodal fusion may then amplify pollution into confident false answers in multi-sentence factual QA.\"},{\"question\":\"What is QIMG-7 and how is it constructed?\",\"answer\":\"QIMG-7 is a controlled paired clean/polluted multimodal retrieval benchmark for multi-sentence factual QA. It spans four datasets, seven image-attack families (metadata, pixel, and style perturbations), and 16 evaluation regimes with 1,760 rows per method.\"},{\"question\":\"How does SATR work and what advantage does it provide over full multimodal fusion?\",\"answer\":\"SATR is a training-free source-aware trust-resolution approach that compares Parametric, Text-only, and Full-MM candidate answers and chooses, composes, or falls back based on structured source reliability fields. The Field-Selector variant achieves the best balanced score (0.816) and improves over Full-MM by 11.7 points, indicating selective trust beats unconditional fusion.\"}]",1784208010,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"trust-before-fusion-qimg-7-and-source-aware-resolution-for-polluted-multimodal-rag","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/trust-before-fusion-qimg-7-and-source-aware-resolution-for-polluted-multimodal-rag/86038/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper address in multimodal RAG systems?","Question",{"text":76,"@type":77},"Retrieved multimodal evidence can be topically relevant but unreliable due to corrupted metadata, misleading images, and edited or adversarial content. Naive multimodal fusion may then amplify pollution into confident false answers in multi-sentence factual QA.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is QIMG-7 and how is it constructed?",{"text":81,"@type":77},"QIMG-7 is a controlled paired clean/polluted multimodal retrieval benchmark for multi-sentence factual QA. It spans four datasets, seven image-attack families (metadata, pixel, and style perturbations), and 16 evaluation regimes with 1,760 rows per method.",{"name":83,"@type":74,"acceptedAnswer":84},"How does SATR work and what advantage does it provide over full multimodal fusion?",{"text":85,"@type":77},"SATR is a training-free source-aware trust-resolution approach that compares Parametric, Text-only, and Full-MM candidate answers and chooses, composes, or falls back based on structured source reliability fields. The Field-Selector variant achieves the best balanced score (0.816) and improves over Full-MM by 11.7 points, indicating selective trust beats unconditional fusion.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]