[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-148668-105":59,"doc-detail-148668-en":130},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":123,"head_meta":125,"extra_data":127,"updated_unix":129},105,"en","does-more-retrieved-evidence-help-visual-retrieval-augmented-generation-with-diffusion-language-models-abstract-and-introduction","Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models - Abstract and Introduction","","Visual retrieval-augmented generation (RAG) often expands the retrieved evidence set to improve coverage, assuming all evidence should be passed to the generator. The work shows this assumption fails for diffusion language models: adding more retrieved pages increases answer-page recall but can reduce answer accuracy due to semantic conflict. A latent-source analysis attributes the mismatch to source-coherence loss in parallel denoising, with interference visible in first-step answer-block distributions. An Entropy-Based Candidate Filter (ECF) selectively admits evidence, improving accuracy across multiple DLMs and visual QA benchmarks.",{"@graph":69,"@context":122},[70,84,105],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/does-more-retrieved-evidence-help-visual-retrieval-augmented-generation-with-diffusion-language-models-abstract-and-introduction/148668/",{"url":83,"name":65,"@type":85,"image":86,"author":91,"headline":65,"publisher":94,"fileFormat":97,"inLanguage":63,"description":67,"dateModified":98,"datePublished":99,"encodingFormat":97,"isAccessibleForFree":100,"interactionStatistic":101},"DigitalDocument",{"url":87,"@type":88,"width":89,"height":90},"https://docshare.wps.com/thumbnails/does-more-retrieved-evidence-help-visual-retrieval-augmented-generation-with-diffusion-language-models-abstract-and-introduction/148668.png","ImageObject",300,407,{"name":92,"@type":93},"Connor ","Person",{"url":74,"name":95,"@type":96},"DocShare","Organization","application/pdf","2026-09-17","2026-08-26",true,{"@type":102,"interactionType":103,"userInteractionCount":34},"InteractionCounter",{"@type":104},"ViewAction",{"@type":106,"mainEntity":107},"FAQPage",[108,114,118],{"name":109,"@type":110,"acceptedAnswer":111},"Why does retrieving more evidence pages not always improve answers in visual DLM-RAG?","Question",{"text":112,"@type":113},"More pages increase answer-page recall, but unconditionally passing all retrieved pages can reduce accuracy because semantically conflicting visual sources interfere during parallel denoising.","Answer",{"name":115,"@type":110,"acceptedAnswer":116},"What does the latent-source analysis reveal about the failure mode?",{"text":117,"@type":113},"It explains the mismatch via source-coherence loss in parallel denoising, where positionwise proposals can combine incompatible visual sources into unsupported answers.",{"name":119,"@type":110,"acceptedAnswer":120},"How does ECF prevent harmful evidence admission while keeping retrieval coverage?",{"text":121,"@type":113},"ECF is a training-free evidence-admission framework that forms multi-granularity evidence units and uses blank-controlled block confidence and retrieval rank to decide which candidates enter the final context.","https://schema.org",{"og:url":83,"og:type":124,"og:title":65,"og:site_name":95,"og:description":67},"article",{"robots":126,"canonical":83},"index,follow",{"doc_id":128,"site_id":62},148668,1787782833,{"code":4,"msg":5,"data":131},{"doc_id":128,"user_id":132,"nickname":92,"user_avatar":133,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":34,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":139,"language":140,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":67,"update_tm":129,"read_time":36},687207022233,"https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d","Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with  \nDiffusion Language Models?  \nJiankun Wang 1 , Yisen Gao2 , Ziwei Zhang 1 ,  \nXingcheng Fu3 , Jiaxin Bai4 , Chen Gao5  \n1 School of Computer Science and Engineering, Beihang University  \n2Department of Computer Science and Engineering, The Hong Kong University of Science and Technology  \n3 School of Computer Science and Engineering, Guangxi Normal University  \n4Department of Computer Science, Hong Kong Baptist University  \n5 School of Computing, National University of Singapore  \narXiv :2608 .07006v 1 [ cs .CL] 7 Aug 2026  \nAbstract  \nVisual retrieval-augmented generation (RAG) commonly expands the retrieved evidence set to improve answer-page coverage, implicitly assuming that all available evidence should be passed to the generator. We show that this assumption does not hold for diffusion language models (DLMs): retrieving more pages increases answer-page recall, whereas unconditionally passing all retrieved pages to the generator often reduces answer accuracy, primarily because of semantic conflict. A latent-source analysis explains this mismatch through source-coherence loss in parallel denoising, where positionwise proposals can combine incompatible visual sources into unsupported answers. We further find that such interference is already visible in the first-step answer-block distribution, making it possible to assess evidence before decoding. To preserve retrieval coverage while limiting harmful visual exposure, we propose the Entropy-Based Candidate Filter (ECF), a training-free evidence-admission framework. To reduce irrelevant content within individual candidates, ECF constructs multi-granularity evidence units; to identify beneficial additional evidence, it uses blank-controlled block confidence and retrieval rank to determine whether and which candidate should enter the final context. Across three multimodal DLMs and five visual QA benchmarks, ECF improves answer accuracy by 2.62 percentage points on average over the strongest fixed top-k input and, with LLaDA2.0-Uni, by 2.37 percentage points on average over the best competing training-free result for each dataset. These results show that broader retrieval benefits visual DLM-RAG through selective evidence admission rather than unconditional evidence expansion.  \nCode is publicly available at [https://github.com/wjkuser/ECF](https://github.com/wjkuser/ECF).  \n1 Introduction  \nMost modern large language models generate text autoregressively and have achieved strong performance across question answering, reasoning and tool-use tasks (Xiao et al. 2026; Ma et al. 2025) . More recently, diffusion language models (DLMs) have attracted increasing attention for their bidirectional conditioning and parallel decoding capabilities (Austin et al. 2021; Hoogeboom et al. 2021; Sahoo et al. 2024; Nie et al. 2025; Gao et al. 2026) . Recent studies (Inclusion AI et al. 2026; You et al. 2025; Ye et al. 2026) have further  \nPreprint. Under review.  \nextended DLMs to multimodal generation, enabling visual inputs to guide the denoising process.  \nHowever, many knowledge-intensive visual questions cannot be answered reliably from model parameters alone, creating a need for external evidence. Retrieval-augmented generation (RAG) addresses this need by conditioning generation on retrieved content (Izacard and Grave 2021; Shi et al. 2024; Asai et al. 2024), and has been extended to visual tasks (Yu et al. 2025; Li et al. 2025) . Visual RAG is challenging because answer-bearing evidence is often localized within a visually dense page, while each retrieved page introducesa complete visual source containing additional structured content. To compensate for imperfect retrieval, existing visual RAG systems, which are predominantly built around autoregressive LLMs, often provide the generator with multiple top-ranked pages to improve answer-page coverage (Yu et al. 2025; Luo et al. 2026) . Recent work has improved the ","cbCaiaLlk8QhzwnU","https://ap.wps.com/l/cbCaiaLlk8QhzwnU","pdf",838466,16,"English","# Abstract\n## Problem and motivation\n## Key experiments and findings\n## Proposed method: ECF\n## Results overview\n# Introduction\n## Background on DLMs and visual retrieval-augmented generation\n## Challenges of visual RAG with parallel denoising\n## Central question and experimental setup\n## Main observations from recall and accuracy trends","[{\"question\":\"Why does retrieving more evidence pages not always improve answers in visual DLM-RAG?\",\"answer\":\"More pages increase answer-page recall, but unconditionally passing all retrieved pages can reduce accuracy because semantically conflicting visual sources interfere during parallel denoising.\"},{\"question\":\"What does the latent-source analysis reveal about the failure mode?\",\"answer\":\"It explains the mismatch via source-coherence loss in parallel denoising, where positionwise proposals can combine incompatible visual sources into unsupported answers.\"},{\"question\":\"How does ECF prevent harmful evidence admission while keeping retrieval coverage?\",\"answer\":\"ECF is a training-free evidence-admission framework that forms multi-granularity evidence units and uses blank-controlled block confidence and retrieval rank to decide which candidates enter the final context.\"}]","Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models - Abstract and Introduction | PDF"]