[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81708-en":3,"doc-seo-81708-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81708,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","PRA-RAG Provably Robust Aggregation in Retrieval-Augmented Generation against Retrieval Corruption","Retrieval-Augmented Generation (RAG) improves Large Language Models by adding external knowledge, yet it stays exposed to poisoning attacks that manipulate retrieved texts and misdirect model outputs. PRA-RAG introduces a provably robust retrieval aggregation method: it samples many combinations of retrieved passages and exploits geometric structures in the embedding space to select a stable, robust subset. The approach provides theoretical bounds on the maximum impact of poisoned retrievals and a quantitative robustness measure. Experiments show attack success reductions to as low as 1% while preserving strong accuracy of 71%.","PRA-RAG: Provably Robust Aggregation in Retrieval-Augmented Generation against Retrieval Corruption  \nXue Tan1,3 , Yi Zheng1 , Chang Huo1 , Yunruo Zhang3 , Yu Liu1,3 , Hao Luan1,3 , Zhuyang Yu1,3 , Xiaoyan Sun2B , Ping Chen3B , Jun Dai2B  \n1 School of Computer Science, Fudan University, Shanghai, China,  \n2Department of Computer Science, Worcester Polytechnic Institute, MA, USA, 3Institute of Big Data, Fudan University, Shanghai, China, Corresponding authors: [pchen@fudan.edu.cn](pchen@fudan.edu.cn), [xsun7@wpi.edu](xsun7@wpi.edu), [jdai@wpi.edu](jdai@wpi.edu)  \narXiv :2607 .000 12v 1 [ cs .IR] 8 May 2026  \nAbstract  \nRetrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by incorporating external knowledge, effectively mitigating their inherent knowledge limitations.  \nHowever, RAG remains vulnerable to poisoning attacks that manipulate retrieved texts to mislead model outputs. Existing defense mechanisms often lack theoretical robustness guarantees and perform unreliably when the LLM has limited knowledge of the retrieved content.  \nIn this work, we propose PRA-RAG, a provably robust retrieval aggregation algorithm designed to defend against poisoning attacks on retrieved texts. PRA-RAG samples multiple combinations of retrieved texts and utilizes geometric structures in the embedding space to identify a robust subset, from which a stable aggregated representation is derived. We provide theoretical bounds on the maximum impact of poisoned retrieved content and establish a quantitative measure of RAG’s robustness.  \nExperiments across multiple benchmarks and RAG architectures demonstrate that PRA-RAG reduces the attack success rate to as low as 1% while maintaining an accuracy of 71%, significantly outperforming representative state-ofthe-art (SOTA) methods.  \n1 Introduction  \nRetrieval-Augmented Generation (RAG) (Lewis et al., 2020) is an advanced generative paradigm that integrates external knowledge databases to effectively address the limitations of LLMs indomain-specific knowledge coverage and access to up-to-date information. In a typical RAG pipeline, when a user submits a query (e.g., “What is the name of the highest mountain?\"), the retriever first encodes it into an embedding vector using a text encoder (e.g., BERT (Kenton and Toutanova, 2019)) and then retrieves the most similar texts from an external knowledge database. These retrieved texts are subsequently provided as context to the LLM to  \nguide and enhance the response generation process. RAG has been widely adopted in various real-world applications due to its strengths in knowledge augmentation and high-quality generation, with prominent examples including ChatGPT (Achiam et al., 2023), Microsoft Bing Chat (Microsoft, 2024), and Google Search AI (Google, 2024) . However, the integration of external knowledge databases further exacerbates concerns regarding the security of LLMs (rocky, 2024 ; BBC, 2024) .  \nRecent studies have shown that injecting malicious content into retrieved texts can steer the LLM to generate responses aligned with the attacker’s intent (e.g., the target answer could be “Fuji\" when the target question is “What is the name of the highest mountain?\"), thereby posing a serious threat to the reliability and security of RAG systems (Zou et al., 2024 ; Greshake et al., 2023 ; Tan et al., 2024) . To counter attacks induced by poisoned texts, AstuteRAG (Wang et al., 2025) and TrustRAG (Zhou et al., 2025) introduce detection-based defenses that leverage the internal knowledge of LLMs to identify and filter maliciously retrieved content. However, their effectiveness is limited in scenarios where the LLM lacks sufficient knowledge to recognize the poisoned inputs. Moreover, existing methods (Wang et al., 2025 ; Zhou et al., 2025 ; Wei et al., 2024 ; Asai et al., 2024) lack a theoretical framework for certifying or quantifying the robustness of RAG systems, leaving their reliability under adversarial conditions largely unverifi","cbCaijvA0zsvU3ZG","https://ap.wps.com/l/cbCaijvA0zsvU3ZG","pdf",4148049,3,1,16,"English","en",105,"# Introduction\n## Retrieval-Augmented Generation and Threat Model\n## Proposed PRA-RAG Approach\n## Key Contributions\n## Evaluation","[{\"question\":\"What problem does PRA-RAG address in RAG systems?\",\"answer\":\"PRA-RAG targets poisoning attacks that manipulate retrieved texts to steer LLM outputs. It aims to reduce the influence of malicious retrieved content during the retrieval stage.\"},{\"question\":\"How does PRA-RAG aggregate retrieved texts to improve robustness?\",\"answer\":\"PRA-RAG samples diverse combinations of retrieved texts, embeds each combination into a geometric space, then selects a minimum-radius ball enclosing more than half the combinations. It uses a weighted average of the selected subset’s embeddings to form a robust representation.\"},{\"question\":\"What theoretical guarantees does PRA-RAG provide?\",\"answer\":\"PRA-RAG models semantic shifts caused by poisoned retrievals and proves an upper bound on that shift. This bound offers a quantitative basis for constraining poisoning impact and measuring robustness.\"}]",1784175537,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"pra-rag-provably-robust-aggregation-in-retrieval-augmented-generation-against-retrieval-corruption","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/pra-rag-provably-robust-aggregation-in-retrieval-augmented-generation-against-retrieval-corruption/81708/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does PRA-RAG address in RAG systems?","Question",{"text":75,"@type":76},"PRA-RAG targets poisoning attacks that manipulate retrieved texts to steer LLM outputs. It aims to reduce the influence of malicious retrieved content during the retrieval stage.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does PRA-RAG aggregate retrieved texts to improve robustness?",{"text":80,"@type":76},"PRA-RAG samples diverse combinations of retrieved texts, embeds each combination into a geometric space, then selects a minimum-radius ball enclosing more than half the combinations. It uses a weighted average of the selected subset’s embeddings to form a robust representation.",{"name":82,"@type":73,"acceptedAnswer":83},"What theoretical guarantees does PRA-RAG provide?",{"text":84,"@type":76},"PRA-RAG models semantic shifts caused by poisoned retrievals and proves an upper bound on that shift. This bound offers a quantitative basis for constraining poisoning impact and measuring robustness.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]