[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86415-en":3,"doc-seo-86415-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86415,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","The Injection Paradox Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection","A reproducible failure mode in RAG-based LLM recommendation systems is introduced as the “Injection Paradox,” where prompt injections embedded in retrieved documents backfire on the attacker by suppressing the target brand below the injection-free baseline. In safety-trained Claude models, injected documents cause sharp drops in recommendation rates, with suppression propagating to other unmodified documents of the same brand. In Claude Opus 4.6, the target brand falls from a 54% baseline to zero top-2 recommendations across 50 trials despite injections in only 1 of 4 brand documents. Contrasting GPT behavior shows increased recommendations under the same injection, pointing to model-family differences. Code, prompts, privacy-preserving per-trial outcomes, and aggregated results are released, enabling replication and mitigation research.","The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection  \nHyunseok Paeng 1  \narXiv :2606 .09204v2 [ cs .LG] 11 Jul 2026  \nAbstract  \nWe present a reproducible failure mode of safety training in RAG-based LLM recommendation, the Injection Paradox, in which prompt injections embedded in retrieved documents backfire against the attacker, suppressing the target brand below the injectionfree baseline. In safety-trained Claude models, documents containing prompt injections suffer a sharp drop in recommendation rate, and this suppression propagates beyond the injected document to unmodified documents of the same brand. In Claude Opus 4 .6, the target brand drops from a 54% baseline to zero top-2 recommendations across all 50 trials, even though only 1 of 4 brand documents in the corpus contains an injection. The directional pattern is reproduced in counterfactual experiments and across three brands. A contrasting result across the GPT models tested, where the same injection instead increases recommendations, suggests model-family differences in how injectionlike context affects recommendation behavior. These findings raise the technical possibility of a reverse-attack scenario in which an adversary embeds injections in a competitor’s documents to suppress the competitor’s brand via safety-sensitive model behavior. Code, prompts, privacy-preserving pertrial outcome records, and aggregate results are released.1  \n1 Independent Researcher. Correspondence to: Hyunseok Paeng \u003C[peanghs@naver.com](peanghs@naver.com) >.  \nAccepted at the ICML 2026 Workshop on Failure Modes in Agentic AI (FAGEN) . Non-archival.  \n1[https://github.com/peanghs/injection-paradox](https://github.com/peanghs/injection-paradox). See [docs/METHODOLOGY.md](docs/METHODOLOGY.md) for the aggregation protocol.  \n1. Introduction  \nLarge language model (LLM)-based recommendation systems are rapidly growing (Nawara & Kashef, 2025) . ChatGPT Search uses retrieved web search results to generate recommendations. In February 2026, Microsoft Security reported that 31 companies across 14 industries had attempted to bias AI-assistant recommendations by planting hidden instructions in clickable links and URL prompt parameters (“AI recommendation poisoning”) (Microsoft Defender Security Research Team, 2026) . In response, LLM providers have adopted safety training (a post-training alignment process that teaches models to refuse harmful or manipulative inputs), most notably OpenAI’s Reinforcement Learning from Human Feedback (RLHF)(Ouyang et al. , 2022) and Anthropic’s Constitutional AI (CAI) (Bai et al. , 2022) . Defenses against direct prompt hacking (jailbreaks) attempted through the user-facing chat interface have advanced considerably, with Anthropic’s Constitutional Classifiers achieving over 95% blocking rates in automated evaluations (Sharma et al. , 2025) . However, the interaction between safety training and indirect prompt injection (malicious instructions embedded in external documents and injected via RAG context) remains insufficiently studied.  \nThe risks of indirect prompt injection have been systematized in the context of LLM-integrated applications (Greshake et al. , 2023); adversarial search engine optimization targeting LLM-driven selection has been demonstrated on production systems (Nestaaset al. , 2025); and cognitive-bias-based manipulation of LLM recommendations (Filandrianos et al. , 2025) and adversarial-text-based RAG poisoning (Zou et al. , 2025) have been explored independently. Yet these studies do not address the interaction between safety training and injection, and to our knowledge no prior work has reported a phenomenon in which safety training itself produces novel side effects beyond its intended defensive function.  \nAccordingly, this study investigates whether the injection defense afforded by safety training produces  \nunintended side effects. We examine whether models, in the process of de","cbCaiv2yw9YPT60G","https://ap.wps.com/l/cbCaiv2yw9YPT60G","pdf",229867,4,1,18,"English","en",105,"# Introduction\n## Research motivation\n## Experimental design and scope\n## Key findings","[{\"question\":\"What is the Injection Paradox described in the document?\",\"answer\":\"The Injection Paradox is a reproducible failure mode where prompt injections embedded in retrieved documents backfire against the attacker, suppressing recommendations below the injection-free baseline. It is observed in safety-trained LLM recommenders using RAG context injection.\"},{\"question\":\"How do safety-trained Claude models respond when injected documents are retrieved?\",\"answer\":\"Injected documents lead to a sharp drop in recommendation rate, and the suppression propagates to other unmodified documents of the same brand. The document reports severe effects in Claude Sonnet and especially Claude Opus.\"},{\"question\":\"What contrasting behavior is reported for GPT models under the same injection?\",\"answer\":\"For the GPT models tested, the same prompt injection increases recommendations rather than suppressing them. This contrast is used to suggest differences between model families in how injection-like context affects recommendation behavior.\"}]",1784211595,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-injection-paradox-brand-level-suppression-in-safety-trained-llm-recommendations-via-rag-context-injection","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/the-injection-paradox-brand-level-suppression-in-safety-trained-llm-recommendations-via-rag-context-injection/86415/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the Injection Paradox described in the document?","Question",{"text":75,"@type":76},"The Injection Paradox is a reproducible failure mode where prompt injections embedded in retrieved documents backfire against the attacker, suppressing recommendations below the injection-free baseline. It is observed in safety-trained LLM recommenders using RAG context injection.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do safety-trained Claude models respond when injected documents are retrieved?",{"text":80,"@type":76},"Injected documents lead to a sharp drop in recommendation rate, and the suppression propagates to other unmodified documents of the same brand. The document reports severe effects in Claude Sonnet and especially Claude Opus.",{"name":82,"@type":73,"acceptedAnswer":83},"What contrasting behavior is reported for GPT models under the same injection?",{"text":84,"@type":76},"For the GPT models tested, the same prompt injection increases recommendations rather than suppressing them. This contrast is used to suggest differences between model families in how injection-like context affects recommendation behavior.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]