[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85808-en":3,"doc-seo-85808-105":28,"detail-sidebar-cat-0-en-105":89},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":11},85808,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","The Effect of Multi-Lingual and Keyword Adversarial Injection on LLM Relevance Judgment","Large language models (LLMs) are increasingly used as automated judges for relevance evaluation in information retrieval, yet their robustness to adversarial manipulation is not well understood, especially in multilingual contexts. This study analyzes cross-lingual prompt injection attacks against LLM-based relevance judgments using TREC Deep Learning collections and two open-weight models. It tests instruction- and content-based injections across eight languages, finding multilingual query-based attacks inflate relevance scores and evade existing defenses. The results show defenses can be adapted to mitigate attacks, but attacks remain easily bypassable, exposing a critical gap in proactive evaluation for LLM-as-a-judge systems.","The Eﬀect of Multi-Lingual and Keyword Adversarial Injection  \non LLM Relevance Judgment  \nNguyen Khoi Vo  \nRMIT University Melbourne, VIC, Australia  \nMark Sanderson  \nRMIT University Melbourne, VIC, Australia  \nTuong Duy Duong  \nRMIT University Melbourne, VIC, Australia  \nOleg Zendel  \nRMIT University Melbourne, VIC, Australia  \narXiv :2607 . 10080v 1 [ cs .IR] 11 Jul 2026  \nAbstract  \nLarge language models (LLMs) are increasingly being used as automated judges for relevance evaluation in information retrieval, yet their robustness to adversarial manipulation remains insuﬃciently understood, particularly in multilingual settings. In this work, we investigate the impact of cross-lingual prompt injection attacks on LLM-based relevance judgments using TREC Deep Learning collections and two open-weight models under established prompting frameworks. We examine both instruction-based and contentbased injection strategies in 8languages spanning diﬀerent resource levels. Our results demonstrate that multilingual query-based injections are highly eﬀective in inﬂating relevance scores while simultaneously evading existing prompt-injection defenses. We further found that, although existing defense mechanisms can be modiﬁed to mitigate such attacks, these injections can be easily adapted to bypass them. These ﬁndings highlight a critical gap in current defense approaches and demonstrate that language generalization can act as an attack vector, underscoring the need for more robust and proactive evaluation frameworks for LLM-as-a-judge systems.  \nKeywords  \nlarge language models, information retrieval, adversarial prompting, relevance judgment, multilingual evaluation  \n1 Introduction  \nRecent advances in LLMs have positioned them as potential alternatives to traditional human-based relevance judgments, leading to their increasing adoption as automated judges in information retrieval (IR) tasks [4, 10] . As LLM judges are deployed in evaluation pipelines—including TREC-style benchmarks and commercial search quality assessment; their susceptibility to adversarial manipulation carries potential consequences: inﬂated relevance scores can distort evaluation outcomes, misguide retrieval system development, and undermine the integrity oflarge-scale automated annotation.  \nPrior work on adversarial manipulation of LLM-based evaluation has been limited. Most studies focus on instruction-based  \nThis work is licensed under a Creative Commons Attribution 4 .0 International License.  \nVulGen’26, Melbourne, VIC, Australia  \n© 2026 Copyright held by the owner/author(s) .  \nprompt injection (e.g., “ignore previous instructions”) [7] . In contrast, only a small number of works have examined content-based manipulation. In particular, Alaoﬁ et al. [1] shows that inserting query keywords into passages can fool LLM-based relevance judgments, while Cuconasu et al. [3] demonstrates that introducing distracting content can signiﬁcantly degrade retrieval-augmented generation (RAG) performance. However, these two lines of work remain largely disconnected: the ﬁrst focuses on keyword-based injection, while the second focuses on performance degradation in RAG systems, without evaluating adversarial implications forLLMas-a-judge settings. Furthermore, limited attention has been given to richer forms of content manipulation, such as query variants or semantically similar but irrelevant text, and to their transferability across languages. In particular, content-based injection strategies – such as the inclusion of query keywords, phrases, or their variants – resemble keyword stuﬃng and black-hat SEO practices, yet their robustness and generality remain insuﬃciently understood. This gap is particularly concerning given recent ﬁndings by Thomas et al. [9], which suggest that LLM-based evaluation can make relevance judgments across languages. This raises the question of whether content-based manipulations can also transfer across languages and remain eﬀective under mul","cbCaic13fwSZFOHZ","https://ap.wps.com/l/cbCaic13fwSZFOHZ","pdf",120611,1,3,"English","en",105,"# Introduction\n# Methodology\n# Results\n# Discussion","[{\"question\":\"What problem does the paper address about LLMs in information retrieval evaluation?\",\"answer\":\"It studies how robust LLM-based judges are to adversarial manipulation when performing relevance judgments in information retrieval, with a focus on multilingual settings.\"},{\"question\":\"What kinds of adversarial injections are investigated?\",\"answer\":\"The paper examines both instruction-based and content-based prompt injection strategies, including multilingual query-based injections using query keywords and variants.\"},{\"question\":\"How do the proposed multilingual attacks affect relevance judgments and existing defenses?\",\"answer\":\"Multilingual query-based injections are highly effective at inflating relevance scores while evading existing prompt-injection defenses; even modified defenses can still be bypassed via easy adaptation of the injections.\"}]",1784206378,{"code":4,"msg":29,"data":30},"ok",{"site_id":24,"language":23,"slug":31,"title":13,"keywords":32,"description":14,"schema_data":33,"social_meta":84,"head_meta":86,"extra_data":88,"updated_unix":27},"the-effect-of-multi-lingual-and-keyword-adversarial-injection-on-llm-relevance-judgment","",{"@graph":34,"@context":83},[35,51,66],{"@type":36,"itemListElement":37},"BreadcrumbList",[38,42,46,48],{"item":39,"name":40,"@type":41,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":43,"name":44,"@type":41,"position":45},"https://docshare.wps.com/document/","Document",2,{"item":47,"name":12,"@type":41,"position":21},"https://docshare.wps.com/document/research-report/",{"item":49,"name":13,"@type":41,"position":50},"https://docshare.wps.com/document/the-effect-of-multi-lingual-and-keyword-adversarial-injection-on-llm-relevance-judgment/85808/",4,{"url":49,"name":13,"@type":52,"author":53,"headline":13,"publisher":55,"fileFormat":58,"inLanguage":23,"description":14,"dateModified":59,"datePublished":60,"encodingFormat":58,"isAccessibleForFree":61,"interactionStatistic":62},"DigitalDocument",{"name":9,"@type":54},"Person",{"url":39,"name":56,"@type":57},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":63,"interactionType":64,"userInteractionCount":20},"InteractionCounter",{"@type":65},"ViewAction",{"@type":67,"mainEntity":68},"FAQPage",[69,75,79],{"name":70,"@type":71,"acceptedAnswer":72},"What problem does the paper address about LLMs in information retrieval evaluation?","Question",{"text":73,"@type":74},"It studies how robust LLM-based judges are to adversarial manipulation when performing relevance judgments in information retrieval, with a focus on multilingual settings.","Answer",{"name":76,"@type":71,"acceptedAnswer":77},"What kinds of adversarial injections are investigated?",{"text":78,"@type":74},"The paper examines both instruction-based and content-based prompt injection strategies, including multilingual query-based injections using query keywords and variants.",{"name":80,"@type":71,"acceptedAnswer":81},"How do the proposed multilingual attacks affect relevance judgments and existing defenses?",{"text":82,"@type":74},"Multilingual query-based injections are highly effective at inflating relevance scores while evading existing prompt-injection defenses; even modified defenses can still be bypassed via easy adaptation of the injections.","https://schema.org",{"og:url":49,"og:type":85,"og:title":13,"og:site_name":56,"og:description":14},"article",{"robots":87,"canonical":49},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":90},[91,95,99,103,108,113,118,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":44,"category_name":92,"show_sort_weight":93,"slug":94},"Story & Novel",90,"story-novel",{"id":45,"doc_module":4,"doc_module_name":44,"category_name":96,"show_sort_weight":97,"slug":98},"Literature",80,"literature",{"id":50,"doc_module":4,"doc_module_name":44,"category_name":100,"show_sort_weight":101,"slug":102},"Exam",70,"exam",{"id":104,"doc_module":4,"doc_module_name":44,"category_name":105,"show_sort_weight":106,"slug":107},5,"Comic",60,"comic",{"id":109,"doc_module":4,"doc_module_name":44,"category_name":110,"show_sort_weight":111,"slug":112},6,"Technology",50,"technology",{"id":114,"doc_module":4,"doc_module_name":44,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":44,"category_name":12,"show_sort_weight":119,"slug":120},30,"research-report",{"id":122,"doc_module":4,"doc_module_name":44,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":44,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":44,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":44,"category_name":135,"show_sort_weight":104,"slug":136},19,"General","general"]