[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85236-en":3,"doc-seo-85236-105":30,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85236,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for Grounded RAG","LLM-as-a-judge evaluation is widely used for retrieval-augmented generation (RAG), yet reusing the same model family as both generator and judge makes self-leniency hard to diagnose. Eval-Pair Matrix presents a controlled metaevaluation protocol for source-grounded RAG. It induces hidden answer-causal contradictions from GaRAGe grounding, generates answers from perturbed passages using GPT, Grok, and Gemini, and has the same models judge against original evidence. Results on validated records show minimal same-model F1 and recall differences, with a limited robust paired gap related to answers avoiding an induced claim; targeted human review finds no genuine false alarms.","Eval-Pair Matrix: Answer-Paired Meta-Evaluation of LLM Judges for  \nGrounded RAG  \nSriram Selvam  \n[selvamsriram@gmail.com](selvamsriram@gmail.com)  \nAnneswa Ghosh  \n[anneswaghosh@gmail.com](anneswaghosh@gmail.com)  \narXiv :2607 . 10626v 1 [ cs .CL] 12 Jul 2026  \nAbstract  \nLLM-as-a-judge evaluation is widely used for retrieval-augmented generation (RAG), but reusing the same model family as both generator and judge makes self-leniency difficult to identify. We introduce Eval-Pair Matrix, a controlled metaevaluation protocol for source-grounded RAG. Starting from GaRAGe questions and grounding passages, we induce one hidden answer-causal contradiction per record, generate answers from perturbed passages with GPT, Grok, and Gemini models, and then use the same models as blind judges to evaluate each answer against the original passages. The experiment contains 300 core records, 897 labeled generator outputs, and 2,683 judge verdicts in a crossed 3 × 3 matrix; the primary analysis uses 275 fully validated records. Instead of comparing diagonal and off-diagonal cells across different answers, we estimate same-model effects by pairing judges on the exact same candidate answer. This changes the interpretation: diagonal and offdiagonal F1 are similar, and the paired same-model recall effect is near zero (−0 .5 pp; 95% clusterbootstrap CI [−2 .7 , +1 .7]) . The only robust paired gap is lower matching-judge flagging for answers that avoided the induced claim (−4 .3 pp) . A targeted human evaluation finds that reviewed apparent false positives are alternate source-error detections, mistakes in labeling whether the induced claim was adopted, or unclear cases; none were adjudicated as genuine false alarms. The lesson is methodological: RAG judge studies should report full matrices, answer-paired effects, behavior strata, and label-task alignment.  \n1 Introduction  \nRetrieval-augmented generation (RAG) systems combine a language model’s parametric knowledge with external passages supplied at inference time (Lewis et al., 2020) . This makes them attractive for  \nknowledge-intensive applications, but it also creates a difficult evaluation problem: an answer maybe fluent, cited, and still contradict the reference passages. In practice, many RAG evaluations now use another LLM as the judge because LLM judges are cheap, fast, and can emit structured rationales (Liu et al., 2023 ; Zheng et al., 2023 ; Kim et al., 2023) . When the generator and judge are from the same model family, evaluation may become circular.  \nExisting LLM-as-a-judge work documents position, verbosity, length, style, familiarity, and selfpreference effects (Dubois et al., 2024 ; Shi et al., 2024 ; Liu et al., 2024 ; Wataoka et al., 2024 ; Ye et al., 2024 ; Spiliopoulou et al., 2025) . This literature also shows that self-preference is an identification problem: a judge may favor its own model’s output because it is better, clearer, more familiar, or more stylistically aligned, rather than because the judge is lenient toward its own model. We ask how this problem changes in source-grounded RAG. If we induce a known source-relative contradiction hidden from the judge, then the judge is not choosing a favorite answer; it is checking whether a candidate answer contradicts the original evidence.  \nWe build the task from GaRAGe, a benchmark with human-written answers and passage-level grounding annotations (Sorodoc et al., 2025) . For each selected example, an LLM perturber chooses one answer-causal target proposition and rewrites every passage that states or entails that claim so the evidence consistently supports a plausible but false replacement value. Generators answer the original question using the perturbed grounding. Judges then evaluate the answer against the unmodified grounding only. Figure 1 summarizes the end-toend design.  \nA raw diagonal analysis compares each judge’s matching-model cell with that judge’s two offdiagonal generator cells. That contrast is useful  \n","cbCaiknWsjLuGQkD","https://ap.wps.com/l/cbCaiknWsjLuGQkD","pdf",959023,4,1,15,"English","en",105,"# Introduction\n# Eval-Pair Matrix pipeline\n# Experimental setup and matrix design\n# Paired vs diagonal/off-diagonal analysis\n# Results and human evaluation\n# Methodological lessons","[{\"question\":\"What problem does Eval-Pair Matrix address in LLM-as-a-judge RAG evaluation?\",\"answer\":\"It addresses the difficulty of detecting self-leniency when the same model family is used for both the generator and the judge, where the judge may favor outputs that are more familiar or aligned.\"},{\"question\":\"How does Eval-Pair Matrix induce and control evaluation contradictions?\",\"answer\":\"Starting from GaRAGe questions and grounded passages, it induces one hidden answer-causal contradiction per record by perturbing passages so the evidence supports a plausible but false replacement, then generates answers from these perturbed passages.\"},{\"question\":\"What does the answer-paired analysis change compared with diagonal vs off-diagonal matrix comparisons?\",\"answer\":\"Instead of comparing different answers across diagonal and off-diagonal cells, it pairs judges on the exact same candidate answer to estimate same-model effects, which alters the interpretation of the matrix outcomes.\"},{\"question\":\"What do the experimental and human evaluation findings indicate about false alarms?\",\"answer\":\"Automatic metrics show small same-model effects for error-detection recall, and targeted human review of apparent false positives concludes none were genuine false alarms; they correspond to source-error detection differences, labeling mistakes, or unclear cases.\"}]",1784201910,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":28},"eval-pair-matrix-answer-paired-meta-evaluation-of-llm-judges-for-grounded-rag","",{"@graph":36,"@context":89},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/eval-pair-matrix-answer-paired-meta-evaluation-of-llm-judges-for-grounded-rag/85236/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Eval-Pair Matrix address in LLM-as-a-judge RAG evaluation?","Question",{"text":75,"@type":76},"It addresses the difficulty of detecting self-leniency when the same model family is used for both the generator and the judge, where the judge may favor outputs that are more familiar or aligned.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Eval-Pair Matrix induce and control evaluation contradictions?",{"text":80,"@type":76},"Starting from GaRAGe questions and grounded passages, it induces one hidden answer-causal contradiction per record by perturbing passages so the evidence supports a plausible but false replacement, then generates answers from these perturbed passages.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the answer-paired analysis change compared with diagonal vs off-diagonal matrix comparisons?",{"text":84,"@type":76},"Instead of comparing different answers across diagonal and off-diagonal cells, it pairs judges on the exact same candidate answer to estimate same-model effects, which alters the interpretation of the matrix outcomes.",{"name":86,"@type":73,"acceptedAnswer":87},"What do the experimental and human evaluation findings indicate about false alarms?",{"text":88,"@type":76},"Automatic metrics show small same-model effects for error-detection recall, and targeted human review of apparent false positives concludes none were genuine false alarms; they correspond to source-error detection differences, labeling mistakes, or unclear cases.","https://schema.org",{"og:url":52,"og:type":91,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":93,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":96},[97,101,105,109,114,119,124,127,132,135,139],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},5,"Comic",60,"comic",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},6,"Technology",50,"technology",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":125,"slug":126},30,"research-report",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":130,"slug":131},9,"Religion & Spirituality",20,"religion-spirituality",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":130,"slug":134},"World Cup","world-cup",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":136,"slug":138},10,"Lifestyle","lifestyle",{"id":140,"doc_module":4,"doc_module_name":46,"category_name":141,"show_sort_weight":110,"slug":142},19,"General","general"]