[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83214-en":3,"doc-seo-83214-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83214,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Evaluating RAG Metrics in Applied Contexts An Experiment Its Findings and Its Limitations","Empirical evaluation of relevance-focused RAG metrics is conducted using a French business-domain question answering dataset built by human annotators. A RAG system’s generated responses and retrieved spans are scored with metrics from four libraries: Ragas, DeepEval, RAGChecker, and Opik, then compared against human evaluator judgments and recall for retrieval quality. Correlations are analyzed, methodology limitations are discussed, and comparisons to prior literature motivate future research directions.","arXiv :2607 .07302v 1 [ cs .CL] 8 Jul 2026  \nEvaluating RAG Metrics in Applied Contexts: An Experiment, Its Findings and Its Limitations  \nQuentin Brabant  \nOrange Research, Lannion, France  \n[quentin. brabant@orange. com](quentin. brabant@orange. com)  \nAbstract  \nThis paper reports an empirical study evaluating the relevance of several RAG metrics. The experiment is based on a question-answering dataset created by human annotators from business data. The generated responses and retrieved spans of a RAG system are scored using evaluation metrics from four libraries (Ragas, DeepEval, RAGChecker, Opik) . These metrics are compared to scores given by two evaluators, as well as to standard metrics such as recall. An analysis of correlations is conducted. Finally, we highlight certain limitations of our methodology, compare it to those used in the literature, and suggest some avenues for future research. This paper is an English translation of a paper originally published in the French-speaking workshop EvalLLM 2026 (Brabant, 2026) .  \n1 Introduction  \nEvaluating and comparing RAG (Retrieval Augmented Generation) systems remains a challenging task today: even when a sufficiently large set of test questions with reference answers is available, automatically evaluating a system’s responses against these references is far from trivial. A popular approach is to use so-called LLMas-a-judge metrics to perform this evaluation. Although these metrics generally seem more relevant than classical metrics such as BLEU, it is difficult to know in advance what the relevance of a specific metric will be on the dataset under consideration, especially since evaluation criteria can vary (relevance, factuality, completeness of responses, etc.) . It is therefore useful, when evaluating a RAG system under development, to conduct an evaluation of available metrics in order to verify that they provide acceptable approximations of the criterion considered, and that they will thus enable reliable comparison of different iterations of the RAG system being developed. Generally, metrics are evaluated by measuring their correlation with scores given by humans.  \nThis article reports an experiment of this type. This experiment is based on a question-answering dataset created by annotators from business data. The responses  \nproduced and the documents retrieved by a RAG system are scored using RAG metrics from four libraries: Ragas 1 (Es et al. , 2024), DeepEval2 , RAGChecker3 (Ru et al. , 2024), and Opik4 . These metrics are compared to reference evaluations: human evaluations for assessing generated responses, and recall for evaluating retrieval.  \nNote that our objective is not to compare the different metrics evaluated, asthe study results are strongly dependent on the choices made during its design. The reported experiments rather aim to apply a methodology in order to test its advantages and limitations. Unfortunately, the question-answering dataset used cannot be made public; however, we share the raw scores and the code used for the statistical analyses5 .  \nThe article is organized as follows. Section 2 describes the application context and the data used, including manual annotation and evaluation processes. Section 3 describes the methodology applied to compare reference evaluations with evaluations produced by metrics from the tested libraries. Section 4 reports and analyzes the results of this experiment. Certain limitations of our methodology are highlighted in Section 5 . In Section 6, our methodology is compared to those employed in the literature. We finally propose some research perspectives aimed at facilitating the application of reliable methodologies for evaluating RAG metrics, in Section 7 .  \n2 Context and Data  \nOur company is developing a RAG solution designed to answer questions in French about a business domain. The developed system processes each given question via two key modules: first, a retriever, whose role is to retrieve releva","cbCaitEOYi1dK0fK","https://ap.wps.com/l/cbCaitEOYi1dK0fK","pdf",247947,2,1,13,"English","en",105,"# Introduction\n## Context and Data\n## The Question-Answering Dataset\n# Methodology\n# Results and Analysis\n# Limitations\n# Comparison With Literature\n# Future Research Directions","[{\"question\":\"What does the experiment evaluate in RAG systems?\",\"answer\":\"It evaluates the relevance of multiple RAG metrics applied to a question-answering task, scoring both generated responses and retrieved spans.\"},{\"question\":\"How is the dataset created for the study?\",\"answer\":\"A question-answering dataset is built from business documents and includes 96 questions, each linked to reference answers and one or more reference spans annotated by company employees.\"},{\"question\":\"How are metric outputs validated in the experiment?\",\"answer\":\"Metric scores are compared to reference evaluations from two evaluators for response quality and to recall for retrieval evaluation, followed by correlation analysis.\"}]",1784185995,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"evaluating-rag-metrics-in-applied-contexts-an-experiment-its-findings-and-its-limitations","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/evaluating-rag-metrics-in-applied-contexts-an-experiment-its-findings-and-its-limitations/83214/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does the experiment evaluate in RAG systems?","Question",{"text":75,"@type":76},"It evaluates the relevance of multiple RAG metrics applied to a question-answering task, scoring both generated responses and retrieved spans.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the dataset created for the study?",{"text":80,"@type":76},"A question-answering dataset is built from business documents and includes 96 questions, each linked to reference answers and one or more reference spans annotated by company employees.",{"name":82,"@type":73,"acceptedAnswer":83},"How are metric outputs validated in the experiment?",{"text":84,"@type":76},"Metric scores are compared to reference evaluations from two evaluators for response quality and to recall for retrieval evaluation, followed by correlation analysis.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]