[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121870-en":3,"doc-seo-121870-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121870,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Towards Explainable Evaluation Metrics for Machine Translation - Conference Paper","Unlike classical lexical overlap metrics such as BLEU, many contemporary machine translation evaluation metrics rely on black-box large language models (e.g., COMET, BERTScore). While they correlate well with human judgments, classical metrics still dominate, partly because their decision processes are more transparent. This concept paper defines properties and goals for explainable MT metrics, synthesizes recent techniques, and reviews generative-model-based explainability approaches such as ChatGPT and GPT-4, concluding with a vision for next-generation interpretable evaluation and explanations.","Towards Explainable Evaluation Metrics for Machine  \nTranslation  \nChristoph Leiter  \nNatural Language Learning Group University of Mannheim  \nB6 26, 68159 Mannheim, Germany  \nPiyawat Lertvittayakumjorn  \nImperial College London  \nMarina Fomicheva  \nUniversity of Sheffield  \nWei Zhao  \nUniversity of Aberdeen  \nHeidelberg Institute for Theoretical Studies  \nYang Gao  \nRoyal Holloway, University of London  \nSteffen Eger  \nUniversity of Mannheim  \n[christoph.leiter@uni-mannheim.de](christoph.leiter@uni-mannheim.de)  \n[pl1515@imperial.ac.uk](pl1515@imperial.ac.uk)  \n[m.fomicheva@sheffield.ac.uk](m.fomicheva@sheffield.ac.uk)[ ](m.fomicheva@sheffield.ac.uk)[wei.zhao@abdn.ac.uk](wei.zhao@abdn.ac.uk)  \n[gaostayyang@google.com](gaostayyang@google.com)  \n[steffen.eger@uni-mannheim.de](steffen.eger@uni-mannheim.de)  \nEditor: Ivan Titov  \nAbstract  \nUnlike classical lexical overlap metrics such as BLEU, most current evaluation metrics for machine translation (for example, COMET or BERTScore) are based on black-box large language models. They often achieve strong correlations with human judgments, but recent research indicates that the lower-quality classical metrics remain dominant, one of the potential reasons being that their decision processes are more transparent. To foster more widespread acceptance of novel high-quality metrics, explainability thus becomes crucial. In this concept paper, we identify key properties as well as key goals of explainable machine translation metrics and provide a comprehensive synthesis of recent techniques, relating them to our established goals and properties. In this context, we also discuss the latest state-of-the-art approaches to explainable metrics based on generative models such as ChatGPT and GPT4 . Finally, we contribute a vision of next-generation approaches, including natural language explanations. We hope that our work can help catalyze and guide future research on explainable evaluation metrics and, mediately, also contribute to better and more transparent machine translation systems.  \nKeywords: evaluation metrics, explainability, interpretability, machine translation, machine translation evaluation  \n1 Introduction  \nThe field of evaluation metrics for Natural Language Generation (NLG), especially machine translation (MT) is in a crisis (Marie et al., 2021) . Despite the development of multiple  \n©2024 Christoph Leiter, Piyawat Lertvittayakumjorn, Marina Fomicheva, Wei Zhao, Yang Gao and Steffen Eger.  \nLicense: CC-BY 4.0, see [https://creativecommons.org/licenses/by/4.0/](https://creativecommons.org/licenses/by/4.0/. Attribution)[. Attribution](https://creativecommons.org/licenses/by/4.0/. Attribution) requirements are provided  \nat [http://jmlr.org/papers/v25/22-0416.html](http://jmlr.org/papers/v25/22-0416.html).  \nLeiter, Lertvittayakumjorn, Fomicheva, Zhao, Gao and Eger  \nhigh-quality evaluation metrics in recent years (e.g., Zhao et al., 2019; Zhang et al., 2020a; Rei et al. , 2020; Sellam et al. , 2020; Yuan et al. , 2021; Rei et al. , 2023a; Kocmi and Federmann, 2023a), the Natural Language Processing (NLP) community appears hesitant to adopt them for assessing NLG systems (Marie et al., 2021; Gehrmann et al., 2023) . Empirical investigations of Marie et al. (2021) indicate that the majority of MT papers relies on surface-level evaluation metrics such as BLEU and METEOR (Papineni et al., 2002; Banerjee and Lavie, 2005), which were created two decades ago, a trend that may allegedly have worsened in recent times. These surface-level metrics cannot (even) measure semantic similarity of their inputs and are thus fundamentally flawed, particularly when it comes to assessing the quality of recent state-of-the-art MT systems (e.g., Peyrard, 2019; Freitag et al., 2022), raising concerns about the credibility of the scientific field. We argue that the potential reasons for this neglect of recent high-quality metrics include: (i) non-enforcement by reviewers; (ii) easier comparison to previ","cbCaiq6LCfBvN7B4","https://ap.wps.com/l/cbCaiq6LCfBvN7B4","pdf",757079,1,49,"English","en",105,"# Introduction\n## Motivation for Explainable MT Evaluation Metrics\n## Benefits of Explainability in AI Evaluation\n## Role of Transparent Decision Processes\n## From Classical Metrics to Black-Box Models","[{\"question\":\"Why do classical metrics like BLEU still dominate machine translation evaluation research?\",\"answer\":\"The document attributes this to factors such as reviewer non-enforcement, easier comparison to prior work, computational inefficiency for newer metrics, and limited trust and transparency in high-quality black-box metrics.\"},{\"question\":\"What does the paper aim to provide regarding explainable evaluation metrics?\",\"answer\":\"It identifies key properties and goals of explainable machine translation metrics and synthesizes recent techniques, connecting them to the established goals and properties.\"},{\"question\":\"How does the paper relate explainability to building trust in evaluation metrics?\",\"answer\":\"It argues that explanations that align with human reasoning and faithfully reflect a metric’s internal decision process can increase acceptance and trust within the research community.\"}]","Towards Explainable Evaluation Metrics for Machine Translation - Conference Paper | PDF",1785807350,123,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"towards-explainable-evaluation-metrics-for-machine-translation-conference-paper","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/towards-explainable-evaluation-metrics-for-machine-translation-conference-paper/121870/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-04",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do classical metrics like BLEU still dominate machine translation evaluation research?","Question",{"text":76,"@type":77},"The document attributes this to factors such as reviewer non-enforcement, easier comparison to prior work, computational inefficiency for newer metrics, and limited trust and transparency in high-quality black-box metrics.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What does the paper aim to provide regarding explainable evaluation metrics?",{"text":81,"@type":77},"It identifies key properties and goals of explainable machine translation metrics and synthesizes recent techniques, connecting them to the established goals and properties.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the paper relate explainability to building trust in evaluation metrics?",{"text":85,"@type":77},"It argues that explanations that align with human reasoning and faithfully reflect a metric’s internal decision process can increase acceptance and trust within the research community.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]