[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84599-en":3,"doc-seo-84599-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84599,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","MetaHOPE: Metaphor Translation Evaluation Framework for Analyzing MT and LLM Translation Errors","MetaHOPE is an error severity-aware annotation framework designed to evaluate metaphor translations and to diagnose translation errors made by neural machine translation systems and open-source large language models. The paper targets the difficulties metaphors pose for NLP due to semantic complexity, contextual dependence, and cultural embedding. Three representative systems—GoogleMT, GPT-5.4, and Hunyuan-7b—are compared on English-to-Chinese and Chinese-to-English using VUAMC and PSUCMC, with MetaHOPE-based error annotation and human post-edited gold references released publicly.","MetaHOPE: Metaphor Translation Evaluation Framework Investigating Open-Source LLMs and State-of-the-Art Neural  \nTranslation Models  \nJiahui Liang 1 , Lifeng Han2 ,3  \n1 Centre for Linguistics, Humanities, Leiden University, NL  \n2 LIACS, Leiden University, NL  \n3 BDS, Leiden University Medical Centre, NL  \n[j.h.l.jiahui@hum.leidenuniv.nl | l.han@lumc.nl](j.h.l.jiahui@hum.leidenuniv.nl | l.han@lumc.nl)  \narXiv :2607 .00848v 1 [ cs .CL] 1 Jul 2026  \n摘要  \nIn this opinion paper, we propose MetaHOPE, an error severity-aware annotation framework for evaluating metaphor translations. Metaphors present challenges for machine translation (MT) and natural language understanding and processing (NLU, NLP), because it presents the features of semantic complexity, contextual dependency, and cultural embeddings that can lead to ambiguity issues for NLP models. To investigate how state-of-the-art NLP models perform on translating metaphors, we select three representative systems, i.e. , GoogleMT, GPT5.4, and Hunyuan-7b as Neural MT (NMT) models and LLMs. We used two human-annotated metaphor corpora, including VUAMC and PSUCMC for English-to-Chinese and Chinese-to-English translation purposes. The original corpora we used are monolingual, where we carried out error annotation using the MetaHOPE framework, and also produced the human post-edited gold reference for bilingual use as a new resource. We believe the MetaHOPE evaluation framework for metaphor translation annotation, the parallel corpora resources, and the error analysis on SOTA automatic translation models can be useful and shed some light for the field of metaphor translation study. We share our resources publicly upon paper acceptance.  \n1 Introduction  \nMetaphors are pervasive in everyday discourse and serve as an essential cognitive tool, enabling people to understand and communicate abstract, complex, and unfamiliar concepts through more concrete and familiar experiences. For example, economic indicators may“soar”or“plummet”, governments may “fight”inflation, and negotiations may “reach a dead end”. These expressions draw on concrete experiences of movement, conflict, and space to convey meanings that extend beyond literal language (Lakoff and Johnson, 1980; Smedinga et al., 2023) . Beyond their semantic complexity, metaphors are also culturally embedded, and their interpretation often requires contextual awareness, sociocultural knowledge, and conceptual reasoning. As a result, they pose challenges for both machine translation (MT) and broader natural language understanding and processing (NLU, NLP) tasks.  \nRecent advances in neural MT (NMT) and large language models (LLMs) have substantially improved translation quality, with some systems achieving performance comparable to human translators on general translation benchmarks (Kocmiet al., 2025) . However, such improvements do not necessarily extend to metaphor translation (Han et al., 2026) . Karakanta et al. (2025) report metaphor translation accuracy rates of only 64-80%, while Wang et al. (2024) find that around 20% of metaphorical expressions remain non-equivalent in translation. A major source of error is overly literal translation, particularly for multi-word expressions (MWEs) such as idioms  \nand collocations, where models often fail to capture the intended figurative meaning (Bhatia et al. , 2023, 2024; Han et al., 2024) . Therefore, to better understand the gap between general MT performance and metaphor translation performance, it is necessary to systematically analyze metaphor translation errors. In addition, existing studies mainly focus on translation strategies (Pedersen, 2017; Zajdel, 2022; Li and Chen, 2025) or translation quality on equivalence, fluency, emotional effect, and authenticity (Wang et al., 2024) . However, finegrained error analysis remains limited. Karakanta et al. (2025) classify issues into meaning, form, andomission, but this framework is relatively coarsegrained and does not address severi","cbCaic3vd5J80qWZ","https://ap.wps.com/l/cbCaic3vd5J80qWZ","pdf",563338,1,18,"English","en",105,"# Introduction\n# Background and Related Work\n## Metaphors and Translation\n## Translation Evaluation Frameworks","[{\"question\":\"What is MetaHOPE, and what problem does it address in metaphor translation evaluation?\",\"answer\":\"MetaHOPE is an error severity-aware annotation framework for metaphor translation. It targets the need for fine-grained analysis of metaphor-specific translation errors beyond overall translation quality.\"},{\"question\":\"Which MT/LLM systems and translation directions are evaluated in the study?\",\"answer\":\"The study evaluates GoogleMT, GPT-5.4, and Hunyuan-7b. It analyzes errors for English-to-Chinese (EN-ZH) and Chinese-to-English (ZH-EN) translation.\"},{\"question\":\"How are errors annotated and what resources are produced for analysis?\",\"answer\":\"The framework adapts HOPE into five metaphor-oriented error categories with a five-level severity scale. Using VUAMC and PSUCMC, the authors generate MetaHOPE error annotations and create human post-edited gold references for bilingual use.\"}]",1784197021,45,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"metahope-metaphor-translation-evaluation-framework-for-analyzing-mt-and-llm-translation-errors","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/metahope-metaphor-translation-evaluation-framework-for-analyzing-mt-and-llm-translation-errors/84599/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is MetaHOPE, and what problem does it address in metaphor translation evaluation?","Question",{"text":75,"@type":76},"MetaHOPE is an error severity-aware annotation framework for metaphor translation. It targets the need for fine-grained analysis of metaphor-specific translation errors beyond overall translation quality.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which MT/LLM systems and translation directions are evaluated in the study?",{"text":80,"@type":76},"The study evaluates GoogleMT, GPT-5.4, and Hunyuan-7b. It analyzes errors for English-to-Chinese (EN-ZH) and Chinese-to-English (ZH-EN) translation.",{"name":82,"@type":73,"acceptedAnswer":83},"How are errors annotated and what resources are produced for analysis?",{"text":84,"@type":76},"The framework adapts HOPE into five metaphor-oriented error categories with a five-level severity scale. Using VUAMC and PSUCMC, the authors generate MetaHOPE error annotations and create human post-edited gold references for bilingual use.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]