[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83501-en":3,"doc-seo-83501-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83501,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","Large Language Models for Multi-Lingual Equivalent Mutant Detection: An Extended Empirical Study","Mutation testing is a widely used software quality technique, but equivalent mutants create unnecessary cost and bias, reducing practical effectiveness. This study investigates whether large language models (LLMs) can improve equivalent mutant detection (EMD), addressing limitations of pure code-analysis approaches and data-scarce machine learning methods. Using 3,302 Java and 1,088 C mutant pairs, experiments compare strategies, evaluate cross-lingual generalization, and benchmark against state-of-the-art baselines. LLM-based models yield higher F1-scores, with fine-tuned code embeddings providing the best accuracy and inference efficiency comparable to prior ML.","arXiv :2607 .005 1 1v 1 [ cs . SE] 1 Jul 2026  \nLarge Language Models for Multi-Lingual Equivalent Mutant Detection: An Extended Empirical Study  \nHONGLIN SHU∗ , Tianjin University, China and Kyushu University, Japan ZHAO TIAN∗ , College of Intelligence and Computing, Tianjin University, China DONG WANG†, College of Intelligence and Computing, Tianjin University, China JUNJI YU, College of Intelligence and Computing, Tianjin University, China JIAZHE ZHANG, College of Intelligence and Computing, Tianjin University, China XUEJIE CAO, College of Intelligence and Computing, Tianjin University, China JUNJIE CHEN, College of Intelligence and Computing, Tianjin University, China YASUTAKA KAMEI, Kyushu University, Japan  \nMutation testing is a powerful technique for ensuring software quality. However, the presence of equivalent mutants introduces unnecessary costs and biases, limiting its practical effectiveness. Although numerous equivalent mutant detection (EMD) methods have been proposed, they often face distinct challenges: pure-code analysis methods can be limited by their reliance on specific compiler infrastructures, while existing machinelearning approaches remain constrained by scarce training data and limited generalization to unseen mutants. Large language models (LLMs) have recently demonstrated remarkable performance across diverse coderelated tasks by better capturing program semantics. Yet their potential for EMD remains largely unexplored, particularly in the multi-lingual context. This paper presents the first comprehensive empirical study on LLMs for EMD, using 3,302 Java and 1,088 C mutant pairs to benchmark against state-of-the-art methods, explore strategy variations, assess efficiency, and evaluate cross-lingual generalization. Experimental results show that LLM-based approaches achieve higher F1-scores than the evaluated traditional methods, with fine-tuned code embedding yielding the highest detection accuracy among the tested strategies. Moreover, LLM-based approaches strike a practical balance between effectiveness and efficiency with inference times comparable to existing machine-learning models. Importantly, fine-tuned LLMs demonstrate measurable generalization across programming languages. These findings establish LLMs as a viable and efficient approach for tackling the longstanding challenge of equivalent mutant detection, offering new directions for advancing mutation testing in practice.  \nCCS Concepts: • Software and its engineering → Software testing and debugging.  \n∗ These authors contributed equally to this work.†Corresponding Author  \nAuthors’ Contact Information: Honglin Shu, Tianjin University, Tianjin, China and and Kyushu University, Fukuoka, Japan, [shu.honglin.167@s.kyushu-u.ac.jp](shu.honglin.167@s.kyushu-u.ac.jp); Zhao Tian, College of Intelligence and Computing, Tianjin University, Tianjin, China, [tianzhao@tju.edu.cn](tianzhao@tju.edu.cn); Dong Wang, College of Intelligence and Computing, Tianjin University, Tianjin, China, [dong_w@tju.edu.cn](dong_w@tju.edu.cn); Junji Yu, College of Intelligence and Computing, Tianjin University, Tianjin, China, [junjiyu@tju.edu.cn](junjiyu@tju.edu.cn); Jiazhe Zhang, College of Intelligence and Computing, Tianjin University, Tianjin, China, [2839197907z@gmail.com](2839197907z@gmail.com); Xuejie Cao, College of Intelligence and Computing, Tianjin University, Tianjin, China, [caoxuejie@tju.edu.cn](caoxuejie@tju.edu.cn); Junjie Chen, College of Intelligence and Computing, Tianjin University, Tianjin, China, [junjiechen@tju.edu.cn](junjiechen@tju.edu.cn); Yasutaka Kamei, Kyushu University, Fukuoka, Japan, [kamei@ait.kyushu-u.ac.jp](kamei@ait.kyushu-u.ac.jp).  \nPermission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for compon","cbCaikwnr7zIqHHj","https://ap.wps.com/l/cbCaikwnr7zIqHHj","pdf",4933059,5,1,37,"English","en",105,"# Introduction\n## Mutation testing overview\n## Challenges of equivalent mutants\n## Motivation for LLM-based EMD","[{\"question\":\"Why does equivalent mutant detection matter in mutation testing?\",\"answer\":\"Equivalent mutants introduce unnecessary computational cost and bias, which limits mutation testing’s practical value. EMD aims to identify and handle these redundant mutants more effectively.\"},{\"question\":\"What data and evaluation setup does the study use for EMD?\",\"answer\":\"The paper benchmarks LLM approaches using 3,302 Java mutant pairs and 1,088 C mutant pairs, comparing against state-of-the-art traditional methods.\"},{\"question\":\"How do LLM-based methods perform compared with traditional EMD approaches?\",\"answer\":\"Experimental results show LLM-based approaches achieve higher F1-scores than evaluated traditional methods. Fine-tuned code embedding provides the highest detection accuracy among the tested strategies while maintaining inference times comparable to existing ML models.\"}]",1784188459,93,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"large-language-models-for-multi-lingual-equivalent-mutant-detection-an-extended-empirical-study","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/large-language-models-for-multi-lingual-equivalent-mutant-detection-an-extended-empirical-study/83501/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why does equivalent mutant detection matter in mutation testing?","Question",{"text":76,"@type":77},"Equivalent mutants introduce unnecessary computational cost and bias, which limits mutation testing’s practical value. EMD aims to identify and handle these redundant mutants more effectively.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What data and evaluation setup does the study use for EMD?",{"text":81,"@type":77},"The paper benchmarks LLM approaches using 3,302 Java mutant pairs and 1,088 C mutant pairs, comparing against state-of-the-art traditional methods.",{"name":83,"@type":74,"acceptedAnswer":84},"How do LLM-based methods perform compared with traditional EMD approaches?",{"text":85,"@type":77},"Experimental results show LLM-based approaches achieve higher F1-scores than evaluated traditional methods. Fine-tuned code embedding provides the highest detection accuracy among the tested strategies while maintaining inference times comparable to existing ML models.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]