[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82543-en":3,"doc-seo-82543-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82543,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","YOMI-Bench：用于评估日语汉字阅读与语音理解的基准","YOMI-Bench is introduced as a dedicated benchmark for evaluating kanji reading and phonological understanding in large language models for Japanese. Japanese kanji often have multiple possible readings that cannot be reliably determined from surface text alone, which leads to empirically low kanji-reading performance. The benchmark defines four targeted tasks and evaluates one multilingual open LLM, four Japanese-specific open LLMs, and five commercial LLMs, showing limited accuracy even for Japanese-focused models and poor results on reading-dependent generation.","YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese  \nRyota Mibayashi 1 , Hiroya Takamura2 , Hitomi Yanaka3 ,4 ,5  \n1 Kobe University 2 National Institute of Advanced Industrial Science and Technology (AIST)  \n3 The University of Tokyo 4 RIKEN 5 Tohoku University  \narXiv :2607 .00664v 1 [ cs .CL] 1 Jul 2026  \nAbstract  \nWe propose YOMI-Bench, a benchmark for evaluating kanji reading and phonological understanding of large language models (LLMs) for Japanese. In Japanese, a single kanji character often has multiple possible readings, making it difficult to infer the correct reading from surface-level text alone. Due to these linguistic characteristics, it is empirically known that LLMs exhibit low performance in kanji reading for Japanese. The proposed YOMIBench consists of four tasks specifically designed to evaluate kanji reading performance in Japanese. In our evaluation using YOMIBench, we assessed one multilingual open LLM, four Japanese-specific open LLMs, and five commercial LLMs. As a result, we found that even Japanese-specific models show low performance, and that commercial models also perform poorly on generation tasks that require consideration of kanji readings.  \n1 Introduction  \nThe multilingual abilities of large language models (LLMs) have been analyzed from various perspectives (Zhu et al., 2024) . Among these, we focus on the ability of LLMs to recognize and utilize the phonological readings of text, which is related to tasks such as grapheme-to-phoneme (G2P) estimation and phoneme-aware text generation. The relationship between written forms and their pronunciations varies substantially across languages, and even among languages that share the same writing system, notable differences can be observed. For example, although Chinese and Japanese share many kanji characters, nearly all kanji characters in Chinese correspond to a single pronunciation, except for approximately 10% of cases (Matsuo et al., 2010) . In contrast, approximately 60% of Japanese kanji have multiple possible readings (see details in Sec. 3.1) , requiring more complex linguistic information for correct interpretation. For instance,  \nFigure 1: The overview of YOMI-Bench.  \nthe kanji character “覚” has three possible readings:“kaku,” “obo,” and “sa.” These readings vary depending on the word in which the character appears, such as “覚醒(kakusei/Awakening),” “覚える(oboeru/Memorize),” and “覚める(sameru/Wake up).” Thus, in order to correctly predict the readings of kanji characters appearing within individual words, models should not only possess the knowledge of the multiple readings of each kanji character, but also predict the correct reading based on the word. For practical situations, the ability to correctly understand the readings of kanji is important for LLMs to solve classification and generation tasks that take rhyming into account, such as proofreading and generating lyrics (Potash et al., 2015 ; Nikolov et al., 2020), rap verses (Xue et al., 2021 ; Mibayashi et al., 2023), and advertising texts (Lei et al., 2022) .  \nHowever, the reading ability of LLMs that considers such linguistic characteristics of Japanese remains largely unexplored. Therefore, as illustrated in Figure 1, we construct YOMI-Bench, a benchmark for evaluating Japanese reading performance of LLMs. YOMI-Bench is a multi-task evaluation set involving seven types of binary/multiple QA and text generation tasks. For example, the question “Generate the reading of the word覚醒” asks LLMs to answer the correct kanji reading as a text generation task, where the gold answer is “kakusei.”We also use YOMI-Bench to evaluate the reading  \nabilities of representative LLMs. The contributions of this study include:  \n• We construct a challenging benchmark involving seven tasks that requires correct understanding of phonological readings in Japanese and release them on GitHub 1 as publicly available linguistic resources.  \n• Using the ","cbCaifiU1nrXbszi","https://ap.wps.com/l/cbCaifiU1nrXbszi","pdf",475134,3,1,5,"English","en",105,"# Introduction\n# Related Work\n## Grapheme-to-Phoneme (G2P) Benchmarks\n## LLM Benchmarks in Japanese","[{\"question\":\"Why is kanji reading in Japanese difficult for LLMs?\",\"answer\":\"A single kanji character can have multiple readings, and the correct one often depends on the word context. Surface text alone is insufficient to infer the reading reliably.\"},{\"question\":\"What tasks does YOMI-Bench include?\",\"answer\":\"YOMI-Bench is a multi-task benchmark covering seven types of binary/multiple QA and text generation tasks designed to evaluate kanji reading and phonological understanding.\"},{\"question\":\"How did different LLM categories perform on YOMI-Bench?\",\"answer\":\"Evaluation includes one multilingual open model, four Japanese-specific open models, and five commercial models. Results indicate even Japanese-specific models show low performance, and commercial models also perform poorly on generation tasks requiring consideration of kanji readings.\"}]",1784181428,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"yomi-bench-a-benchmark-for-evaluating-kanji-reading-and-phonological-understanding-of-llms-for-japanese","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/yomi-bench-a-benchmark-for-evaluating-kanji-reading-and-phonological-understanding-of-llms-for-japanese/82543/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is kanji reading in Japanese difficult for LLMs?","Question",{"text":75,"@type":76},"A single kanji character can have multiple readings, and the correct one often depends on the word context. Surface text alone is insufficient to infer the reading reliably.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What tasks does YOMI-Bench include?",{"text":80,"@type":76},"YOMI-Bench is a multi-task benchmark covering seven types of binary/multiple QA and text generation tasks designed to evaluate kanji reading and phonological understanding.",{"name":82,"@type":73,"acceptedAnswer":83},"How did different LLM categories perform on YOMI-Bench?",{"text":84,"@type":76},"Evaluation includes one multilingual open model, four Japanese-specific open models, and five commercial models. Results indicate even Japanese-specific models show low performance, and commercial models also perform poorly on generation tasks requiring consideration of kanji readings.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":22,"slug":137},19,"General","general"]