[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83980-en":3,"doc-seo-83980-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83980,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Revisiting the Relation Between Language Model Perplexity and ASR Word Error Rate for Modern End-to-End Speech Recognition","Language model (LM) perplexity (PPL) has long been used as a proxy for automatic speech recognition (ASR) word error rate (WER), often described as an approximately linear relationship in log-log space. Modern end-to-end ASR systems challenge this link because they include internal language modeling, are frequently evaluated without external LMs, and can incorporate neural LMs or LLMs via multiple decoding strategies. This study re-examines PPL–WER linearity, tests external-LM usefulness, analyzes encoder context effects, and evaluates how internal language modeling subtraction alters the observed relation.","Revisiting the Relation Between Language Model Perplexity and ASR Word Error Rate for Modern End-to-End Speech Recognition  \nMohammad Zeineldeen∗†, Albert Zeyer∗†, Haoran Zhang†, Robin Schmitt∗†, Ralf Schl¨uter∗†, Hermann Ney∗†  \n∗ AppTek. ai GmbH, Aachen, Germany  \n† Machine Learning and Human Language Technology Group, Faculty of Computer Science,  \nRWTH Aachen University Aachen, Germany  \n{zeineldeen, zeyer, schmitt, schlueter, [ney](ney}@ml.rwth-aachen.de)[}](ney}@ml.rwth-aachen.de)[@ml.rwth-aachen.de](ney}@ml.rwth-aachen.de)  \narXiv :2607 .056 12v 1 [ cs .CL] 6 Jul 2026  \nAbstract—Language model (LM) perplexity (PPL) has historically been used as a proxy for automatic speech recognition (ASR) word error rate (WER), with prior work reporting an approximately linear relation in log-log space. Modern end-toend ASR systems challenge this assumption because they already contain internal language modeling capacity, are often evaluated without external language models, and can now be combined with neural LMs and large language models (LLMs) through different recognition strategies. This paper revisits the relation between PPL and WER for modern ASR systems. We study whether external LMs still improve current end-to-end ASR systems, whether the PPL-WER relation remains linear in loglog space, how encoder context length affects this relation, and how LLM perplexities fit into the trend observed for standard neural LMs. We further investigate internal language modeling (ILM) in attention-based encoder-decoder systems and show that ILM subtraction changes the observed PPL-WER relation, indicating that the decoder’s internal LM must be considered when interpreting the effect of external LM quality.  \nIndex Terms—speech recognition, language models, perplexity, shallow fusion, large language models  \nI. INTRODUCTION & RELATED WORK  \nLanguage models (LMs) have long been central to automatic speech recognition (ASR) . Perplexity (PPL) is commonly used as an intrinsic LM metric, tracing back to its introduction as a measure of speech task difficulty [1] . However, ASR performance is measured by word error rate (WER), which also depends on the acoustic model, decoding strategy, pruning, and scale tuning. Thus, lower LM perplexity does not automatically translate into lower WER. Earlier studies, mostly using n-gram LMs [2], reported experimental evidence for a correlation between perplexity and WER across several ASR tasks [3]–[5] . The work in [4] systematically studied this question and found that WER and PPL are approximately linearly related in log-log space. The same log-log linear relation was later observed in [6] for hybrid ASR models on Quaero and LibriSpeech dev-clean using both n-gram and long shortterm memory (LSTM) LMs [7] . This observation has been influential because it suggests that relative PPL improvements can be translated into expected WER improvements through a power-law relation.  \nLarge-scale ASR studies also support stronger external LMs. The study in [8] showed that increasing LM scale and using large text corpora yields consistent WER reductions. Later work extended LM integration to neural LMs and endto-end ASR. Shallow fusion combines an external LM with the ASR model during beam search [9], while cold fusion incorporates a pretrained LM during training [10] . Several integration strategies were compared in [11] for attentionbased encoder-decoder (AED) ASR and found shallow fusion to be a strong and simple first-pass decoding method across multiple conditions. Neural Transformer LMs [12] have also been applied successfully to hybrid and end-to-end ASR, including lattice rescoring and shallow fusion [13] .  \nModern end-to-end ASR systems make the PPL-WER relation less direct. In Connectionist temporal classification (CTC)  \n[14], recurrent neural network transducer (RNN-T) [15], and AED [16] systems, acoustic and language modeling components are no longer cleanly separated as in classical hybrid ASR. AED and tra","cbCaiv9fNX6ELS1n","https://ap.wps.com/l/cbCaiv9fNX6ELS1n","pdf",366335,4,1,"English","en",105,"# Abstract\n# Introduction & Related Work\n## Classical PPL–WER correlation\n## External LM integration strategies\n## Why end-to-end ASR weakens the proxy assumption\n## Motivation for systematic re-examination","[{\"question\":\"Why might LM perplexity no longer predict ASR word error rate in modern end-to-end systems?\",\"answer\":\"End-to-end ASR models learn internal language modeling from speech-text training data, so external LM effects can interact with the decoder’s internal LM rather than mapping directly to perplexity.\"},{\"question\":\"What does the paper investigate about the PPL–WER relationship?\",\"answer\":\"It tests whether external LMs still improve modern end-to-end ASR, whether the PPL–WER relation remains approximately linear in log-log space, and how encoder context length influences the relation.\"},{\"question\":\"How does internal language model subtraction affect the observed PPL–WER trend?\",\"answer\":\"The paper shows that ILM subtraction changes the PPL–WER relation, indicating that the decoder’s internal LM must be considered when interpreting how external LM quality impacts WER.\"}]",1784191825,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"revisiting-the-relation-between-language-model-perplexity-and-asr-word-error-rate-for-modern-end-to-end-speech-recognition","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":20},"https://docshare.wps.com/document/revisiting-the-relation-between-language-model-perplexity-and-asr-word-error-rate-for-modern-end-to-end-speech-recognition/83980/",{"url":51,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why might LM perplexity no longer predict ASR word error rate in modern end-to-end systems?","Question",{"text":74,"@type":75},"End-to-end ASR models learn internal language modeling from speech-text training data, so external LM effects can interact with the decoder’s internal LM rather than mapping directly to perplexity.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What does the paper investigate about the PPL–WER relationship?",{"text":79,"@type":75},"It tests whether external LMs still improve modern end-to-end ASR, whether the PPL–WER relation remains approximately linear in log-log space, and how encoder context length influences the relation.",{"name":81,"@type":72,"acceptedAnswer":82},"How does internal language model subtraction affect the observed PPL–WER trend?",{"text":83,"@type":75},"The paper shows that ILM subtraction changes the PPL–WER relation, indicating that the decoder’s internal LM must be considered when interpreting how external LM quality impacts WER.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]