[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82419-en":3,"doc-seo-82419-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82419,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Tokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASR","Lightweight speech recognition models are crucial for edge deployment, yet highly optimized architectures like Moonshine can fail on morphologically rich non‑Latin languages such as Bengali. The failure stems from an English-centric byte-level tokenizer that splits Bengali words into long, high-fertility byte chains, causing catastrophic autoregressive collapse during inference. A vocabulary transplantation pipeline replaces the decoder vocabulary with BanglaBERT WordPiece and resizes token embeddings. Experiments reduce tokenizer fertility from 9.16 to 1.30, cut autoregressive sequence length by 85.8%, and eliminate decoding instability. On the 882-hour Lipi-Ghor dataset, the modified model attains 21.54% WER with an RTF of 0.0053, offering a scalable cross-script adaptation blueprint without costly pre-training.","Tokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient  \nBengali ASR  \nSanjid Hasan 1 Md. Abdur Rahman 2  \narXiv :2607 .09598v 1 [ cs .CL] 10 Jul 2026  \nAbstract  \nLightweight speech recognition models are critical for edge deployment, yet highly optimized architectures like Moonshine often fail on morphologically rich, non-Latin languages such as Bengali. This study identifies the root cause of this failure as the model’s English-centric bytelevel tokenizer, which fragments Bengali words into high-fertility byte chains and triggers catastrophic autoregressive collapse during inference.  \nTo resolve this, a novel vocabulary transplantation pipeline is proposed to replace the decoder vocabulary with the native-script BanglaBERT WordPiece vocabulary and resize the corresponding token embedding matrix. Experimental results demonstrate a reduction in token fertility from  \n9.16 to 1.30 . By decreasing autoregressive sequence length by 85.8%, decoding instability is entirely mitigated. When evaluated on the 882-hour Lipi-Ghor dataset, the modified architecture achieves a competitive 21.54% Word Error Rate (WER) and a Real-Time Factor (RTF) of 0 .0053. Ultimately, this research provides a scalable, reproducible blueprint for cross-script adaptation of compact ASR models without the need for resource-intensive pre-training.  \n1. Introduction  \nRecent advancements in Automatic Speech Recognition (ASR) have achieved remarkable performance through largescale, self-supervised learning. However, these models often rely on broad-coverage tokenizers designed for highresource languages, which perform suboptimally on morphologically rich languages like Bengali. This mismatch results in high token fertility, leading to increased sequence length and autoregressive collapse during inference.  \n1 Khulna University of Engineering & Technology, Bangladesh 2Military Institute of Science and Technology, Bangladesh. Correspondence to: Sanjid Hasan \u003C[hasan2203090@stud.kuet.ac.bd](hasan2203090@stud.kuet.ac.bd) >.  \nProceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \nIn this paper, we address this bottleneck by proposing a Tokenizer Transplantation pipeline. Rather than training from scratch, we surgically adapt a lightweight base architecture (e.g., Moonshine) by replacing its vocabulary with a native-script BanglaBERT tokenizer. We show that by decoupling acoustic representation learning from language modeling, we can stabilize the decoder and significantly improve performance on the 882-hour Lipi-Ghor dataset. Our approach provides a scalable blueprint for adapting efficient, pre-trained architectures to under-represented languages, ensuring that the benefits of ASR reach linguistically diverse communities. The full implementation details for our framework are available in our repository.1  \n2. Related Work  \n2.1. Speech Representation Learning  \nThe current ASR landscape is dominated by frameworks such as Baevski et al. (2020) and Babu et al. (2022), which leverage self-supervised learning for robust feature extraction. While models like Radford et al. (2023) have pushed the boundaries of accuracy via large-scale supervision, their computational footprint remains prohibitive for edge devices. Consequently, lightweight architectures like Moonshine (Useful Sensors, 2024; King & Sabra, 2025) have emerged as essential for local, resource-constrained deployment.  \n2.2. Vocabulary Adaptation  \nTokenization quality is a primary determinant of model performance (Rust et al., 2021) . Prior efforts, such as WECHSEL (Minixhofer et al., 2022) and FOCUS (Dobler & de Melo, 2023), have proposed innovative ways to initialize subword embeddings for cross-lingual transfer. However, these methods primarily address text-based Large Language Models (LLMs) . Our work extends these principles to ASR by identifying ”tokenizer transplantation” as a critical re","cbCaigifaWXs4BnI","https://ap.wps.com/l/cbCaigifaWXs4BnI","pdf",1438852,1,5,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n## Speech Representation Learning\n## Vocabulary Adaptation\n## Bengali ASR\n# Tokenizer Fertility and Decoding Instability\n# Methodology","[{\"question\":\"Why do lightweight ASR models like Moonshine struggle with Bengali?\",\"answer\":\"Their English-centric byte-level tokenizer fragments Bengali words into long byte chains, producing high tokenizer fertility that makes autoregressive decoding unstable during inference.\"},{\"question\":\"What is the proposed solution in Tokenizer Transplantation?\",\"answer\":\"It replaces the decoder vocabulary with the native-script BanglaBERT WordPiece vocabulary and resizes the token embedding matrix, reducing token fertility and stabilizing decoding.\"},{\"question\":\"How much does the approach improve tokenizer fertility, sequence length, and inference stability?\",\"answer\":\"Tokenizer fertility drops from 9.16 to 1.30, autoregressive sequence length is reduced by 85.8%, and decoding instability is fully mitigated.\"}]",1784180265,13,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"tokenizer-transplantation-mitigating-autoregressive-collapse-in-edge-efficient-bengali-asr","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/tokenizer-transplantation-mitigating-autoregressive-collapse-in-edge-efficient-bengali-asr/82419/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do lightweight ASR models like Moonshine struggle with Bengali?","Question",{"text":75,"@type":76},"Their English-centric byte-level tokenizer fragments Bengali words into long byte chains, producing high tokenizer fertility that makes autoregressive decoding unstable during inference.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the proposed solution in Tokenizer Transplantation?",{"text":80,"@type":76},"It replaces the decoder vocabulary with the native-script BanglaBERT WordPiece vocabulary and resizes the token embedding matrix, reducing token fertility and stabilizing decoding.",{"name":82,"@type":73,"acceptedAnswer":83},"How much does the approach improve tokenizer fertility, sequence length, and inference stability?",{"text":84,"@type":76},"Tokenizer fertility drops from 9.16 to 1.30, autoregressive sequence length is reduced by 85.8%, and decoding instability is fully mitigated.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":21,"slug":137},19,"General","general"]