[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83888-en":3,"doc-seo-83888-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83888,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Listen Think Transcribe Continuous Latent Test-Time Scaling for ASR","End-to-end ASR models transcribe in a single pass, leaving no decoder opportunity to revisit difficult inputs. LatentASR introduces continuous latent test-time scaling on top of a fully frozen ASR backbone via two parameter-efficient modules: a Latent Adapter that iteratively refines a small set of latent prefix positions with bounded, stabilized updates, and a Value Head that predicts usefulness and halts early. Using only ~4M added parameters and a minimal 500-utterance activation set, LatentASR reduces WER on clean benchmarks and yields larger gains on accented, code-switched speech while skipping about half the utterances to save compute.","Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR  \nHo Lam Chung 1 ,2, Yiming Chen3, Dau-Cheng Lyu2, Hsiao-Tsung Hung2, Hung-yi Lee 1  \n1National Taiwan University, Taiwan 2ASUS, Taiwan 3National University of Singapore,  \nSingapore  \n[holam.chung@protonmail.com](holam.chung@protonmail.com) , [yiming.chen@u.nus.edu](yiming.chen@u.nus.edu) , ricer   [lu@asus.com](lu@asus.com) ,  \nAlexht [Hung@asus.com](Hung@asus.com) , [tlkagkb93901106@gmail.com](tlkagkb93901106@gmail.com)  \narXiv :2607 .0505 1v 1 [ cs . SD] 6 Jul 2026  \nAbstract  \nEnd-to-end ASR models transcribe in a single pass, leaving no room for the decoder to revisit hard inputs. We propose LatentASR, a parameter-efficient method that adds continuous latent test-time scaling to a frozen ASR backbone. Two small trainable modules drive it: a Latent Adapter that iterativelyrefines a few latent prefix positions through bounded, stabilized updates, and a Value Head that predicts whether extra computation will help and halts the loop early. The Qwen3-ASR-0.6B backbone stays fully frozen, and we train only ∼4M extra parameters. We activate this loop with a deliberately small, diverse 500-utterance training set. Under this minimaldata regime, standard adaptation methods all regress: full finetuning, LoRA, and prompt tuning each increase WER. LatentASR is the only tested method that reduces WER on both clean benchmarks (FLEURS −2 .54% and VoxPopuli −0 .47% relative) . The reductions are concentrated on intrinsically hard inputs. On accented and code-switched speech (ASCEND), LatentASR achieves a 16.0% relative CER reduction. Across 30 FLEURS languages (23,049 utterances), the multilingual WER decreases uniformly across resource tiers, confirming that the adapter generalizes without overfitting. Dynamic halting preserves most of the clean-set reduction at a fraction of the compute, skipping roughly half of all utterances at the entry gate. Our results show that a small, carefully chosen activation set can switch on test-time scaling inside a frozen ASR model without corrupting the model itself, converting fixed per-utterance compute into input-dependent compute where it is most needed. Index Terms: automatic speech recognition, latent test-time scaling, continuous thought, parameter-efficient finetuning  \n1. Introduction  \nEnd-to-end automatic speech recognition (ASR) models map audio directly to text in a single forward pass [1, 2] . This design is simple and effective. However, it forces a single left-to-right decoder to handle acoustic disambiguation, language modeling, and error correction at the same time.  \nRecent work has shown that allocating extra computation at inference time can substantially improve model performance. This paradigm is broadly known as test-time scaling, and includes explicit chain-of-thought reasoning [3], parallel sampling, and iterative refinement. Within this line, latent test-time scaling has drawn growing attention. For example, Coconut replaces discrete chain-of-thought tokens with continuous hidden states that are iteratively fed back into the model [4] . QuietSTaR trains models to generate and exploit implicit thoughts to improve next-token prediction [5] . Pause tokens provide extra compute steps without requiring explicit intermediate text [6] . Overall, these findings suggest that latent compute scaling is particularly beneficial on challenging inputs.  \nMotivated by this, we ask a simple question. Can an ASR decoder benefit from latent test-time scaling before it commits to a transcript?  \nApplying continuous latent scaling to ASR introduces two challenges absent in standard NLP tasks. First, ASR is a direct transcription task with no intermediate reasoning trajectory to distill, unlike Coconut and Quiet-STaR which rely on explicit chain-of-thought rationales. The model must discover how to use extra latent computation entirely unsupervised. Second, modern ASR backbones are typically kept frozen for efficiency and zero-","cbCaiaRNXfWs2u1M","https://ap.wps.com/l/cbCaiaRNXfWs2u1M","pdf",561739,3,1,9,"English","en",105,"# Abstract\n# Introduction\n## Problem motivation and challenges\n## Proposed method: LatentASR\n## Key contributions","[{\"question\":\"What performance and compute benefits are reported for LatentASR?\",\"answer\":\"LatentASR reduces WER on clean benchmarks and significantly improves CER on accented and code-switched speech. Dynamic halting preserves most of the clean-set gain while skipping roughly half the utterances, reducing average steps substantially without losing accuracy gains.\"}]",1784191241,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"listen-think-transcribe-continuous-latent-test-time-scaling-for-asr","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/listen-think-transcribe-continuous-latent-test-time-scaling-for-asr/83888/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What performance and compute benefits are reported for LatentASR?","Question",{"text":75,"@type":76},"LatentASR reduces WER on clean benchmarks and significantly improves CER on accented and code-switched speech. Dynamic halting preserves most of the clean-set gain while skipping roughly half the utterances, reducing average steps substantially without losing accuracy gains.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,119,122,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]