[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84814-en":3,"doc-seo-84814-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84814,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","REDDIT：用回放式分布编辑纠正ASR模型生成的时间戳漂移而不遗忘","Modern autoregressive ASR systems generate timestamps as decoded tokens, enabling timestamped transcription without frame-level aligners. However, generated timestamps drift across long non-speech spans: the transcript can stay linguistically plausible while the time axis shifts away from the audio. Benchmarks with gap and long-gap settings across 15 systems reveal this non-speech-induced drift. Naive timestamp-corrected fine-tuning can improve targeted alignment yet cause forgetting. REDDIT is a lightweight two-stage post-training method that corrects timestamps while matching the frozen base distribution on non-timestamp tokens, avoiding catastrophic forgetting. Experiments on Whispertiny show large mIoU gains and reduced out-of-domain alignment errors with minimal parameter updates.","REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing  \nCheng-Kang Chou∗1, Ming-To CHUANG∗1, Ke-Han Lu 1 , Chan-Jan Hsu2 , Hung-yi Lee3  \n1National Taiwan University 2Carnegie Mellon University 3NTU Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)  \narXiv :2607 .05364v 1 [ cs .CL] 6 Jul 2026  \nAbstract—Modern autoregressive ASR systems can emit timestamps as decoded tokens, enabling timestamped transcription without frame-level aligners or inference-time post-processing. We show that these generated timestamps can drift across long non-speech spans: the transcript may remain plausible, but the decoded time axis drifts away from the audio. We study this non-speech-induced timestamp drift with self-built gap and long-gap benchmarks across 15 evaluated timestamp-producing ASR and audio-language systems. Naive timestamp-corrected fine-tuning improves alignment but can severely degrade non-target ASR behavior, exposing a forgetting problem. We propose REDDIT (REplay-based Distribution eDITing), a lightweight two-stage post-training framework that corrects timestamps while avoiding this catastrophic forgetting: it first edits timestamp targets under the model’s own replayed decoder context while matching the frozen base distribution on non-timestamp tokens, then applies a short edited-prefix refinement stage. In this framework, we construct correction supervision without human transcripts or human timestamp annotations by combining VAD-trimmed speech spans with inserted non-speech gaps and known concatenation offsets. On Whispertiny, 34.9 hours of targeted correction audio used and only 1.6% of model parameters updated, raising long-gap mIoU from 38.7% to 95.0% and reducing mixed-gap out-of-domain AAS from 2752 ms to 223 ms while preserving CV-en MER at 41.3% (versus 524.2% for ordinary SFT decoder tuning).  \nIndex Terms—automatic speech recognition, model-generated timestamps, non-speech gap, replay-based learning, post-training, catastrophic forgetting  \nI. INTRODUCTION  \nAutomatic speech recognition (ASR) and audio-language models increasingly need to answer not only what was said, but also when it was said. Autoregressive speech systems expose this interface by generating time as part of the decoded output, either as timestamp tokens in ASR models such as Whisper or as word-level and text-form temporal predictions in speech-aware language models [1], [3]–[5] . This self-generated timestamp interface avoids the need for a separate VAD, frame-level aligner, or forced-alignment module at inference time.  \nThis convenience introduces a failure mode that transcript-centric metrics can miss. Speech evaluation already extends beyond lexical content, e.g., tonality [46]; here, the missing dimension is temporal placement. Because the time axis is generated by the same decoder that emits text, a model can produce plausible content while assigning it to the wrong region of the audio. We focus on non-speech-induced timestamp drift: when speech follows a long silent or non-speech span, self-generated timestamps may place the following speech too early, too late, or under a globally displaced time axis, rendering subtitle timing unusable despite having an acceptable word error rate or task-level transcript quality.  \nA likely reason this failure is under-tested is the mismatch between common training and long-gap deployment conditions. Timestamped ASR examples often begin near speech onset, and VAD-based trimming can further increase speech density [11] . During training, this practice may reduce the model’s exposure to long non-speech prefixes and strengthen a prior for early timestamp tokens near 0s.  \nPost-hoc alignment pipelines can refine boundaries after decoding, but they operate at a different level than correcting the model’s native  \n* Equal contribution.  \npredictions. VAD segmentation, forced alignment, and attentionbased alignment add inference-time m","cbCaic3uMvizO43x","https://ap.wps.com/l/cbCaic3uMvizO43x","pdf",386342,3,1,"English","en",105,"# Introduction\n## Non-speech-induced timestamp drift\n## Limitations of post-hoc alignment and naive fine-tuning\n## REDDIT approach and contributions","[{\"question\":\"什么是ASR中“非语音诱导的时间戳漂移”？\",\"answer\":\"当长时间静音或非语音段之后出现语音时，模型自生成的时间戳可能把后续语音放到过早/过晚或整体时间轴偏移的位置，导致字幕时间不可靠，即使词错率或转写内容质量仍可能看起来合理。\"},{\"question\":\"为什么仅做时间戳校正的朴素微调会带来“遗忘”问题？\",\"answer\":\"朴素的时间戳校正微调可能在目标时间戳数据上改善对齐，但会在非目标ASR行为上退化，说明难点不只是时间戳监督，而是需要对时间编辑具有抗遗忘能力。\"},{\"question\":\"REDDIT如何在纠正时间戳的同时避免灾难性遗忘？\",\"answer\":\"REDDIT采用两阶段后训练：先在模型回放的解码器上下文下编辑时间戳目标，并将非时间戳token的分布锚定到冻结基模型；再从阶段一的检查点进行短的编辑前缀精修，从而实现时间轴纠正而不过度破坏原有行为。\"}]",1784198414,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"reddit-replay-based-distribution-editing-for-timestamp-drift-correction-in-asr-without-forgetting","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":20},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/reddit-replay-based-distribution-editing-for-timestamp-drift-correction-in-asr-without-forgetting/84814/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"什么是ASR中“非语音诱导的时间戳漂移”？","Question",{"text":74,"@type":75},"当长时间静音或非语音段之后出现语音时，模型自生成的时间戳可能把后续语音放到过早/过晚或整体时间轴偏移的位置，导致字幕时间不可靠，即使词错率或转写内容质量仍可能看起来合理。","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"为什么仅做时间戳校正的朴素微调会带来“遗忘”问题？",{"text":79,"@type":75},"朴素的时间戳校正微调可能在目标时间戳数据上改善对齐，但会在非目标ASR行为上退化，说明难点不只是时间戳监督，而是需要对时间编辑具有抗遗忘能力。",{"name":81,"@type":72,"acceptedAnswer":82},"REDDIT如何在纠正时间戳的同时避免灾难性遗忘？",{"text":83,"@type":75},"REDDIT采用两阶段后训练：先在模型回放的解码器上下文下编辑时间戳目标，并将非时间戳token的分布锚定到冻结基模型；再从阶段一的检查点进行短的编辑前缀精修，从而实现时间轴纠正而不过度破坏原有行为。","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]