[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82893-en":3,"doc-seo-82893-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82893,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Context Aware ASR for Mandarin Technical Lectures","Technical lectures combine Mandarin speech with embedded English technical terms, where those terms carry key meaning while occupying few characters. As a result, character error rate (CER) can remain low even when the crucial terms are misrecognized. The study evaluates whether lecture context improves technical-term recognition by building a term-rich Mandarin AI/ML lecture benchmark and introducing term-centric metrics. A two-pass, reference-free decoding method uses a self-built glossary from first-pass hypotheses to improve term recall across multiple ASR backbones while maintaining or reducing CER.","Context-Aware ASR for Mandarin Technical Lectures  \nHo-Lam Chung 1 ,2, Yiming Chen2, Hung-yi Lee 1  \n1National Taiwan University, Taiwan 2ASUS  \n[holam.chung@protonmail.com](holam.chung@protonmail.com) , [yiming.chen@u.nus.edu](yiming.chen@u.nus.edu) , [tlkagkb93901106@gmail.com](tlkagkb93901106@gmail.com)  \narXiv :2607 .05058v 1 [ cs . SD] 6 Jul 2026  \nAbstract  \nTechnical lectures mix Mandarin speech with English technical terms. These terms carry the core meaning of the lecture, yet they occupy few characters. Character error rate (CER) therefore hides their recognition failures. We study whether lecture context helps recognize these terms. We build a termrich Mandarin AI/ML lecture benchmark, and we define termcentric metrics that measure technical-term recognition directly. We then propose a two-pass, reference-free decoding method. The first pass runs segment-only ASR. We extract the most frequent technical terms from the first-pass hypotheses, and we prompt the recognizer with this self-built glossary in the second pass. Across five ASR backbones, the first-pass glossary raises term recall for every model and holds or lowers CER on all five. On Breeze-ASR-25 it lifts term recall from 52 .50% to 60.13% while lowering CER, and a hybrid that adds a small external term list reaches 62.05% recall and 82.73% term precision. Lecture context, recovered from the model’s own output, is a practical signal for technical-term recognition. Term-centric evaluation exposes errors that CER misses.  \nIndex Terms: speech recognition, code-switching, contextual biasing, technical terms, Mandarin  \n1. Introduction  \nLecture ASR powers subtitles, search, and note-taking for students. In AI/ML lectures, bilingual technical terms anchor the content. Terms such as RAG, SWE-bench, AI Agent, token, and embedding name the concepts the speaker explains. A student who reads “RIG” instead of “RAG” loses the point of the segment.  \nThese terms occupy a small fraction of the characters. A transcript can reach low CER while still misrecognizing the key term. The Mandarin around the term is fluent, so CER stays low. The usable content, however, is wrong. CER does not capture this failure.  \nLecture terms are bursty. Once the speaker introduces AI Agent, the term recurs across many later segments. The lecture title, the previous transcript, and the repeated terms all point to the same vocabulary. Segment-only decoding discards this signal. It recognizes each segment in isolation, so it cannot use the strong context that the lecture provides.  \nWe address this gap by building a term-rich benchmark, measuring term recognition directly, and supplying a self-built glossary as decoding context.  \nOur contributions are as follows.  \n• We release a term-rich Mandarin AI/ML lecture ASR benchmark. It contains 8,888 technical-term occurrences over a 5 .01-hour code-switched test set.  \n• We define term-centric evaluation. It reports term recall, term precision, term F1, and term error rate, so it separates technical-term accuracy from overall CER.  \n• We propose a two-pass, reference-free glossary prompt. It extracts a lecture glossary from first-pass hypotheses, and it improves term recall across five ASR backbones while holding or lowering CER on all of them. A hybrid that adds a small external glossary reaches 62.05% term recall on Breeze-ASR-25, closing part of the remaining headroom.  \n2. Related Work  \nCode-switching Mandarin-English ASR. Mandarin-English code-switching is a long-standing challenge for ASR. Public corpora such as SEAME [1], the ASRU 2019 challenge set [2], and ASCEND [3] established Mandarin-English benchmarks, but they target conversational rather than lecture speech. Recent work fine-tunes large models on Mandarin with English terms [4], and distills language-specific knowledge for codeswitching [5] . These methods improve the acoustic model. We instead keep the model frozen, and we add context at decoding time.  \nContextual biasing. Contex","cbCaidHUb0WcBr7J","https://ap.wps.com/l/cbCaidHUb0WcBr7J","pdf",198974,1,5,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n# Benchmark","[{\"question\":\"Why can character error rate (CER) fail to reflect technical-term recognition errors in Mandarin technical lectures?\",\"answer\":\"Technical terms carry the lecture’s core meaning but occupy few characters. The Mandarin around them can be recognized fluently, keeping CER low even when the key terms are wrong.\"},{\"question\":\"How does the proposed two-pass, reference-free method build context for decoding?\",\"answer\":\"The first pass performs segment-only ASR, then extracts frequent technical terms from its hypotheses to form a self-built glossary. The second pass uses this glossary as decoding context, without relying on an external reference list.\"},{\"question\":\"What improvements does the glossary prompting bring, and how is evaluation performed beyond CER?\",\"answer\":\"Across five ASR backbones, the glossary improves term recall while holding or lowering CER. Term-centric evaluation reports term recall, precision, F1, and term error rate to expose errors that CER can miss.\"}]",1784183748,13,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"context-aware-asr-for-mandarin-technical-lectures","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/context-aware-asr-for-mandarin-technical-lectures/82893/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why can character error rate (CER) fail to reflect technical-term recognition errors in Mandarin technical lectures?","Question",{"text":74,"@type":75},"Technical terms carry the lecture’s core meaning but occupy few characters. The Mandarin around them can be recognized fluently, keeping CER low even when the key terms are wrong.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does the proposed two-pass, reference-free method build context for decoding?",{"text":79,"@type":75},"The first pass performs segment-only ASR, then extracts frequent technical terms from its hypotheses to form a self-built glossary. The second pass uses this glossary as decoding context, without relying on an external reference list.",{"name":81,"@type":72,"acceptedAnswer":82},"What improvements does the glossary prompting bring, and how is evaluation performed beyond CER?",{"text":83,"@type":75},"Across five ASR backbones, the glossary improves term recall while holding or lowering CER. Term-centric evaluation reports term recall, precision, F1, and term error rate to expose errors that CER can miss.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,108,113,118,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":105,"show_sort_weight":106,"slug":107},"Comic",60,"comic",{"id":109,"doc_module":4,"doc_module_name":45,"category_name":110,"show_sort_weight":111,"slug":112},6,"Technology",50,"technology",{"id":114,"doc_module":4,"doc_module_name":45,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":119,"slug":120},30,"research-report",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":21,"slug":136},19,"General","general"]