[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86466-en":3,"doc-seo-86466-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86466,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Efficiently Adapting Spoken Language Models for the Singaporean Context","Spoken language models (SLMs) combine speech perception with reasoning, yet adapting them to sensitive, multilingual domains remains insufficiently studied. The work adapts an open-source SLM for Singapore’s Home Team context across five speech tasks and four official languages, using LoRA fine-tuning, a surrogate text-QA dataset to reduce catastrophic forgetting, and a multi-task objective extending CoBa reweighting to speech. It introduces HTD-multilingualQA (504,853 samples) and evaluates HTMoonstone (5B), which matches or exceeds larger SLMs by up to 7× on most tasks while degrading original speech QA by under 2%.","Efficiently Adapting Spoken Language Models for the Singaporean Context  \nLanguage AI R&D, xData  \nHome Team Science & Technology Agency (HTX), Singapore  \nNg Jia Sheng Jason  \narXiv :2607 . 10092v 1 [ cs .CL] 11 Jul 2026  \nAbstract  \nSpoken language models (SLMs) unify speech perception and reasoning, but adapting them to sensitive domains is underexplored, especially when the original training data is inaccessible and the use case demands multilingual, spokenquery interaction. We adapt an open-source SLM to the Singaporean Home Team context across five speech tasks in Singapore’s four official languages, combining LoRA fine-tuning, a surrogate text-QA dataset that guards against catastrophic forgetting, and a multi-task objective that adapts the CoBa reweighting scheme to speech. We also build HTD-multilingualQA, a 504,853 sample multilingual QA dataset in text and spoken form. The resulting HTMoonstone (5B) matches or outperforms SLMs up to 7× its size on most tasks, attains the best accent and gender recognition among all models evaluated, and loses under 2% of its original speech QA ability.  \n1 Introduction  \nThe Home Team Science and Technology Agency (HTX) is a statutory board under Singapore’s Ministry of Home Affairs (MHA) . It serves as a “force multiplier” for the Home Team departments, applying science and technology to strengthen homeland security and public safety. Among the use cases HTX receives from these departments are speech tasks such as automatic speech recognition (ASR) and spoken question answering (QA), the latter forming the core of speech-to-speech conversational AI applications. Because much of this work involves highly sensitive data, closed-source or proprietary models are often unsuitable, which is why we turn to open-source alternatives.  \nHTX’s Language AI R&D team has been developing and adapting open-source foundational speech models, such as those for ASR. However,  \nthe field is moving towards spoken language models (SLMs) (Arora et al., 2025), which handle a wide range of cognitive and perception speech tasks (Peng et al., 2024) . Several open-source SLMs already exist, including Kimi-Audio (Kimi Team, 2025) and Audio-Flamingo-Next (Ghosh et al., 2026), as well as models built for the Singaporean context, such as MERaLiON-2 (MERaLiON Team, 2024) . However, because our operating environment and use cases are specific to the Home Team context, a suitable model must either be built from scratch or adapted from an existing open-source SLM. Training from scratch demands far more audio data, which is costly to acquire, so we favour adaptation. The SLM domain is still nascent, and few studies document how to adapt these models effectively for new tasks while preserving their original performance, particularly when the full dataset used to build them is unavailable.  \nAs we have use cases involving the use of text context and spoken queries and no Singaporean multilingual spoken QA training dataset currently exists, the development of a resource that covers Singaporean English, Mandarin, Bahasa Melayu, and Tamil is essential for adapting models to handle such spoken QA task effectively. Furthermore, adapting models without full access to their original training data often risks catastrophic forgetting. Our adaptation strategy is designed to mitigate this risk, ensuring we maximise performance on both perceptive and cognitive speech tasks in the Singaporean domain while preserving the model’s foundational capabilities.  \nIn this paper, we address these challenges and introduce:  \n1. a strategy to effectively adapt SLMs to the Singaporean context for multiple speech tasks;  \n2. HTD-multilingual-QA: a multilingual, multi-turn QA training dataset in both text and spoken form, grounded in Singaporean Home Team context that covers the 4 Singaporean languages-English, Mandarin, Bahasa Melayu, and Tamil for training SLMson spoken QA tasks; and  \n3. HT-Moonstone, an adapted 5B-parameter SLM that performs ","cbCaiktReBc40Dxg","https://ap.wps.com/l/cbCaiktReBc40Dxg","pdf",471502,4,1,10,"English","en",105,"# Introduction\n## Background and motivation\n## Key contributions\n# Related Work\n## Rise of Spoken Language Models\n## Singaporean Audio-Text Datasets","[{\"question\":\"Why is adapting spoken language models to the Singaporean Home Team context challenging?\",\"answer\":\"The domain involves highly sensitive data and typically requires multilingual spoken-query interaction. Additionally, the original training datasets for existing models are often inaccessible, making effective adaptation without losing prior abilities difficult.\"},{\"question\":\"What strategy is used to adapt an open-source SLM while mitigating catastrophic forgetting?\",\"answer\":\"The approach combines LoRA fine-tuning, a surrogate text-QA dataset designed to guard against catastrophic forgetting, and a multi-task objective that adapts the CoBa reweighting scheme to speech.\"},{\"question\":\"What are HTD-multilingualQA and HTMoonstone, and how do they perform?\",\"answer\":\"HTD-multilingualQA is a 504,853-sample multilingual QA dataset provided in both text and spoken form. HTMoonstone is an adapted 5B SLM that matches or outperforms SLMs up to 7× larger on most tasks, achieves the best accent and gender recognition among evaluated models, and shows under 2% degradation in original speech QA ability.\"}]",1784211892,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"efficiently-adapting-spoken-language-models-for-the-singaporean-context","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/efficiently-adapting-spoken-language-models-for-the-singaporean-context/86466/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is adapting spoken language models to the Singaporean Home Team context challenging?","Question",{"text":75,"@type":76},"The domain involves highly sensitive data and typically requires multilingual spoken-query interaction. Additionally, the original training datasets for existing models are often inaccessible, making effective adaptation without losing prior abilities difficult.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What strategy is used to adapt an open-source SLM while mitigating catastrophic forgetting?",{"text":80,"@type":76},"The approach combines LoRA fine-tuning, a surrogate text-QA dataset designed to guard against catastrophic forgetting, and a multi-task objective that adapts the CoBa reweighting scheme to speech.",{"name":82,"@type":73,"acceptedAnswer":83},"What are HTD-multilingualQA and HTMoonstone, and how do they perform?",{"text":84,"@type":76},"HTD-multilingualQA is a 504,853-sample multilingual QA dataset provided in both text and spoken form. HTMoonstone is an adapted 5B SLM that matches or outperforms SLMs up to 7× larger on most tasks, achieves the best accent and gender recognition among evaluated models, and shows under 2% degradation in original speech QA ability.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]