[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85086-en":3,"doc-seo-85086-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85086,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","When Synthetic Speech Is All You Have: Better Call GRPO","LLM-based ASR for regulated domains like banking is constrained by privacy, making real speech expensive and legally difficult to collect, so synthetic text-to-speech (TTS) becomes an attractive alternative. Synthetic speech, however, remains acoustically mismatched and prior work largely relies on supervised fine-tuning (SFT). Group Relative Policy Optimization (GRPO) enables critic-free reinforcement learning that rewards low-WER hypotheses, yielding a 40% relative WER reduction from SFT (36.71%→22.09%) and up to 45% with SFT-then-GRPO.","When Synthetic Speech Is All You Have:  \nBetter Call GRPO  \nShashi Kumar1 ,2,⋆, Yanis Labrak1,⋆ ,  \nHasindri Watawana 1 ,2 , Sergio Burdisso 1 , Esa´u Villatoro-Tello 1 , Kadri Hacio˘glu3 , Petr Motlicek1 ,4 , Andreas Stolcke3  \n1 Idiap Research Institute, Martigny, Switzerland  \n2 EPFL, Lausanne, Switzerland; 3 Uniphore, U.S.A.  \n4 Brno University of Technology, Brno, Czech Republic  \narXiv :2607 .08409v 1 [ cs .CL] 9 Jul 2026  \nAbstract—LLM-based ASR adapted to regulated domains such as banking is bottlenecked by privacy: real speech is costly and legally constrained to collect, making synthetic text-to-speech (TTS) an attractive substitute. Yet synthetic speech stays acoustically mismatched with real recordings, and work on this gap has stayed within supervised fine-tuning (SFT). We instead turn to reinforcement learning, and show that Group Relative Policy Optimization (GRPO) extracts far more from the same synthetic speech than SFT. Synthetic-only adaptation of the model with GRPO, a critic-free method rewarding low-WER hypotheses, reduces WER by 40% relative to SFT (36.71%→22.09%), and an SFT-then-GRPO combination pushes this further to 45%. Wetrace the gain to behavior rather than representation: GRPO reduces insertion errors by improving stopping calibration and speech-to-text alignment by better anchoring attention to audio, leaving early-layer representations intact. When synthetic speech is the main resource, reinforcement learning should be preferred over supervised fine-tuning.  \nIndex Terms—GRPO, Reinforcement Learning, Synthetic Speech, Text-to-Speech, Domain Adaptation, Automatic Speech Recognition, Speech LLM, Low-Resource  \nI. INTRODUCTION  \nSpeech recordings, such as customer–agent telephone calls in regulated domains such as banking are among the hardest data to collect at scale, because privacy is the central concern and regulations such as the GDPR and the EU AI Act [1] restrict how such recordings, which may qualify as biometric data, can be stored, shared, and reused for training due to personally identifiable and financial information they carry.  \nThese constraints slow the adoption of speech technologies in high-stakes domains, as systems must be adapted to indomain speech data, yet sourcing such recordings remains slow, costly, and tightly regulated.  \nTo sidestep this problem, the community has explored substituting real in-domain audio with speech synthesized using text-to-speech (TTS) systems from available transcripts [2]–[6] . However, synthetic speech still differs acoustically from real recordings [7], and this severe mismatch keeps the resulting performance gain sub-optimal, motivating continuation of work on closing the synthetic-to-real gap [8], [9] . A recent work convolves synthetic speech with room impulse responses  \n⋆ Equal contribution.  \n(RIRs) to reproduce real channel characteristics, matching the real-speech baseline while using only a quarter of the real data [10] .  \nThese works typically build on pre-trained large language model (LLM) architectures adapted to the speech modality by plugging in a speech encoder [11]–[16], yet they all share the same limitation of remaining within the supervised fine-tuning (SFT) paradigm. Meanwhile, the broader NLP community has moved beyond cross-entropy and successfully leverages synthetic data [17]–[19] through reinforcement learning (RL) approaches such as Direct Preference Optimization (DPO) [20] and Group Relative Policy Optimization (GRPO) [21] .  \nThese observations, combined with the acoustic irregularities of synthetic speech, lead us to hypothesize that this reliance on SFT is what caps the performance achievable with synthetic speech. Token-level cross-entropy forces the model to reproduce the reference token by token, so local synthetic artifacts, such as: unnatural prosody, phonetic glitches, or imperfect RIR simulation, get absorbed into fluent but acoustically unsupported continuations, producing hallucinated insertions. RL inst","cbCaisBUK4gH7rSB","https://ap.wps.com/l/cbCaisBUK4gH7rSB","pdf",892150,2,1,7,"English","en",105,"# Introduction\n## Privacy and data constraints in regulated domains\n## Synthetic speech and the synthetic-to-real gap\n## Why supervised fine-tuning may cap performance\n# Method\n## LLM-based ASR architecture and experimental setup\n# Key questions\n## Effectiveness of GRPO vs SFT\n## Value of adding real data\n## Scaling with synthetic data\n## Reward function choices\n## Synthetic data subset selection\n## Why GRPO outperforms SFT","[{\"question\":\"Why is synthetic speech used for ASR in regulated domains?\",\"answer\":\"Because privacy regulations make real speech costly and tightly constrained to collect, store, share, and reuse for training. Synthetic TTS from available transcripts can replace hard-to-obtain real audio.\"},{\"question\":\"What does GRPO change compared with supervised fine-tuning (SFT) for synthetic speech adaptation?\",\"answer\":\"GRPO performs critic-free reinforcement learning by scoring a group of sampled hypotheses and rewarding those that outperform the group using task metrics like WER. This directly optimizes sequence-level outcomes rather than forcing token-by-token cross-entropy replication.\"},{\"question\":\"How much WER improvement does GRPO achieve with synthetic speech?\",\"answer\":\"Using synthetic-only adaptation with GRPO reduces WER by 40% relative to SFT (36.71%→22.09%). Combining SFT-then-GRPO improves this further to a 45% relative reduction.\"}]",1784200998,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"when-synthetic-speech-is-all-you-have-better-call-grpo","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/when-synthetic-speech-is-all-you-have-better-call-grpo/85086/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is synthetic speech used for ASR in regulated domains?","Question",{"text":75,"@type":76},"Because privacy regulations make real speech costly and tightly constrained to collect, store, share, and reuse for training. Synthetic TTS from available transcripts can replace hard-to-obtain real audio.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does GRPO change compared with supervised fine-tuning (SFT) for synthetic speech adaptation?",{"text":80,"@type":76},"GRPO performs critic-free reinforcement learning by scoring a group of sampled hypotheses and rewarding those that outperform the group using task metrics like WER. This directly optimizes sequence-level outcomes rather than forcing token-by-token cross-entropy replication.",{"name":82,"@type":73,"acceptedAnswer":83},"How much WER improvement does GRPO achieve with synthetic speech?",{"text":84,"@type":76},"Using synthetic-only adaptation with GRPO reduces WER by 40% relative to SFT (36.71%→22.09%). Combining SFT-then-GRPO improves this further to a 45% relative reduction.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]