[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83758-en":3,"doc-seo-83758-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83758,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","TokAN Accent Normalization Using Self Supervised Speech Tokens","Accent normalization (AN) converts nonnative (L2) accented speech into standard (L1) speech while preserving speaker identity. Prior approaches either demand naturally recorded parallel L1–L2 data or degrade quality when trained with synthesized supervision. TokAN introduces a token-based framework using self-supervised discrete speech tokens from a jointly trained L1–L2 VQ tokenizer, enabling token-to-token conversion without synthetic supervisory speech. A GRPO reinforcement learning stage and a flow-matching synthesizer improve accent reduction and intelligibility, with duration-aware control for dubbing and live casting.","TokAN: Accent Normalization Using Self-Supervised Speech Tokens  \nQibing Bai, Student Member, IEEE, Shuai Wang, Senior Member, IEEE, Yuhan Du, Bohan Li,  \nYannan Wang, and Haizhou Li, Fellow, IEEE  \narXiv :2607 .03928v 1 [ cs . SD] 4 Jul 2026  \nAbstract—Accent normalization (AN) seeks to convert nonnative (L2) accented speech into standard (L1) speech while preserving speaker identity. The current techniques either require naturally recorded parallel L1–L2 speech for training, or suffer from quality degradation when supervised by synthesized targets. In this paper, we present TokAN, a token-based accent normalization framework that operates on self-supervised discrete speech tokens extracted from a L1–L2 jointly trained vector-quantization (VQ) tokenizer, without the need of synthetic supervisory speech. An autoregressive encoder-decoder model performs token-to-token conversion, translating L2-accented token sequences into the tokens of standard voice. We also introduce reinforcement learning (RL) post-training based on Group Relative Policy Optimization (GRPO), using word error rate and accent classifier confidence as complementary rewards. A non-autoregressive flow-matching synthesizer recovers the Melspectrogram from the converted tokens, conditioned on the source speaker embedding. We also develop a flow-matching duration predictor that supports total-duration-aware synthesis, making TokAN applicable to duration-critical tasks such as voice dubbing and live casting. Experiments on seven English accents demonstrate that TokAN reduced the word error rate from 12.40% to 9.89% after supervised fine-tuning, and further to 9.23% after RL post-training, consistently outperforming frame-toframe, direct flow-matching, and prompt-based token-conversion baselines in terms of accent reduction and intelligibility.  \nIndex Terms—Accent conversion, discrete speech tokens, vector quantization, duration control, reinforcement learning.  \nI. INTRODUCTION  \nACCENT conversion (AC) seeks to alter speech from one  \naccent to another while preserving the speaker’s characteristics. A particularly important special case is accent normalization (AN), also referred to as foreign accent conversion (FAC) [1], which converts non-native (L2) accented speech into a native (L1) accented form. AN technology enables a wide range of applications, including pronunciation training for language learners [2], authentic multimedia dubbing [3], and personalized text-to-speech systems [4] .  \n(Corresponding authors: Shuai Wang and Haizhou Li.)  \nQ. Bai is with the School of Data Science (SDS), The Chinese University of Hong Kong, Shenzhen (CUHKSZ), China, and with Tencent Ethereal Audio Lab, Tencent, Shenzhen, China.  \nS. Wang is with the School of Intelligence Science and Technology, Nanjing University, Suzhou, China, and with Shenzhen Loop Area Institute, Shenzhen, China.  \nY. Du is with the School of Intelligence Science and Technology, Nanjing University, Suzhou, China.  \nB. Li is with the X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University, Shanghai, China.  \nH. Li is with School of Artificial Intelligence (SAI), CUHKSZ, China, with Shenzhen Research Institute of Big Data, China, and with Shenzhen Loop Area Institute, China.  \nY. Wang is with Tencent Ethereal Audio Lab, Tencent, China. Samples: [https://p1ping.github.io/TokAN-Samples](https://p1ping.github.io/TokAN-Samples).  \nCode: [https://github.com/P1ping/TokAN](https://github.com/P1ping/TokAN).  \nEarly deep learning approaches for AN are referencebased [1], [5]–[7], relying on native accent speech samples to generate accent-neutral representations via phonetic posteriorgram (PPG) features [1], [5], [7] or native text-to-speech (TTS) [6] . VEVO [8] leverages speech tokens and accent prompts for prompt-based accent style transfer. However, reference or prompt speech requirements at inference limit deployment in practice.  \nReference-free methods [9]–[11] eliminate this requirement by dire","cbCailP99WXZgDM7","https://ap.wps.com/l/cbCailP99WXZgDM7","pdf",1186690,4,1,13,"English","en",105,"# Introduction\n## Accent conversion and accent normalization background\n## Reference-based and reference-free methods\n## Limitations of supervised synthetic targets and parallel corpora\n## Motivation for self-supervised discrete speech tokens","[{\"question\":\"What problem does TokAN address in speech technology?\",\"answer\":\"TokAN targets accent normalization by transforming L2 accented speech into an L1 standard accent while preserving the speaker’s identity.\"},{\"question\":\"How does TokAN avoid relying on synthetic supervisory speech?\",\"answer\":\"TokAN operates on self-supervised discrete speech tokens extracted from a jointly trained L1–L2 VQ tokenizer, and uses token-to-token conversion with no need for synthetic supervisory speech.\"},{\"question\":\"How is TokAN improved after initial training?\",\"answer\":\"TokAN applies reinforcement learning post-training using Group Relative Policy Optimization (GRPO), combining word error rate and accent classifier confidence as complementary rewards.\"}]",1784190249,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"tokan-accent-normalization-using-self-supervised-speech-tokens","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/tokan-accent-normalization-using-self-supervised-speech-tokens/83758/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does TokAN address in speech technology?","Question",{"text":75,"@type":76},"TokAN targets accent normalization by transforming L2 accented speech into an L1 standard accent while preserving the speaker’s identity.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does TokAN avoid relying on synthetic supervisory speech?",{"text":80,"@type":76},"TokAN operates on self-supervised discrete speech tokens extracted from a jointly trained L1–L2 VQ tokenizer, and uses token-to-token conversion with no need for synthetic supervisory speech.",{"name":82,"@type":73,"acceptedAnswer":83},"How is TokAN improved after initial training?",{"text":84,"@type":76},"TokAN applies reinforcement learning post-training using Group Relative Policy Optimization (GRPO), combining word error rate and accent classifier confidence as complementary rewards.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]