[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82387-en":3,"doc-seo-82387-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82387,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","FreyaTTS Technical Report","FreyaTTS is a compact, tokenizer-free Turkish-first text-to-speech model for highly reliable and efficient conversational synthesis. It uses an 183.2M-parameter non-autoregressive conditional flow-matching Diffusion Transformer (DiT) operating in the frozen continuous-latent space of the frozen AudioVAE2, keeping 16 kHz encoding and 48 kHz decoding fixed. The framework enables rule-free end-to-end modeling from 92 Turkish characters without phonemizers, parallel denoising to avoid autoregressive error buildup, and production-hardening post-training for stable speaker identity and short-utterance coverage. On Freya-TR-Eval, it reports WER 8.0% and CER 3.0%, with real-time factor 0.11 on consumer GPUs and fast CPU deployment. Model weights, code, and benchmark are released under Apache-2.0.","arXiv :2607 .09530v 1 [ cs .CL] 10 Jul 2026  \nFreyaTTS Technical Report  \nFreya Team  \nJuly 13, 2026  \n Project: [https://github.com/freyavoiceai/FreyaTTS](https://github.com/freyavoiceai/FreyaTTS)  \n Model: [https://huggingface.co/freyavoice/Freya-TTS](https://huggingface.co/freyavoice/Freya-TTS)  \nAbstract  \nWe introduce FreyaTTS, a compact, tokenizer-free, Turkish-first text-to-speech model designed for highly reliable and efficient conversational synthesis. FreyaTTS is a 183. 2 M-parameter non-autoregressive conditional flow-matching Diffusion Transformer (DiT) that operates in the frozen continuous-latent space of the frozen AudioVAE2 (Zhou et al., 2026), reused unmodified (16 kHz encode, 48 kHz decode); holding the codec fixed lets the model devote all of its capacity to the text-to-latent map while inheriting 48 kHz reconstruction for free. We advance the framework across three key dimensions: (i) Rule-Free End-to-End Modeling, by driving generation end-to-end from a 92-symbol Turkish character vocabulary with no phonemizer, grapheme-to-phoneme frontend, or discrete speech tokenizer, so that agglutinative morphology, vowel harmony, and the spoken form of numbers and acronyms are learned directly from audio; (ii) Non-Autoregressive Parallel Denoising, which avoids the left-to-right error accumulation of autoregressive decoders by predicting and denoising the entire latent sequence in parallel over a predicted duration; and (iii) Production-Hardening Post-Training, utilizing a two-stage post-training recipe–a single-speaker voice lock that stabilizes speaker identity (collapsing cross-generation 􀀛0 standard deviation from 74. 9 Hz to 5. 0 Hz) followed by short-utterance coverage to resolve isolated-token and short-phrase failures. On our Freya-TR-Eval benchmark, FreyaTTS achieves a band-matched WER of 8. 0% and CER of 3. 0%, outperforming larger open-source systems at a fraction of their parameter size. With a real-time factor of 0.11 on consumer GPUs and the ability to run faster than real time on a laptop CPU, the model is highly optimized for resource-constrained edge deployment. We release the model weights, the training and inference code, and the evaluation benchmark under the Apache-2.0 license.  \n2  \nContents  \n1 Introduction 3  \n2 Related Work 4  \n2.1 Speech Synthesis Paradigms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4  \n2.2 Turkish, Multilingual, and Efficient TTS . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5  \n3 Methodology 6  \n3. 1 Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6  \n3.2 Frozen AudioVAE2 Latent Space . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7  \n3.3 The FreyaTTS Diffusion Transformer . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7  \n3.4 Pretraining from Scratch . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8  \n3.5 Post-training . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9  \n3.6 Inference . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 9  \n4 Experiments and Results 10  \n4.1 Experimental Setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 10  \n4.2 Main Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11  \n4.3 Speaker Consistency and the Voice-Lock . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12  \n4.4 Inference Efficiency . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12  \n5 Conclusion and Future Work 13  \n1 Introduction  \nText-to-speech (TTS) has advanced from producing merely intelligible speech toward generating natural, expressive, and controllable audio (Shen et al., 2018; Ren et al., 2020), driven by t","cbCaioNSVjzxGhHU","https://ap.wps.com/l/cbCaioNSVjzxGhHU","pdf",530585,3,1,16,"English","en",105,"# Contents\n## Introduction\n## Related Work\n## Methodology\n## Experiments and Results\n## Conclusion and Future Work","[{\"question\":\"What is FreyaTTS designed to achieve?\",\"answer\":\"FreyaTTS targets highly reliable and efficient conversational text-to-speech, with strong performance in Turkish while remaining compact for resource-constrained deployment.\"},{\"question\":\"How does FreyaTTS handle text without phonemizers or discrete speech tokens?\",\"answer\":\"It performs rule-free end-to-end modeling directly from a 92-symbol Turkish character vocabulary, learning agglutinative morphology, vowel harmony, and spoken forms of numbers and acronyms from audio.\"},{\"question\":\"What techniques improve synthesis quality and robustness after training?\",\"answer\":\"It uses a production-hardening two-stage post-training approach: a single-speaker voice lock to stabilize speaker identity and short-utterance coverage to resolve failures on isolated tokens and brief phrases.\"}]",1784180076,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"freyatts-technical-report","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/freyatts-technical-report/82387/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is FreyaTTS designed to achieve?","Question",{"text":75,"@type":76},"FreyaTTS targets highly reliable and efficient conversational text-to-speech, with strong performance in Turkish while remaining compact for resource-constrained deployment.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does FreyaTTS handle text without phonemizers or discrete speech tokens?",{"text":80,"@type":76},"It performs rule-free end-to-end modeling directly from a 92-symbol Turkish character vocabulary, learning agglutinative morphology, vowel harmony, and spoken forms of numbers and acronyms from audio.",{"name":82,"@type":73,"acceptedAnswer":83},"What techniques improve synthesis quality and robustness after training?",{"text":84,"@type":76},"It uses a production-hardening two-stage post-training approach: a single-speaker voice lock to stabilize speaker identity and short-utterance coverage to resolve failures on isolated tokens and brief phrases.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]