[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84796-en":3,"doc-seo-84796-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84796,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","Unified Audio Intelligence Without Regressing on Text Intelligence","Audio intelligence covers understanding, reasoning, and generation for both audio and speech. This work presents Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM derived from the text-only Nemotron-Cascade-2-30B-A3B. Audex uses a single Transformer decoder that encodes audio, projects it into the text embedding space, and treats text tokens and quantized audio output tokens uniformly. Training uses large-scale audio-text datasets plus multi-stage supervised learning, Cascade RL, and on-policy distillation. The model attains state-of-the-art audio understanding, ASR/translation, text-to-speech, audio generation, and speech-to-speech while preserving text reasoning, alignment, knowledge, long-context, and agentic capabilities with minimal or no regression, and releases checkpoints for research.","arXiv :2607 .05 196v2 [ cs .CL] 7 Jul 2026  \nUnified Audio Intelligence Without Regressing on Text Intelligence  \nZhifeng Kong∗ , Sang-gil Lee* , Jaehyeon Kim* , Boxin Wang* , Zihan Liu, Sungwon Kim, Yang Chen, Arushi Goel, Rajarshi Roy, Wenliang Dai, Zhuolin Yang, Yangyi Chen, Dongfu Jiang, Sreyan Ghosh, Tuomas Rintamaki, Andrew Tao, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, Wei Ping*†  \nAbstract  \nAudio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2- 30B-A3B, a strong text-only MoE LLM. Audex adopts a simple unified design with a single Transformer decoder: audio inputs are encoded and projected into the text embedding space, while text tokens and quantized audio output tokens are treated uniformly during generation. This architecture enables strong audio-text fusion, seamless multimodal generation, and compatibility with standard LLM training and inference infrastructure. For training, we meticulously curate audio-text datasets comprising 157.4B audio tokens and 320.5B text tokens. We apply multi-stage supervised training on these datasets, followed by text-only Cascade RL and multi-domain on-policy distillation. Audex delivers state-of-the-art audio understanding, speech recognition and translation, text-to-speech, audio generation, and speech-to-speech generation, while preserving very compelling reasoning, alignment, knowledge, long-context, and agentic capabilities of its text-only LLM backbone with marginal or no regression. We release the model checkpoints to facilitate open research.  \n Nemotron-Labs-Audex-30B-A3B: the audio-text to audio-text model.  \n Nemotron-Labs-Audex-2B: the smaller 2B model uses the same training recipe as the 30B-A3B model.  \n1. Introduction  \nAudio, including speech, music and environmental sound, is central to how humans perceive, communicate, and experience the world, from understanding the physical environment and engaging in spoken conversation to expressing emotions beyond written language and appreciating music. As such, audio intelligence, which involves understanding, reasoning about, and generating both audio and speech, is an indispensable modality for building artificial general intelligence (AGI) .  \nIn recent years, substantial efforts have been devoted to building audio LLMs that excel at audio and speech understanding, including Audio Flamingo (Ghosh et al., 2025b, 2026; Goel et al., 2025; Kong et al., 2024a), Qwen2-Audio (Chu et al., 2024), Kimi-Audio (Ding et al., 2025), Voxtral (Liu et al., 2025a), and Step-Audio 2 (Wu et al., 2025a) . Other works have explored unified LLMs capable of both audio understanding and generation, including Qwen-Omni (Team, 2026; Xu et al., 2025a,b), MiMo-Audio (Zhang et al., 2025a), and UALM (Tian et al., 2026) .  \nIt is worth noting that LLMs capable of handling multimodal inputs and outputs, including audio and vision, especially those supporting multimodal generation, often exhibit noticeable regressions on text benchmarks. For vision-only inputs, the NVLM study (Dai et al., 2024) documented this issue and addressed it by training on a mixture of high-quality text-only and vision-text data. For models with multimodal output capabilities, the issue is even more challenging. For example, even when the multimodal output domain is constraint to speech generation, Qwen3-Omni (Xu et al., 2025b) and Qwen3.5-Omni (Team, 2026) show degradation on key text  \n∗ Equal contribution, with authors listed in reverse alphabetical order by first name.  \n†Leads the effort. Correspondence to: {zkong, sanggill, jaehyeonk, boxinw, [wping}@nvidia.com](wping}@nvidia.com).  \nbenchmarks relative to their text-only counterparts – Qwen3 (Yang et al., 2025) and Qwen3.5 (Qwen-Team, 2026) . 1 Such degradation is concerning because much of a model’s intelligence, including reasoning ability, knowledge, an","cbCaimPaiqEAXzoe","https://ap.wps.com/l/cbCaimPaiqEAXzoe","pdf",687869,1,41,"English","en",105,"# Abstract\n# Introduction\n## Problem: multimodal regressions on text\n## Approach: unified audio-text model\n## Training pipeline and objectives","[{\"question\":\"What is Nemotron-Labs-Audex-30B-A3B (Audex) and what problem does it address?\",\"answer\":\"Audex is a unified audio-text LLM built on a strong text-only MoE backbone. It targets the common issue where multimodal models degrade performance on text benchmarks compared with text-only counterparts.\"},{\"question\":\"How does Audex unify audio and text during generation?\",\"answer\":\"Audex uses a single Transformer decoder. Audio is encoded and projected into the text embedding space, and generation treats text tokens and quantized audio output tokens uniformly.\"},{\"question\":\"What training methods are used to preserve text intelligence while adding audio capability?\",\"answer\":\"Training uses curated audio-text datasets with multi-stage supervised training, followed by text-only Cascade RL and multi-domain on-policy distillation, aiming for strong audio performance with minimal regression on text capabilities.\"}]",1784198306,103,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"unified-audio-intelligence-without-regressing-on-text-intelligence","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/unified-audio-intelligence-without-regressing-on-text-intelligence/84796/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is Nemotron-Labs-Audex-30B-A3B (Audex) and what problem does it address?","Question",{"text":75,"@type":76},"Audex is a unified audio-text LLM built on a strong text-only MoE backbone. It targets the common issue where multimodal models degrade performance on text benchmarks compared with text-only counterparts.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Audex unify audio and text during generation?",{"text":80,"@type":76},"Audex uses a single Transformer decoder. Audio is encoded and projected into the text embedding space, and generation treats text tokens and quantized audio output tokens uniformly.",{"name":82,"@type":73,"acceptedAnswer":83},"What training methods are used to preserve text intelligence while adding audio capability?",{"text":84,"@type":76},"Training uses curated audio-text datasets with multi-stage supervised training, followed by text-only Cascade RL and multi-domain on-policy distillation, aiming for strong audio performance with minimal regression on text capabilities.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]