[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85105-en":3,"doc-seo-85105-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85105,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","It Takes Few to TANGO A Quantized Distributed Model for Binaural Speech Enhancement","Neural network–based multichannel speech enhancement delivers strong results but is constrained by heavy computation and memory needs on resource-limited devices. The paper studies low-precision inference for TANGO, a hybrid distributed binaural system that combines neural mask estimation with spatial filtering. It evaluates post-training quantization and quantization-aware training for the neural parts, tracks how mask quantization errors propagate, and shows spatial filtering compensates for most errors, enabling a simplified MN-TANGO model with comparable final performance. Compact versions achieve 4.65 MMAC/s and 0.177 MB using INT8 quantization, ERB compression, and grouped recurrent layers.","It Takes Few to TANGO: A Quantized Distributed Model for Binaural Speech Enhancement  \nZahra Benslimane ∗†, Pierre Chouteau∗ , Martyna Poreba ∗ , Fabrice Auzanneau ∗ , Michal Szczepanski ∗ , Fabian Chersi ∗ , Romain Serizel †  \n∗ Universit Paris-Saclay, CEA, List, F-91120 Palaiseau, France  \n† Universit de Lorraine, CNRS, Inria, LORIA, F-54000 Nancy, France  \n[zahra-hafida.benslimane@cea.fr](zahra-hafida.benslimane@cea.fr)  \narXiv :2607 .08645v 1 [ cs . SD] 9 Jul 2026  \nAbstract—Neural network-based multichannel speech enhancement systems achieve strong enhancement performance, but their computational and memory requirements limit deployment on resource-constrained devices. This paper investigates lowprecision inference for TANGO, a hybrid distributed binaural speech enhancement system combining neural mask estimation with spatial filtering. We evaluate post-training quantization and quantization-aware training for the neural components, and analyze how quantization errors in the mask estimators propagate through the downstream spatial filtering stage. Our analysis shows that, although quantization degrades intermediate mask estimates, the spatial filtering stage compensates formost quantization-induced errors. Leveraging this robustness, we simplify TANGO into MN-TANGO, reducing both model size and computational complexity while maintaining comparable final performance. By combining INT8 weight-and-activation quantization with ERB compression and grouped recurrent layers, the most compact MN-TANGO reaches 4.65 MMAC/sand 0. 177 MB.  \nIndex Terms—Speech enhancement, quantization-aware training, recurrent neural networks, low-compute.  \nI. INTRODUCTION  \nDeep learning approaches to speech enhancement (SE) have achieved strong performance, but they often rely on large and computationally expensive models. This limits their deployment on resource-constrained devices, such as embedded systems and hearing aids, where low-latency and low-power inference are critical. To address this limitation, a growing body of work has investigated model compression for neural SE.  \nEarly studies mainly focused on reducing model storage through weight compression. Wu et al. [1] combined channel pruning with k-means clustering to quantize the weights of a time-domain fully convolutional network. Similarly, Tan and Wang [2] applied sparse regularization, iterative pruning, and k-means-based quantization to several architectures, including temporal convolutional networks, and gated convolutional recurrent networks (GCRNs) . Other works explored reduced floating-point representations. Hsu et al. [3] introduced the Exponent-Only Floating-Point Quantized Neural Network (EOFP-QNN), which quantizes the mantissa and exponent separately. Lin et al. [4] went further by discarding the mantissa entirely and retaining only the sign and exponent bits, achieving about 81% model compression.  \nWhile these methods reduce memory footprint, they primarily target weights. Activations, inputs, and outputs often remain in floating point, so inference may still require costly floating-point arithmetic. This limits efficiency on low-power hardware, such as microcontrollers and neural processing units, which are typically optimized for integer pipelines such as INT8 . To address this limitation, Fedorov et al. [5] proposed TinyLSTMs, combining structured pruning with quantizationaware training (QAT) [6] to quantize both weights and activations to 8 bits, while keeping the model outputs on 16-bit. Recent studies further showed that activation and I/O standard  \nQAT can be more challenging than weight quantization alone, especially at high input signal-to-noise ratios (SNRs) . To mitigate this effect,[7] introduced a residual correction branch to compensate for quantization errors. This approach was later extended to source separation using a knowledge-distillationbased loss for quantization-sensitive samples [8] .  \nDespite these advances, most quantization studies for SE fo","cbCaiscHKJuO7ruM","https://ap.wps.com/l/cbCaiscHKJuO7ruM","pdf",463939,1,7,"English","en",105,"# Introduction\n# Background\n## Baseline TANGO Architecture","[{\"question\":\"What problem does the paper address for binaural speech enhancement models?\",\"answer\":\"The paper targets the deployment limits of speech enhancement systems on resource-constrained devices, where computational and memory costs can be too high for low-latency, low-power inference.\"},{\"question\":\"How does quantization affect TANGO, and why does performance remain strong?\",\"answer\":\"Quantization degrades intermediate mask estimates, but the downstream spatial filtering stage compensates for most quantization-induced errors, preserving most enhancement gains.\"},{\"question\":\"What is MN-TANGO and how does it reduce complexity?\",\"answer\":\"MN-TANGO simplifies the original two-stage TANGO by keeping only the second spatial-filtering stage, reducing model size and computational complexity while maintaining comparable final performance.\"}]",1784201135,18,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"it-takes-few-to-tango-a-quantized-distributed-model-for-binaural-speech-enhancement","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/it-takes-few-to-tango-a-quantized-distributed-model-for-binaural-speech-enhancement/85105/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address for binaural speech enhancement models?","Question",{"text":75,"@type":76},"The paper targets the deployment limits of speech enhancement systems on resource-constrained devices, where computational and memory costs can be too high for low-latency, low-power inference.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does quantization affect TANGO, and why does performance remain strong?",{"text":80,"@type":76},"Quantization degrades intermediate mask estimates, but the downstream spatial filtering stage compensates for most quantization-induced errors, preserving most enhancement gains.",{"name":82,"@type":73,"acceptedAnswer":83},"What is MN-TANGO and how does it reduce complexity?",{"text":84,"@type":76},"MN-TANGO simplifies the original two-stage TANGO by keeping only the second spatial-filtering stage, reducing model size and computational complexity while maintaining comparable final performance.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]