[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85136-en":3,"doc-seo-85136-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85136,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Listen to the Features Voice Anonymization Driven by Content Embedding Matching over Signal Reconstruction","The paper introduces a voice anonymization model that prioritizes preserving speech content over producing naturalistic audio. It decodes content embeddings taken from a frozen pretrained wav2vec2 encoder into an anonymized signal via vector quantization and a HiFi-GAN vocoder. Both components are trained on LibriTTS without waveform reconstruction loss or speaker embedding mapping. Training enforces embedding matching between original and anonymized signals while an adversarial speaker classification branch with gradient reversal removes speaker-specific information. Results report very low ASR WER and strong privacy rankings, with partial emotion preservation and audible anonymized speech despite reconstruction-free training.","LISTEN TO THE FEATURES: VOICE ANONYMIZATION DRIVEN BY CONTENT EMBEDDING MATCHING OVER SIGNAL RECONSTRUCTION  \nAUTHOR VERSION  \nAdrien Schneider 1,, Kacper Zabkowski 1, Anderson Augusma 1,,  , Frédérique Letué2,,  ,  \nMaria Camila Pinzon 1, Dominique Vaufreydaz 1,  ,  \n1 Univ. Grenoble Alpes, CNRS, Grenoble INP, LIG, 38000 Grenoble, France  \n2 Univ. Grenoble Alpes, CNRS, Grenoble INP, LJK, 38000 Grenoble, France  \narXiv :2607 .09767v1 [ ee ss . SP] 7 Jul 2026  \nABSTRACT  \nThe paper presents a voice anonymization model focusing on preserving content rather than producing realistic speech. It relies on content embeddings extracted from a frozen pretrained wav2vec2 encoder. These embeddings are decoded into an anonymized signal using vector quantization anda HiFi-GAN vocoder, both trained on LibriTTS without any waveform reconstruction loss or speaker embedding mapping. The training objective enforces that embeddingsof the anonymized signal match those of the original one. While training, an auxiliary speaker classification branch with a gradient reversal layer is used to discard speakerspecific information. Results show that this straightforward embedding-based approach achieves very low WER (2.53) with an anonymization performance (EER 13.39) ranking within first level for VPC. Notably, emotions are partially preserved (UAR 43.91), even without a supporting training objective, while the anonymized voice is audible without reconstruction loss.  \nKeywords: voice anonymization, speech recognition, content embedding driven reconstruction  \n1 Introduction  \nReproducible research is important for accelerating scientific progress either in Artificial Intelligence, in Social Science research or in Computational Social Science. In these contexts, sharing data is as important as sharing source code but it becomes increasingly complex due to legal constraints, such as the GDPR and the AI Act in Europe, or the need to address ethical considerations. These rules are beyond discussion, as protecting the privacy of individuals is critical in the current numerical world but one must consider their impact on science discovery due to the limitations they impose on research data sharing.  \nAmong standard usages of privacy algorithms, anonymization of audio and video data is one way to leverage data sharing. Numerous high-value scientific datasets containing identifiable voices and faces are collected by researchers worldwide. If these data are sufficiently anonymized, they can be shared to the research community and thus can benefit  \nmany other researchers. To achieve this, anonymization must preserve the semantic content of speech, emotions, interpersonal interactions, and other social signals expressed in recorded videos, while maintaining the value of the collected data. A long-term objective for privacy-safe reproducible research would be to make available lightweight anonymization models capable of recording speech in an already anonymized form during corpus collection.  \nThis research on voice anonymization submitted to the Voice Privacy Challenge 2026 [1] is part of a broader project about social-aware video anonymization targeting both audio and video signals. In this first proposal, this research focuses on voice anonymization preserving what is said while evaluating preservation of voice emotions. The originality lies in the ability of the model to learn to generate an anonymized audio segment that, once processed by a given speech encoder, would give the same latent content representation than the original voice segment. There are no constraints nor losses within the training process to force the generation of a realistic, human-sounding speech signal. The motivation is that, with automated downstream tasks in mind, such as Automatic Speech Recognition (ASR) or Speech Emotion Recognition (SER), the anonymization model only needs to generate any signal that carries meaningful information, without expecting it to be clean speech. The pro","cbCaiiiwOdcm0BpV","https://ap.wps.com/l/cbCaiiiwOdcm0BpV","pdf",2504167,4,1,7,"English","en",105,"# Introduction\n## Motivation and privacy context\n## Voice Privacy Challenge 2026 tasks and metrics\n# Method Overview\n## Content-embedding driven anonymization\n## Adversarial speaker information removal","[{\"question\":\"What is the core idea behind the proposed voice anonymization model?\",\"answer\":\"The model generates an anonymized audio segment such that its content embeddings, produced by a given speech encoder, match those of the original voice segment.\"},{\"question\":\"How does training avoid reconstructing realistic speech while still preserving useful information?\",\"answer\":\"Training uses no waveform reconstruction loss and no speaker embedding mapping; instead, it enforces embedding-level matching between original and anonymized signals, while a gradient-reversal speaker classifier discards speaker-specific information.\"},{\"question\":\"Which evaluation tasks and metrics are used in the Voice Privacy Challenge 2026?\",\"answer\":\"Evaluation uses ASR and SER tasks in English, with Word Error Rate (WER) and Unweighted Average Recall (UAR) for performance, and Equal Error Rate (EER) for privacy under a semi-informed attacker using an ASV model fine-tuned on anonymized data.\"}]",1784201311,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"listen-to-the-features-voice-anonymization-driven-by-content-embedding-matching-over-signal-reconstruction","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/listen-to-the-features-voice-anonymization-driven-by-content-embedding-matching-over-signal-reconstruction/85136/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the core idea behind the proposed voice anonymization model?","Question",{"text":75,"@type":76},"The model generates an anonymized audio segment such that its content embeddings, produced by a given speech encoder, match those of the original voice segment.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does training avoid reconstructing realistic speech while still preserving useful information?",{"text":80,"@type":76},"Training uses no waveform reconstruction loss and no speaker embedding mapping; instead, it enforces embedding-level matching between original and anonymized signals, while a gradient-reversal speaker classifier discards speaker-specific information.",{"name":82,"@type":73,"acceptedAnswer":83},"Which evaluation tasks and metrics are used in the Voice Privacy Challenge 2026?",{"text":84,"@type":76},"Evaluation uses ASR and SER tasks in English, with Word Error Rate (WER) and Unweighted Average Recall (UAR) for performance, and Equal Error Rate (EER) for privacy under a semi-informed attacker using an ASV model fine-tuned on anonymized data.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]