[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85887-en":3,"doc-seo-85887-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85887,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","PC Mix Partial Component Audio Spoofing Detection under Mixed Speech and Environmental Sound Conditions","PC-Mix introduces partial-component audio spoofing detection under realistic mixed conditions where bonafide and spoofed segments co-exist across speech and environmental-sound components. The dataset constructs bonafide and partially spoofed environmental components and mixes them with speech signals from an existing partial-spoof dataset, yielding audio with localized manipulation in either or both components. Standardized evaluation protocols and a joint learning framework optimize frame-level detection across speech, environmental sound, and mixed audio. Results show mixed conditions increase difficulty and that matched-condition training outperforms direct transfer.","PC-Mix: Partial-Component Audio Spoofing Detection under Mixed Speech and Environmental  \nSound Conditions  \nZhenshan Zhang∗ , Xueping Zhang∗ , Linxi Li†, Yechen Wang†, Ming Li‡*  \n∗ Digital Innovation Research Center, Duke Kunshan University, Kunshan, China †OfSpectrum, Inc., Los Angeles, USA  \n‡School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China  \n* Corresponding author  \narXiv :2607 . 10345v 1 [ cs . SD] 11 Jul 2026  \nAbstract—Recent studies on partial audio spoofing mainly focus on studio-recorded speech with temporal localization of spoofed segments. However, these studies often overlook realistic conditions where spoofed and bonafide segments simultaneously coexist across speech and environmental sound components. In this paper, we present PC-Mix, the first dataset for partialcomponent spoofing detection, where either or both audio components may be partially spoofed. In PC-Mix, bonafide and partially spoofed environmental-sound components are first constructed and mixed with speech signals from an existing partial-spoof dataset, producing audio in which either or both components may be locally manipulated. This design addresses two major gaps in existing partial spoofing benchmarks: the lack of realistic environmental sounds in speech partial spoofing scenarios and the absence of partial spoofing detection for environmental sound components. We further establish standardized evaluation protocols and design a joint learning framework to optimize spoofing detection across speech, environmental sound, and mixed audio. Experiments highlight the increased difficulty introduced by mixed conditions. The results demonstrate that training under matched target conditions is more effective than directly transferring models trained on speech or environmental sound components.  \nIndex Terms—Audio Anti-Spoofing, Partial Spoofing Detection, Component-Level Audio Anti-Spoofing, Joint Learning  \nI. INTRODUCTION  \nAudio anti-spoofing has become an important task in trustworthy speech and audio processing, as increasingly realistic synthetic and manipulated audio continues to pose serious security risks to real-world applications [1]–[3] .  \nSpoofed audio can be generated or manipulated using increasingly powerful techniques, including voice conversion methods [4], [5], zero-shot speech synthesis methods [6], [7], and audio editing methods [8]–[10], posing significant threats to authentication systems and downstream applications.  \nMost research [11]–[14] only focuses on utterance-level spoof, ignoring localized manipulation within an utterance and treating the entire signal as either fully real or fully spoofed. More recently, attention has shifted toward partial spoofing, where only localized temporal regions of an utterance are manipulated while the rest remains bonafide [15], [16] . This scenario is more realistic, as modern audio editing tools enable fine-grained manipulation of selected segments without  \naltering the entire recording. Existing benchmarks [15], [17]–[21] for partial-spoofing and temporal deepfake localization have shown that localized attacks are particularly challenging to detect.  \nHowever, these studies predominantly focus on the speech component alone and typically assume clean or studiorecorded conditions, overlooking the role of environmental background sounds [11], [12], [22], [23] . In real-world scenarios, audio recordings are inherently composed of multiple acoustic components, including foreground speech and environmental sounds. Recent work has begun to explore component-level spoofing, where spoofing may target either the speech component, the environmental component, or both [24], [25] . This formulation reflects practical scenarios where adversaries may manipulate only selected audio components such as speech or environmental sound or possibly both. Such mixed conditions introduce additional complexity, as detection systems must reason about multiple interacting sou","cbCaieDZGKo5MW5E","https://ap.wps.com/l/cbCaieDZGKo5MW5E","pdf",433570,4,1,"English","en",105,"# Introduction\n## Audio anti-spoofing background and threat model\n## Shift from utterance-level to partial spoofing\n## Need for component-level and mixed-condition evaluation\n# Proposed task and dataset\n## Partial-component spoof evaluation protocol\n## PC-Mix dataset construction and mix generation\n## Original recordings for natural-vs-mixed discrimination\n# Joint learning framework and evaluation\n## Two-stage training strategy\n## Unified partial-component spoofing objective\n# Contributions","[{\"question\":\"What does PC-Mix consider as the core detection task?\",\"answer\":\"It targets partial-component audio spoofing detection, where localized spoofing may occur in the speech component, the environmental-sound component, or both within the same mixed acoustic scene.\"},{\"question\":\"How are the PC-Mix samples constructed?\",\"answer\":\"Bonafide and partially spoofed environmental-sound components are built and then mixed with speech signals from an existing partial-spoof dataset, producing mixtures where either or both components can contain local manipulations. Original recordings are also included for comparison.\"},{\"question\":\"Why do mixed speech-and-environmental conditions make the problem harder?\",\"answer\":\"Detection systems must jointly reason about multiple interacting sources and localized manipulations across components, rather than handling a single homogeneous signal.\"}]",1784206964,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"pc-mix-partial-component-audio-spoofing-detection-under-mixed-speech-and-environmental-sound-conditions","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":20},"https://docshare.wps.com/document/pc-mix-partial-component-audio-spoofing-detection-under-mixed-speech-and-environmental-sound-conditions/85887/",{"url":51,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What does PC-Mix consider as the core detection task?","Question",{"text":74,"@type":75},"It targets partial-component audio spoofing detection, where localized spoofing may occur in the speech component, the environmental-sound component, or both within the same mixed acoustic scene.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How are the PC-Mix samples constructed?",{"text":79,"@type":75},"Bonafide and partially spoofed environmental-sound components are built and then mixed with speech signals from an existing partial-spoof dataset, producing mixtures where either or both components can contain local manipulations. Original recordings are also included for comparison.",{"name":81,"@type":72,"acceptedAnswer":82},"Why do mixed speech-and-environmental conditions make the problem harder?",{"text":83,"@type":75},"Detection systems must jointly reason about multiple interacting sources and localized manipulations across components, rather than handling a single homogeneous signal.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]