[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82805-en":3,"doc-seo-82805-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82805,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","Doppelganger Sound Effects and Their Synthetic Twins","Audio-conditioned generators produce synthetic sound effects alongside real recordings, yet existing benchmarks rarely test whether an embedding can match a specific synthetic clip to the exact real source it was generated from. Doppelganger is introduced as a benchmark for cross-domain sound-effect correspondence, using 10,420 real clips across 34 everyday events with audio-conditioned synthetic twins, plus a controlled 7-class corpus. Standard encoders fail across the boundary, class-supervised invariance generalizes poorly, and pair-based training recovers the true source ~80% on held-out events. Matching is generator-specific and not preserved by unrelated generators.","arXiv :2607 .04337v 1 [ cs . SD] 5 Jul 2026  \nDOPPELGANGER Sound Effects and Their Synthetic Twins  \nElliott Ash  \nETH Zürich  \n[ashe@ethz.ch](ashe@ethz.ch)  \n[https://github.com/elliottash/doppelganger](https://github.com/elliottash/doppelganger)  \nAbstract  \nAudio-conditioned generators now produce synthetic sound effects from real recordings, so the real and synthetic versions of an event increasingly coexist in sound libraries and in the corpora used to train audio models—yet no benchmark measures whether a representation can match a synthetic clip to the specific real recording it was generated from. I introduce Doppelganger, a benchmark for matching sound effects across the synthetic–real boundary, pairing 10 ,420 real clips across 34 everyday sound events each with an audio-conditioned synthetic twin, alongside a controlled 7-class corpus. Off-the-shelf audio encoders do not cross the boundary cleanly. Making the embedding ignore the boundary the standard way—training it on sound-event labels—works on familiar sounds but backfires on new ones, dropping below the untrained encoder. Training on the pairs instead—a clip and its own synthetic twin—generalizes. On sound events held out of training, it recovers the exact real source about 80% of the time (up from 61% untrained; chance 0.03%), whereas no objective meaningfully improves category-level recognition on those unseen events. The learned matching is specific to one generator—it survives changes to that generator’s settings but not a switch to a different generator, and collapses for the text-only generators tested. A human annotation baseline (49 listeners) lands well above chance but below the models on the same trials.  \nSynthetic twins fool people into calling them real about 29% of the time, yet a generator-specific detector separates these audio-conditioned twins from real recordings perfectly.  \n1 Introduction  \nAudio-conditioned generators now turn real sound effects into synthetic variants, so the real and synthetic versions of an event increasingly coexist—side by side in sound libraries, and mixed together in the web-scale corpora used to train audio models. This raises questions the field’s usual tools do not answer—questions not about whether two clips share a sound event, but about whether one specific synthetic clip is the counterpart of one specific real recording. Single-domain embedding benchmarks [Turian et al., 2022, Heigold et al., 2026] score whether an embedding groups similarsounding clips, but treat all audio as one domain. Distributional generative metrics such as Fréchet Audio Distance [Kilgour et al., 2019] ask whether a generator’s output distribution resembles real audio, not whether a given output preserved the event it was conditioned on. And synthetic-audio detection [Ouajdi et al., 2024, Yin et al., 2025] asks only real-or-fake, discarding the correspondence between a synthetic clip and its source.  \nI argue these are facets of one capability—representing the identity of a sound event across the synthetic–real boundary (Figure 1)—and pose it as retrieval:  \nFigure 1: The synthetic–real matching task and the instance-versus-category comparison on unseen sound  \nevents.  \nNotes. Left: each real recording is paired with an audio-conditioned synthetic twin of the same event; I ask whether an embedding can represent the event identity while ignoring how it was rendered. Middle: an invariant embedding mixes the two domains within each event cluster, while a sensitive one splits them—the two axes this benchmark separates. Right: on sound types never seen during training, an embedding trained to match each clip to its own synthetic twin (red) finds the exact real twin far more often than the untrained baseline (dotted line), while no training meaningfully improves retrieval of other clips of the same kind. Matching a specific sound to its twin transfers to new sound types; recognizing the broad sound event does not.  \nCan a representation rec","cbCailnd0kjqmgKU","https://ap.wps.com/l/cbCailnd0kjqmgKU","pdf",1009338,2,1,19,"English","en",105,"# Introduction\n## The synthetic–real matching task\n## Transfer as the objective\n## What the work enables","[{\"question\":\"What problem does the Doppelganger benchmark address?\",\"answer\":\"It evaluates whether a representation can match a specific synthetic sound effect to its exact corresponding real recording across the synthetic–real boundary, not just whether clips share a general sound category.\"},{\"question\":\"How is Doppelganger constructed in terms of data?\",\"answer\":\"It pairs 10,420 real clips across 34 everyday sound events with audio-conditioned synthetic twins, and includes a controlled 7-class corpus for additional evaluation.\"},{\"question\":\"What training approach works best for cross-boundary matching?\",\"answer\":\"Training on corresponding pairs—each real clip matched to its own synthetic twin—generalizes better than training an embedding to be invariant via sound-event labels, which degrades performance on unseen events.\"}]",1784183060,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"doppelganger-sound-effects-and-their-synthetic-twins","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/doppelganger-sound-effects-and-their-synthetic-twins/82805/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the Doppelganger benchmark address?","Question",{"text":75,"@type":76},"It evaluates whether a representation can match a specific synthetic sound effect to its exact corresponding real recording across the synthetic–real boundary, not just whether clips share a general sound category.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is Doppelganger constructed in terms of data?",{"text":80,"@type":76},"It pairs 10,420 real clips across 34 everyday sound events with audio-conditioned synthetic twins, and includes a controlled 7-class corpus for additional evaluation.",{"name":82,"@type":73,"acceptedAnswer":83},"What training approach works best for cross-boundary matching?",{"text":84,"@type":76},"Training on corresponding pairs—each real clip matched to its own synthetic twin—generalizes better than training an embedding to be invariant via sound-event labels, which degrades performance on unseen events.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]