[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84340-en":3,"doc-seo-84340-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84340,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Best-of-N TTS Evaluation is Confounded by ASR Family Alignment","Best-of-N (BoN) inference improves content consistency in zero-shot text-to-speech by selecting from N candidates using an automatic speech recognition (ASR) verifier. The work uncovers a key evaluation confound: verifier quality depends strongly on which ASR family performs judging. On LibriSpeech-PC testclean with F5-TTS, verifier rankings reverse across Whisper, wav2vec 2.0, and HuBERT, and same-family pairings yield more oracle headroom. Cross-family rank ensembles achieve the lowest mean WER across three evaluators and recommend cross-evaluator triangulation.","Best-of-N TTS Evaluation is Confounded by ASR Family Alignment  \nTaehyung Yu 1 Seongjae Kang 1  \narXiv :2607 .08256v 1 [ cs .CL] 9 Jul 2026  \nAbstract  \nBest-of-N (BoN) inference improves content consistency in zero-shot text-to-speech by selecting from N candidates with an automatic speech recognition (ASR) verifier. We identify an underexplored evaluation confound: a verifier’s apparent quality depends strongly on which ASR family judges it. On LibriSpeech-PC testclean (Meister et al., 2023) with F5-TTS (Chenet al., 2025), verifier rankings reverse across Whisper, wav2vec 2.0, and HuBERT evaluators, and same-family verifier–evaluator pairs recover 2– 3 × more oracle headroom than cross-family pairs despite near-identical representations (linear CKA 0.978)—a pattern consistent with identity- or lineage-level coupling rather than representational overlap. We propose two cross-family rank ensembles (rank-averaging and conjunctive maxrank) that attain the lowest mean WER across three independent evaluators—1 .61% at N=10 (−12% relative to F5-TTS)—with no measurable degradation under automatic SIM-o/UTMOS metrics; the best single verifier drives WER from 2.06% to 1.72%(−16 .5%) under the official F5-TTS evaluator. We recommend cross-evaluator triangulation as default reporting practice.  \n1. Introduction  \nRecent flow-matching zero-shot TTS systems—F5-TTS (Chen et al., 2025), E2 TTS (Eskimez et al., 2024), CosyVoice 2 (Du et al., 2024), MaskGCT (Wang et al., 2025), Seed-TTS (Anastassiou et al., 2024), NaturalSpeech 3 (Ju et al., 2024)—produce speech that is indistinguishable from human recordings in naturalness and speaker similarity, yet still produce word-level content errors on a non-trivial fraction of utterances. Best-of-N (BoN) inference is the standard inference-time remedy: synthesize N candidates with different random initializations  \n1 KAIST, Daejeon, South Korea. Correspondence to: Taehyung Yu \u003C[taehyung.yu@kaist.ac.kr](taehyung.yu@kaist.ac.kr) >.  \nAccepted at ICML 2026 Workshop on Machine Learning for Audio, Seoul, South Korea. Copyright 2026 by the author(s) .  \nand select the one whose ASR-decoded transcript matches the target text most closely.  \nBoN is reported to reduce WER by 10–30% relative across recent flow-matching systems, but the literature is inconsistent on which verifier ASR to use—reported choices include wav2vec 2.0 (Baevski et al., 2020), Whisper (Radford et al., 2023), and its distillations (Gandhi et al., 2023)—and the evaluation ASR is usually a single fixed model. The choice of evaluator is a design decision that has so far received no systematic attention.  \nIn this paper we run a four-way evaluator ablation spanning the Whisper (Radford et al., 2023), wav2vec 2.0 (Baevski et al., 2020), and HuBERT (Hsu et al., 2021) families (§3) over a shared set of BoN candidates. Our contributions are the following.  \n• We document that BoN verifier rankings are systematically evaluator-dependent: the same generated outputs are ranked in opposing directions by evaluators from different ASR families, so the verifier preferred under one family can lose under another (§4) .  \n• We quantify the magnitude of this confound on LibriSpeech-PC test-clean and show that the (verifier, evaluator) family pairing is a much larger lever on reported WER than the choice between common verifier checkpoints, while a large oracle headroom remainsunexploited (§4) .  \n• We rule out audio-encoder representation similarity (linear CKA) as the dominant explanation; the pattern is instead consistent with identity-or lineage-level coupling between verifier and evaluator, a speech analog of LLM-as-a-judge self-bias (§5) .  \n• We propose cross-family rank ensembles that select candidates by aggregating verifiers across ASR families, and recommend cross-evaluator triangulation—reporting WER under at least two ASR families with disjoint training lineages—as a default reporting practice (§4.4) .  \n2. Related Work  \nInference-time remedie","cbCaim9TrZJp5OiH","https://ap.wps.com/l/cbCaim9TrZJp5OiH","pdf",234313,3,1,6,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n# Setup","[{\"question\":\"What evaluation confound does the paper identify for Best-of-N (BoN) TTS reranking?\",\"answer\":\"The paper shows that a verifier’s apparent quality depends strongly on which ASR family evaluates it, causing verifier rankings to reverse across different ASR families.\"},{\"question\":\"How do verifier–evaluator ASR family pairings affect reported WER?\",\"answer\":\"Same-family verifier–evaluator pairs recover substantially more oracle headroom than cross-family pairs, making the (verifier, evaluator) family pairing a large driver of reported WER.\"},{\"question\":\"What methods and reporting practice does the paper recommend?\",\"answer\":\"It proposes cross-family rank ensembles (rank-averaging and conjunctive maxrank) to reduce mean WER and recommends cross-evaluator triangulation by reporting results under at least two ASR families with disjoint training lineages.\"}]",1784194930,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"best-of-n-tts-evaluation-is-confounded-by-asr-family-alignment","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/best-of-n-tts-evaluation-is-confounded-by-asr-family-alignment/84340/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What evaluation confound does the paper identify for Best-of-N (BoN) TTS reranking?","Question",{"text":75,"@type":76},"The paper shows that a verifier’s apparent quality depends strongly on which ASR family evaluates it, causing verifier rankings to reverse across different ASR families.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do verifier–evaluator ASR family pairings affect reported WER?",{"text":80,"@type":76},"Same-family verifier–evaluator pairs recover substantially more oracle headroom than cross-family pairs, making the (verifier, evaluator) family pairing a large driver of reported WER.",{"name":82,"@type":73,"acceptedAnswer":83},"What methods and reporting practice does the paper recommend?",{"text":84,"@type":76},"It proposes cross-family rank ensembles (rank-averaging and conjunctive maxrank) to reduce mean WER and recommends cross-evaluator triangulation by reporting results under at least two ASR families with disjoint training lineages.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]