[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85037-en":3,"doc-seo-85037-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85037,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents","Empirical reliability is assessed for Gemini models used as audio judges that score full-duplex agent conversations directly from raw stereo waveforms. Evidence is built on Gemini 2.5 Flash as a ground-truth reference, validated against three calibrated human raters over 209 stereo sessions. Ratings cover eight production dimensions across 152 conversations spanning accent-and-condition strata, plus 57 adversarial defect-injected clips, with agreement and sensitivity analyzed per dimension to support substitution where reliability holds.","arXiv :2607 .07985v 1 [ cs .CL] 8 Jul 2026  \nA Reliability Assessment of LALM Audio Judges for  \nFull-Duplex Voice Agents  \nA. Sayyad, J. Emmons, S. Jones, T. Lin, H. Krishnan  \nSalesforce Applied AI Research, e Verse team  \nAbstract  \nWe report the empirical reliability of Gemini models as audio judges that score full-duplex agent conversations directly from the raw stereo waveform, tested across three models in the Gemini family: 2.5 Flash, 3.5 Flash, and 3.1 Pro. Our primary evidence base uses Gemini 2.5 Flash as the ground-truth model, validated against three calibrated human raters on 209 stereo sessions, scored on 8 production dimensions: 152 full-duplex conversations across 13 accent-and-condition strata, together with 57 adversarial defect-injected clips. The evidence for Gemini 2.5 Flash is consistent across three tests. (i) On 5 of 8 dimensions the LALM-human Spearman ρ departs from the pairwise human-human ρ by at most 0.07, and on 7 of 8 dimensions the two quantities’ 95 percent bootstrap confidence intervals overlap.  \n(ii) The LALM agrees with the three-rater human mean within 1 point on 60 to  \n92 percent of sessions on 6 of 8 dimensions. (iii) On 45 of 48 (defect, dimension)  \ncells the LALM is as sensitive as humans or better under Newcombe-Wilson 95 percent confidence intervals, though most of these are underpowered nulls rather than demonstrated parity. Rank-ordering ability transfers across the Gemini family: 3.5 Flash improves simple agreement to 8 of 8 dimensions, while 3.1 Pro rates several dimensions markedly lower than humans despite comparable rank correlation. A model swap should be re-validated on calibration specifically, not assumed from rank-correlation alone. We identify four areas where deployment requires care, and we estimate that human rating alone for our current evaluation cadence costs roughly two orders of magnitude more than the equivalent LALM workload. The data presented here provides a defensible empirical basis for deploying the LALMas a substitute or fourth rater on the dimensions where the evidence supports it.  \nKeywords: LALM-as-judge, audio language models, validation, voice agents, production deployment, full-duplex audio.  \n1 Background  \nOur voice-agent product surface relies on two production audio judges, AgentSpeechFidelity and ConversationalAudioQuality, which between them emit ratings on eight dimensions for every evaluated session. To date these judges have been validated informally against single raters on small  \nsamples. This study addresses a single question: whether LALM-generated ratings agree with human ratings closely enough to substitute for one or more human raters in a production audioevaluation pipeline. We frame substitutability as an empirical, per-dimension property rather thana global one, and quantify it against a three-rater human reference standard.  \nRecent work in the LALM-as-judge literature reports strong correlations between LALM judges and human Mean Opinion Score (MOS) on isolated text-to-speech utterances [1, 2, 3] . To our knowledge, this is the first study to evaluate a LALM as a judge of raw audio from enterprise fullduplex agent conversations against a multi-rater human reference standard. We do so over stereo agent-client conversations where channel content, turn-taking dynamics, and conversational disfluencies all influence the rating distribution. We extend this line of work to full-duplex conversational audio using a production evaluation stack.  \n2 Related Work  \nLALM-as-judge for speech. AudioJudge [1] judges speech attributes (pronunciation, rate, quality) with high system-level human correlation; ALLD [2] produces descriptive MOS judgements; and SpeechQualityLLM [3] predicts dimension-wise MOS. Broad benchmarks such as AIR-Bench [10] score free-form outputs with a text LLM rather than validating an audio judge against humans. All operate on isolated utterances, and none reports multi-rater human validation over full-duplex convers","cbCaioYsjXKKQMf1","https://ap.wps.com/l/cbCaioYsjXKKQMf1","pdf",808807,1,28,"English","en",105,"# Background\n## Voice-agent audio judges\n## Study purpose\n# Related Work\n## LALM-as-judge for speech\n## Full-duplex evaluation\n## Reliability versus validity\n# Study Design\n## Dataset and dimensions\n## Human annotation protocol\n## LALM scoring approach","[{\"question\":\"What does the study evaluate about LALM audio judges?\",\"answer\":\"It evaluates whether LALM-generated ratings for full-duplex, raw stereo agent-client conversations agree with human ratings closely enough to substitute for one or more human raters in an evaluation pipeline.\"},{\"question\":\"How was reliability evidence collected and validated?\",\"answer\":\"Gemini 2.5 Flash is used as the ground-truth model and is validated against three calibrated human raters across 209 stereo sessions. Ratings are computed on eight production dimensions, including clean accent strata and codec-degraded conditions, plus adversarial defect-injected clips.\"},{\"question\":\"Does agreement transfer across different Gemini models?\",\"answer\":\"Rank-ordering ability transfers across the Gemini family, but model swaps require re-validation. The study reports different behavior across models: one model improves agreement across all dimensions, while another rates some dimensions markedly lower than humans.\"}]",1784200539,71,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"a-reliability-assessment-of-lalm-audio-judges-for-full-duplex-voice-agents","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/a-reliability-assessment-of-lalm-audio-judges-for-full-duplex-voice-agents/85037/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does the study evaluate about LALM audio judges?","Question",{"text":75,"@type":76},"It evaluates whether LALM-generated ratings for full-duplex, raw stereo agent-client conversations agree with human ratings closely enough to substitute for one or more human raters in an evaluation pipeline.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How was reliability evidence collected and validated?",{"text":80,"@type":76},"Gemini 2.5 Flash is used as the ground-truth model and is validated against three calibrated human raters across 209 stereo sessions. Ratings are computed on eight production dimensions, including clean accent strata and codec-degraded conditions, plus adversarial defect-injected clips.",{"name":82,"@type":73,"acceptedAnswer":83},"Does agreement transfer across different Gemini models?",{"text":84,"@type":76},"Rank-ordering ability transfers across the Gemini family, but model swaps require re-validation. The study reports different behavior across models: one model improves agreement across all dimensions, while another rates some dimensions markedly lower than humans.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]