[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84501-en":3,"doc-seo-84501-105":29,"detail-sidebar-cat-0-en-105":82},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84501,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","Multi-task Learning Is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition","Second-language (L2) speech recognition often requires systems to predict both what is pronounced and what is intended in canonical written form. Multi-task learning (MTL) is commonly used to share representations across outputs, yet controlled experiments on Korean and English show this assumption fails. MTL improves meaning transcription while degrading surface transcription, especially for English, with degradation increasing with surface-meaning divergence measured by Levenshtein edit distance. Encoder and decoder analyses attribute effects to encoder-level representational entanglement and constrained surface dual-output decoding, motivating entanglement-mitigating MTL designs for dual-output L2 ASR.","Multi-task Learning is Not Enough: Representational Entanglement in Dual-output Second Language Speech Recognition  \nSeung Hwan Cho 1 Young-Min Kim 1 2  \narXiv :2606 .06065v4 [ cs .CL] 11 Jul 2026  \nAbstract  \nSecond-language (L2) speech recognition often requires transcriptions of pronunciations and intended meanings. Multi-task learning (MTL) is a natural approach because it assumes that shared representations benefit both outputs. However, this paper shows that this assumption does not hold across Korean and English. MTL improves meaning but degrades surface transcription, especially in English, where the degradation scales with surface-meaning divergence measured by Levenshtein edit distance. Encoder analysis links these patterns to encoder-level entanglement, with Korean preserving disentangled representations while English produces nearly identical ones. Cross-output decoder analysis shows that the meaning dual-output decoder adapts with a unique representation, while the surface dualoutput decoder remains constrained by the encoder. These findings motivate the design of MTL frameworks that mitigate encoder-level entanglement to reduce surface degradation in dual-output L2 automatic speech recognition.  \n1. Introduction and Related Work  \nHuman speech exhibits systematic differences between what is actually pronounced (surface-level) and the canonical written form (meaning-oriented) of an utterance. These differences reflect language-specific phonological phenomena, such as coarticulation, phonological reduction, and liaison (Ernestus & Warner, 2011) . The phonological gap is more pronounced in second-language (L2) speech, where speakerspecific deviations are common (Munro, 2021) . Therefore, speech recognition systems for L2 speakers must recover  \n1Department of Industrial Data Engineering, Hanyang University, Seoul, South Korea 2 School of Interdisciplinary Industrial Studies, Hanyang University, Seoul, South Korea. Correspondence to: Young-Min Kim \u003C[yngmnkim@hanyang.ac.kr](yngmnkim@hanyang.ac.kr) >.  \nAccepted at the 43rd International Conference on Machine Learning Workshop on Machine Learning for Audio, Seoul, South Korea.  \n2026. Copyright 2026 by the author(s) .  \nboth transcription forms from a single acoustic signal to enable targeted feedback in language learning and pronunciation assessment applications (Eskenazi, 2009) .  \nMulti-task learning (MTL) is a natural approach for dualoutput (DO), encompassing auxiliary MTL where one task supports another and joint MTL where tasks are learned with equal importance (Ruder, 2017) . In automatic speech recognition (ASR), joint connectionist temporal classification (CTC)-attention training (Kim et al., 2017 ; Watanabe et al., 2017) and intermediate-layer CTC (Nozaki & Komatsu, 2021) use auxiliary CTC to improve alignment and training stability. Joint MTL approaches generate distinct target sequences from the same acoustic input. Examples include dual-decoder models for ASR and speech translation (Leet al., 2020) and unified diarization-separation-recognition systems (Shakeel et al., 2025) . However, the effectiveness of joint MTL for DO L2 ASR remains unexamined, despite the fact that the two outputs share linguistic content.  \nThis paper challenges the assumption that joint MTL benefits both outputs in DO ASR by conducting controlled experiments on Korean and English L2 speech. The results show that joint MTL produces asymmetric output tradeoffs that depend on language, with the source localized to encoder-level representational entanglement. Our contribution is twofold. First, we demonstrate the languagedependent behavior of joint MTL for DO L2 ASR. Second, we identify the underlying mechanisms through encoderand cross-output decoder analyses. These findings motivate the development of structured approaches to mitigate this entanglement.  \n2. Method  \nTo isolate the effect of joint training on shared representations, we compare single-output (SO) models, which are tr","cbCaicoyL79SPPYe","https://ap.wps.com/l/cbCaicoyL79SPPYe","pdf",446962,1,5,"English","en",105,"# Abstract\n# Introduction and Related Work\n# Method\n## Problem Formulation\n## Architecture and Training\n# Experiments","[{\"question\":\"What underlying mechanism is proposed to explain the tradeoff between outputs?\",\"answer\":\"The encoder-level analysis links the observed behavior to representational entanglement. Korean tends to preserve disentangled representations, while English produces nearly identical ones, constraining the surface dual-output decoder.\"}]",1784196126,13,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":77,"head_meta":79,"extra_data":81,"updated_unix":27},"multi-task-learning-is-not-enough-representational-entanglement-in-dual-output-second-language-speech-recognition","",{"@graph":35,"@context":76},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/multi-task-learning-is-not-enough-representational-entanglement-in-dual-output-second-language-speech-recognition/84501/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70],{"name":71,"@type":72,"acceptedAnswer":73},"What underlying mechanism is proposed to explain the tradeoff between outputs?","Question",{"text":74,"@type":75},"The encoder-level analysis links the observed behavior to representational entanglement. Korean tends to preserve disentangled representations, while English produces nearly identical ones, constraining the surface dual-output decoder.","Answer","https://schema.org",{"og:url":51,"og:type":78,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":80,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":83},[84,88,92,96,100,105,110,113,118,121,125],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Comic",60,"comic",{"id":101,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},6,"Technology",50,"technology",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":111,"slug":112},30,"research-report",{"id":114,"doc_module":4,"doc_module_name":45,"category_name":115,"show_sort_weight":116,"slug":117},9,"Religion & Spirituality",20,"religion-spirituality",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":119,"show_sort_weight":116,"slug":120},"World Cup","world-cup",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":122,"slug":124},10,"Lifestyle","lifestyle",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":21,"slug":128},19,"General","general"]