[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84877-en":3,"doc-seo-84877-105":28,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":13,"seo_description":14,"update_tm":26,"read_time":27},84877,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models","Audio-language models compress speech-encoder hidden states via a Querying Transformer (Q-Former) connector before passing tokens to a large language model. The work exposes two compression failures: query collapse where connector output vectors align to a single direction, and information loss where different speakers yield nearly indistinguishable outputs, erasing paralinguistic cues like identity, gender, and prosody. ORCA splits queries into orthogonal groups to prevent collapse, improving SAKURA multi-hop reasoning by 26.4 points to 75.2% and cutting query redundancy 12× while boosting cross-speaker variance 75×.","Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models  \nHo-Lam Chung 1,2 , Ke-Han Lu 1 , Yi-Cheng Lin 1 , Guan-Ting Lin 1 , Yiming Chen2 , Hung-yi Lee 1  \n1 Graduate Institute of Communication Engineering, National Taiwan University, Taipei, Taiwan  \n2ASUS AICS, Taipei, Taiwan  \narXiv :2607 .060 14v 1 [ cs . SD] 7 Jul 2026  \nAbstract—Audio-language models compress a speech encoder’s output through a Querying Transformer (Q-Former) connector before feeding it to a large language model. We identify two failures in this compression. The connector’s output vectors collapse to a single direction, and different speakers produce nearly indistinguishable outputs, with paralinguistic cues such as speaker identity, gender, and prosody lost along the way. Our method, ORCA, reverses this collapse by splitting the queries into groups whose outputs are constrained to point in different directions. On SAKURA multi-hop reasoning, ORCA gains 26.4 points over an identically trained 4B baseline, reaching 75.2%(vs. 49.0% for the 8B Audio Flamingo-3). At the connector level, the same change cuts query redundancy by 12 × and raises crossspeaker variance by 75 ×.  \nIndex Terms—audio-language models, paralinguistic reasoning, connector design, orthogonal subspaces.  \nI. INTRODUCTION  \nAn audio-language model (ALM) extends a large language model (LLM) to accept audio input alongside text. The standard pipeline has three components: an audio encoder (e.g., Whisper [1]) that maps audio to a sequence of hidden states, a connector that compresses this sequence into a fixed number of tokens, and an LLM (e.g., Llama [2], Qwen [3]) that reasons over the combined audio–text prefix. A widely used connector design is the Q-Former [4], in which K learnable queries cross-attend to the encoder output and produce a Ktoken summary for the LLM [5]–[9] . This makes the connector an information bottleneck: it determines the upper bound of audio information available to all downstream modules, and what it discards, no amount of LLM adaptation can recover. In practice, this compression fails to preserve much of the paralinguistic information in the signal. Speech carries semantic content (what is said) and paralinguistic cues such as speaker identity, gender, and prosody (how it is said) .  \nRecent benchmarks show that ALMs fall well below human performance on paralinguistic reasoning [10]–[13], and controlled experiments reveal a systematic preference for textual over acoustic cues [14], [15] . Upstream of the connector, the encoder’s intermediate representations retain measurable paralinguistic information [16] . Yet in §III, we show that the connector’s output is highly redundant, with very little variation across speakers, pointing to the compression itself as a key site of loss.  \nThe shape of this loss has a vivid analogue in the figure of Procrustes from Greek myth. Procrustes was an innkeeper with a single iron bed, and he forced every visitor onto it: those too tall were trimmed, those too short were stretched, so all guests left the same length as the bed. In the ALM pipeline, the bed is the language-modeling objective: it rewards only what helps the LLM predict text. The K-token bottleneck makes the fit tighter still, because with only K vectors to represent the full signal, there is no spare room for information the loss does not reward. Acoustic variation that aligns with text prediction fits the bed and survives; variation that does not is cut away. The result is a connector whose output is shaped by one objective alone, leaving non-semantic acoustic detail nowhere to live.  \nWe identify two measurable symptoms of this Procrustean compression. The first is query collapse: the connector’s K output vectors converge to near-identical directions. The second is information loss: different speakers reading the same sentence produce nearly interchangeable connector outputs. We argue that query collapse is the upstream cause. When ","cbCaia7ffDU5ZvjT","https://ap.wps.com/l/cbCaia7ffDU5ZvjT","pdf",556489,1,"English","en",105,"# Introduction\n## Procrustean compression analogy\n## Symptoms: query collapse and information loss\n## ORCA approach\n# Related Work","[{\"question\":\"What problem does the Q-Former connector cause in audio-language models?\",\"answer\":\"The connector can collapse its query outputs into a single direction and compress away speaker-dependent paralinguistic information, making different speakers’ outputs nearly indistinguishable.\"},{\"question\":\"How does ORCA (Groupwise Orthogonal Connector) address query collapse?\",\"answer\":\"ORCA partitions the K queries into G groups and constrains group outputs to point in different directions using a learned orthogonality regularizer, recovering separation without attribute-specific supervision.\"},{\"question\":\"What improvements does ORCA show on downstream tasks and connector-level metrics?\",\"answer\":\"On SAKURA multi-hop reasoning, ORCA reaches 75.2%, a +26.4 point gain over an identically trained 4B baseline. At the connector level, it reduces query redundancy by 12× and increases cross-speaker variance by 75×.\"}]",1784198963,20,{"code":4,"msg":29,"data":30},"ok",{"site_id":23,"language":22,"slug":31,"title":13,"keywords":32,"description":14,"schema_data":33,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":26},"escaping-the-procrustean-bed-groupwise-orthogonal-connectors-for-audio-language-models","",{"@graph":34,"@context":84},[35,52,67],{"@type":36,"itemListElement":37},"BreadcrumbList",[38,42,46,49],{"item":39,"name":40,"@type":41,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":43,"name":44,"@type":41,"position":45},"https://docshare.wps.com/document/","Document",2,{"item":47,"name":12,"@type":41,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":41,"position":51},"https://docshare.wps.com/document/escaping-the-procrustean-bed-groupwise-orthogonal-connectors-for-audio-language-models/84877/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":39,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does the Q-Former connector cause in audio-language models?","Question",{"text":74,"@type":75},"The connector can collapse its query outputs into a single direction and compress away speaker-dependent paralinguistic information, making different speakers’ outputs nearly indistinguishable.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does ORCA (Groupwise Orthogonal Connector) address query collapse?",{"text":79,"@type":75},"ORCA partitions the K queries into G groups and constrains group outputs to point in different directions using a learned orthogonality regularizer, recovering separation without attribute-specific supervision.",{"name":81,"@type":72,"acceptedAnswer":82},"What improvements does ORCA show on downstream tasks and connector-level metrics?",{"text":83,"@type":75},"On SAKURA multi-hop reasoning, ORCA reaches 75.2%, a +26.4 point gain over an identically trained 4B baseline. At the connector level, it reduces query redundancy by 12× and increases cross-speaker variance by 75×.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":44,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":45,"doc_module":4,"doc_module_name":44,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":44,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":44,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":44,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":44,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":44,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":44,"category_name":124,"show_sort_weight":27,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":27,"doc_module":4,"doc_module_name":44,"category_name":127,"show_sort_weight":27,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":44,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":44,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]