[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84934-en":3,"doc-seo-84934-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84934,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts","Recognizing sentiment polarity, positive or negative, from speech is difficult because it depends on both vocal characteristics and the meaning of spoken words. A multimodal approach is introduced by combining audio with text information: transcripts are generated automatically using ASR, then translated into multiple languages using machine translation. Audio and multilingual text are fused through cascaded cross-modal transformer blocks. Knowledge distillation transfers the multimodal teacher into an audio-only student, improving performance without added inference cost. Experiments and ablations confirm that both automatic transcripts and translations help.","arXiv :2607 .066 1 1v 1 [ cs .CL] 7 Jul 2026  \nAudio Sentiment Analysis via Distillation and Cross-Modal Integration of  \nGenerated Multilingual Transcripts  \nAndrei-George Durduna , Victor Constantinescub , Radu Tudor Ionescua,b,∗  \na Department of Computer Science, University of Bucharest, 15G Iuliu Maniu, Bucharest 061075, Romania b Department of Data Science, PPC Romania, 30 Mircea Voda, Bucharest 030667, Romania  \nAbstract  \nAutomatically recognizing the sentiment, positive or negative, from speech is a challenging task, requiring both the analysis of vocal inflections and the interpretation of uttered words. Recent solutions rely on audio foundation models to solve the task, but it remains unclear if such models can take all aspects into account. To this end, we propose amultimodal solution that integrates audio and text information via cross-modal transformers, where text transcripts are automatically generated via an automatic speech recognition (ASR) tool. Moreover, we create multiple text modalities by automatically translating the transcripts into multiple languages via machine translation tools. Audio and multilingual text features are combined via a cascaded architecture comprising cross-modal transformer blocks that integrate modalities one by one. We further distill knowledge from the multimodal model, called teacher, into a unimodal (audio only) model, called student. We conduct experiments on a large-scale dataset, demonstrating that the automatically generated textual information can bring significant performance boosts in multimodal sentiment polarity classification. Our ablation study confirms that both automatic transcripts and automatic translations are helpful. Moreover, we show that the audio-only model can be enhanced via distillation, boosting performance without any computational overhead during inference. To reproduce the reported results, we publicly release our code at [https://github.com/andreidurdun/cross-modal-audio-sentiment](https://github.com/andreidurdun/cross-modal-audio-sentiment).  \nKeywords: multimodal learning; audio sentiment analysis; audio polarity classification; cross-modal transformer  \n1. Introduction  \nSentiment polarity classification from speech is an actively studied task [2, 11, 28, 31, 44], having a broad range of real-world applications, such as customer service call analysis [5], virtual assistant adaptation [1], mental health monitoring [42], vehicle driver monitoring [52] and gaming experience adaptation [25], among others. As for most speech processing tasks nowadays, deep learning models [14, 18, 37, 41], especially audio foundation models [3, 9, 37, 38], represent the mainstream solution for polarity classification, due to their typically high accuracy levels. However, the task remains challenging, as it involves the concurrent analysis of several aspects, including tone (which transmits emotion), pitch (which conveys excitement or calmness), word usage (which reflects the polarity of the transmitted message), irony / sarcasm (which may often indicate the opposite opinion than the transmitted message) . Disentangling these aspects is a key pathway towards improving performance, yet this area remains largely underexplored in current literature.  \nTo this end, we propose a novel knowledge distillation (KD) pipeline [15], where the teacher model benefits from disentangled audio and text information. While the audio modality is readily available, the text transcript is not directly  \n∗ Corresponding author.  \nE-mail address: [radu.ionescu@fmi.unibuc.ro](radu.ionescu@fmi.unibuc.ro)  \nPreprint submitted to Elsevier July 9, 2026  \n\n|  | Audio modality\u003Cbr>Student |  |\n| --- | --- | --- |\n\n|  | Tea(her |  |\n| --- | --- | --- |\n| \u003Cbr>Automati( Spee(h Re(ognition\u003Cbr>\u003Cbr>\u003Cbr> |  |  |\n|  |  |  |\n\nRoBERTa  \nRoBERTuito  \nCross-Modal Transformer  \nClassi'(ation Head  \nKnowledge distillation  \nCross-Modal Transformer  \nCross-Modal Transformer  \n\n| [CLS] |\n| --- |\n|  |\n|  |\n\n\n| [CLS]","cbCaipCIoi5jciQL","https://ap.wps.com/l/cbCaipCIoi5jciQL","pdf",375342,1,10,"English","en",105,"# Introduction\n## Problem setting and motivation\n## Proposed multimodal pipeline\n## Knowledge distillation and evaluation context","[{\"question\":\"How are text inputs obtained for the multimodal sentiment model?\",\"answer\":\"Text transcripts are generated automatically from speech using an ASR system. The transcripts are then translated into multiple languages with machine translation to create additional text modalities.\"},{\"question\":\"How are audio and multilingual text modalities combined in the proposed system?\",\"answer\":\"The approach uses cascaded cross-modal transformer blocks that integrate modalities one by one until a joint representation is formed.\"},{\"question\":\"What does the knowledge distillation step achieve for the final audio-only model?\",\"answer\":\"A multimodal teacher model transfers knowledge to an audio-only student model. This improves audio-only sentiment polarity performance without increasing computational overhead during inference.\"}]",1784199482,25,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"audio-sentiment-analysis-via-distillation-and-cross-modal-integration-of-generated-multilingual-transcripts","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/audio-sentiment-analysis-via-distillation-and-cross-modal-integration-of-generated-multilingual-transcripts/84934/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How are text inputs obtained for the multimodal sentiment model?","Question",{"text":75,"@type":76},"Text transcripts are generated automatically from speech using an ASR system. The transcripts are then translated into multiple languages with machine translation to create additional text modalities.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are audio and multilingual text modalities combined in the proposed system?",{"text":80,"@type":76},"The approach uses cascaded cross-modal transformer blocks that integrate modalities one by one until a joint representation is formed.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the knowledge distillation step achieve for the final audio-only model?",{"text":84,"@type":76},"A multimodal teacher model transfers knowledge to an audio-only student model. This improves audio-only sentiment polarity performance without increasing computational overhead during inference.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]