[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83708-en":3,"doc-seo-83708-105":29,"detail-sidebar-cat-0-en-105":82},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83708,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","Taste-aware Music Retrieval from Audio Embeddings","Crossmodal correspondences between sound and taste are supported in psychology and neuroscience but remain underused in content-based multimedia retrieval. This work formalizes taste-from-audio prediction as a benchmark for music information retrieval using a perceptually validated multi-source corpus. Ten frozen audio encoders from four HEAR families are compared under a shared multi-task regression head, with gated late-fusion as a configurable variant. Evaluation uses absolute error and rank correlation, showing strong systems predict five tastes with macro RMSE 0.134 and retrieval that tracks group consensus better than text baselines.","Taste-aware music retrieval from audio embeddings  \nMatteo Spanio  \nUniversity of Padua, Padua, Italy [spanio@dei.unipd.it](spanio@dei.unipd.it)  \nAntonio Rod  \nUniversity of Padua, Padua, Italy [roda@dei.unipd.it](roda@dei.unipd.it)  \narXiv :2607 .03296v 1 [ cs . SD] 3 Jul 2026  \nAbstract—Crossmodal correspondences between sound and taste are well established in psychology and neuroscience, but largely absent from content-based multimedia retrieval. We formalise taste-from-audio prediction as a content-based music information retrieval benchmark over a perceptually validated multi-source corpus, comparing ten frozen audio encoders from the four HEAR families under a shared multi-task regression head, with gated late-fusion as a configurable variant. In order to assess the effectiveness of the models, we compute absolute error and rank correlation. The strongest systems predict the five tastes within a macro RMSE of 0. 134; on held-out real music their error is less than half a single rater’s deviation from the consensus (RMSE 0.13 vs. 0.28), so the model tracks the group consensus more closely than an average human rater, and well below the previous state of the art baseline (0 .219). On absolute error the encoders are statistically flat, with a single VGGISH matching the best fusion, but gated late-fusion’s advantage is confined to rank correlation (macro Pearson r 0.724 vs. 0.666). Operationalised as a content-based retrieval index, the predicted taste space ranks a 309-item pool far more faithfully than a CLAP-text baseline, which sits at chance; ridge probes and an audio-bandstop knockout read the strongest representations against documented sound–taste correspondences.  \nIndex Terms—music information retrieval, crossmodal learning, content-based retrieval, sonic seasoning  \nI. INTRODUCTION  \nCrossmodal correspondences between sound and taste are one of the clearest examples of stable associations between audition and the chemical senses. High pitch, consonance, and bright timbre are repeatedly associated with sweetness, while lower pitch, roughness, and darker timbre shift listeners toward bitterness and sourness [1]–[3] . These effects already motivate sonic-seasoning applications in restaurants, advertising, and multisensory design [4], yet they remain peripheral to mainstream music information retrieval (MIR) and multimedia benchmarking.  \nThis gap matters for content-based multimedia indexing. If taste judgments can be predicted from audio in a reproducible way, they become a new semantic axis for organising collections, querying music beyond genre and mood, and recommendation scenarios such as “a similar track but sweeter”. The task also creates an unusual test-bed for explainable multimedia learning: a model is useful only if it scores well and if its behaviour can be compared with empirical findings from psychology and neuroscience.  \nTaste tagging from audio is not a completely novel task. Guedes et al. introduced the Taste & Affect Music Database [5]; Rodriguez fine-tuned five separate Audio Spectrogram Transformer (AST) regressors, one per taste, on a curated  \nFig. 1. Proposed model architecture. One or more frozen audio encoders produce per-encoder embeddings that are concatenated and re-weighted by a learned per-encoder gate; the gated representation feeds a shared two-layer MLP whose sigmoid head outputs a 5-D taste vector in [0 , 1]5 .  \n257-song soundtracks corpus and used them to label the FMA corpus at scale [6]; Spanio et al. extended the corpus, validated the labels perceptually, and explored generative variants [7],[8] . Prior work established the feasibility of taste prediction from audio, with evaluation centred on direct regression error. The present paper expands it into a broader content-based MIR setting; we make three contributions.  \n• We formalise taste-from-audio as a content-based MIR benchmark over a perceptually validated multi-source corpus, with a frozen-encoder protocol across ten enc","cbCaiemXxIwMqjvp","https://ap.wps.com/l/cbCaiemXxIwMqjvp","pdf",744373,3,1,"English","en",105,"# Introduction\n# Related Works","[{\"question\":\"What problem does the paper address?\",\"answer\":\"The results show the strongest models predict the five tastes with macro RMSE 0.134, and on held-out real music their error is less than half the deviation of a single human rater from consensus. Retrieval using the predicted taste space ranks items more faithfully than a CLAP-text baseline.\"}]",1784189899,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":77,"head_meta":79,"extra_data":81,"updated_unix":27},"taste-aware-music-retrieval-from-audio-embeddings","",{"@graph":35,"@context":76},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":20},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/taste-aware-music-retrieval-from-audio-embeddings/83708/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does the paper address?","Question",{"text":74,"@type":75},"The results show the strongest models predict the five tastes with macro RMSE 0.134, and on held-out real music their error is less than half the deviation of a single human rater from consensus. Retrieval using the predicted taste space ranks items more faithfully than a CLAP-text baseline.","Answer","https://schema.org",{"og:url":50,"og:type":78,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":80,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":83},[84,88,92,96,101,106,111,114,118,121,125],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":97,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},5,"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":45,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":45,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":28,"slug":117},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":119,"show_sort_weight":28,"slug":120},"World Cup","world-cup",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":122,"slug":124},10,"Lifestyle","lifestyle",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":97,"slug":128},19,"General","general"]