[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122556-en":3,"doc-seo-122556-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122556,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","LLMs can read music, but struggle to hear it - An evaluation of core music perception tasks","Multimodal Large Language Models (MLLMs) claim “musical understanding,” yet most evaluations conflate listening with score reading. This work benchmarks three state-of-the-art LLMs—Gemini 2.5 Pro, Gemini 2.5 Flash, and Qwen2.5-Omni—across Syncopation Scoring, Transposition Detection, and Chord Quality Identification. Variability is analyzed by contrasting audio versus MIDI, comparing zero-shot to few-shot exposure, and testing reasoning strategies including a LogicLM-style schema-and-solver approach. Results show near-ceiling performance on MIDI but substantial accuracy drops on audio, with reasoning and few-shot prompting offering limited gains.","LLMs can read music, but struggle to hear it.  \nAn evaluation of core music perception tasks  \nBrandon James Carone [bcarone@nyu.edu](bcarone@nyu.edu)  \nDepartment of Psychology, Music and Audio Research Laboratory (MARL), Center for Language, Music, and Emotion (CLaME), New York University  \nIran R. Roman [i.roman@qmul.ac.uk](i.roman@qmul.ac.uk)  \nSchool of Electronic Engineering and Computer Science, Queen Mary University of London  \nPablo Ripoll´es [pripolles@nyu.edu](pripolles@nyu.edu)  \nDepartment of Psychology, Music and Audio Research Laboratory (MARL), Center for Language, Music, and Emotion (CLaME), New York University  \nAbstract  \nMultimodal Large Language Models (MLLMs) claim “musical understanding,” yet most evaluations conflate listening with score reading. We benchmark three SOTA LLMs (Gemini 2.5 Pro, Gemini 2.5 Flash, and Qwen2.5-Omni) across three core music skills: Syncopation Scoring (rhythm perception), Transposition Detection (melody perception), and Chord Quality Identification (harmony perception) . Moreover, we separate three sources of variability: (i) perceptual limitations (by contrasting audio recordings vs. symbolic MIDI inputs), (ii) exposure to prior examples (zero-vs. few-shot manipulations), and (iii) reasoning strategies (Standalone, Chain of Thought, LogicLM) . For the latter we adapt LogicLM, a framework combining LLMs with symbolic solvers to perform structured reasoning. In LogicLM, LLMs act as perceptual formulators, generating strict, machinecheckable schemas (onset grids, interval sequences) that deterministic solvers execute with self-refinement. Our results reveal a clear perceptual gap: models perform near ceiling on MIDI but show substantial accuracy drops on audio. Reasoning and few-shot prompting offer minimal gains. This is expected for MIDI, where performance reaches saturation, but more surprising for audio, where LogicLM, despite near-perfect MIDI accuracy, remains notably brittle. Among models, Gemini Pro achieves the highest performance across most conditions. Transposition yields the highest accuracies across models, while Chord Identification scores slightly below Syncopation. Overall, current systems reason well over symbols (MIDI) but do not yet “listen” reliably from audio, with reasoning strategies having little impact over accuracy. Our method and dataset make the perception–reasoning boundary explicit and offer actionable guidance for building robust, audio music systems.  \nKeywords: Audio Large Language Models, Multimodal Large Language Models, Music Understanding, Benchmarking and Evaluation, Schema-Guided Reasoning, LogicLM  \n1. Introduction  \nRecent advances in foundation models have extended their reach beyond text to multimodal architectures that process audio, vision, and language in a unified framework. Models such as Alibaba’s Qwen2.5-Omni, trained on a range of audio tasks (Xu et al., 2025), and Google’s Gemini 2.5 family, which incorporates advanced multimodal integration for real-time interactions (Comanici et al. , 2025), exemplify this new generation. One problem with many multimodal LLMs is that they boast “generic hearing abilities” and “music understanding”, yet struggle with tasks as simple as recognizing a well-known tune when transposed to a different key (i.e., singing the same song at a higher or lower pitch) or played on another  \n© 2026 B.J. Carone, I.R. Roman & P. Ripoll´es.  \nCarone Roman Ripoll´es  \ninstrument. For example, after being fed a simple piano rendition of “Happy Birthday” and told that the melody represents that tune regardless of which key it is played in, they fail to recognize it when played in a different key or on another instrument (in our own pilot tests, Gemini 2.5 guesses that the transposed version of “Happy Birthday” played at the same tempo is “Twinkle Twinkle Little Star” and Qwen2.5-Omni responded with “Lose Yourself” by Eminem) . On the other hand, most people with Western enculturation recognize “Happy Birthday” a","cbCaiq0yGgNxMm1D","https://ap.wps.com/l/cbCaiq0yGgNxMm1D","pdf",616729,1,27,"English","en",105,"# Introduction\n## Multimodal LLMs and “music understanding” claims\n## Limitations of existing audio and music benchmarks\n## Need for listening-based music structure evaluation\n# Abstract\n## Benchmark design and evaluated music skills\n## Sources of variability and reasoning strategies","[{\"question\":\"Which music perception tasks are evaluated in the benchmark?\",\"answer\":\"The benchmark covers Syncopation Scoring (rhythm perception), Transposition Detection (melody perception), and Chord Quality Identification (harmony perception).\"},{\"question\":\"How does the study separate perceptual limitations from other sources of variability?\",\"answer\":\"It contrasts audio recordings with symbolic MIDI inputs, tests zero-shot versus few-shot exposure to prior examples, and evaluates reasoning strategies including a LogicLM-style method.\"},{\"question\":\"What is the key finding about model performance on MIDI versus audio?\",\"answer\":\"Models perform near the ceiling on MIDI but show substantial accuracy drops on audio, and reasoning strategies produce limited improvements for accuracy.\"}]","LLMs can read music, but struggle to hear it - An evaluation of core music perception tasks | PDF",1785811281,68,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"llms-can-read-music-but-struggle-to-hear-it-an-evaluation-of-core-music-perception-tasks","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/llms-can-read-music-but-struggle-to-hear-it-an-evaluation-of-core-music-perception-tasks/122556/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Which music perception tasks are evaluated in the benchmark?","Question",{"text":75,"@type":76},"The benchmark covers Syncopation Scoring (rhythm perception), Transposition Detection (melody perception), and Chord Quality Identification (harmony perception).","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the study separate perceptual limitations from other sources of variability?",{"text":80,"@type":76},"It contrasts audio recordings with symbolic MIDI inputs, tests zero-shot versus few-shot exposure to prior examples, and evaluates reasoning strategies including a LogicLM-style method.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the key finding about model performance on MIDI versus audio?",{"text":84,"@type":76},"Models perform near the ceiling on MIDI but show substantial accuracy drops on audio, and reasoning strategies produce limited improvements for accuracy.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]