[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120486-en":3,"doc-seo-120486-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120486,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Deep learning’s shallow gains - a comparative evaluation of algorithms for automatic music generation","Deep learning methods are widely regarded as state of the art, yet their advantage for automatic music generation (AMG) using symbolic tokens in a target style has not been conclusively demonstrated. A listening study compares multiple music generation systems across six musical dimensions: stylistic success, aesthetic pleasure, repetition/self-reference, melody, harmony, and rhythm. Models include deep learning algorithms and non-deep methods. In controlled 30-s excerpts, Bayesian non-parametric hypothesis testing evaluates differences and finds deep learning does not clearly outperform classical baselines, while human-composed examples remain significantly ahead.","[eprints@whiterose.ac.uk](eprints@whiterose.ac.uk)[ ](eprints@whiterose.ac.uk)[https://eprints.whiterose.ac.uk](https://eprints.whiterose.ac.uk)  \nUniversities of Leeds, Sheffield and York  \nDeposited via The University of York.  \nWhite Rose Research Online URL for this paper:  \n[https://eprints.whiterose.ac.uk/id/eprint/233324/](https://eprints.whiterose.ac.uk/id/eprint/233324/)  \nVersion: Published Version  \nArticle:  \nYIN, ZONGYU, REUBEN PARIS, FEDERICO, Stepney, Susan et al. (2023) Deep learning’s shallow gains: a comparative evaluation of algorithms for automatic music generation. Machine Learning. pp. 1785-1822. ISSN: 0885-6125  \n[https://doi.org/10.1007/s10994-023-06309-w](https://doi.org/10.1007/s10994-023-06309-w)  \nReuse  \nThis article is distributed under the terms of the Creative Commons Attribution (CC BY) licence. This licence allows you to distribute, remix, tweak, and build upon the work, even commercially, as long as you credit the authors for the original work. More information and the full terms of the licence here: [https://creativecommons.org/licenses/](https://creativecommons.org/licenses/)  \nTakedown  \nIf you consider content in White Rose Research Online to be in breach of UK law, please notify us by  \nemailing [eprints@whiterose.ac.uk](eprints@whiterose.ac.uk) including the URL of the record and the reason for the withdrawal request.  \nDeep learning’s shallow gains: a comparative evaluation of algorithms for automatic music generation  \nZongyu Yin1 · Federico Reuben2 · Susan Stepney1 · Tom Collins2,3  \nReceived: 25 June 2021 / Revised: 20 October 2022 / Accepted: 27 January 2023 /  \nPublished online: 21 March 2023 © The Author(s) 2023  \nAbstract  \nDeep learning methods are recognised as state-of-the-art for many applications of machine learning. Recently, deep learning methods have emerged as a solution to the task of automatic music generation (AMG) using symbolic tokens in a target style, but their superiority over non-deep learning methods has not been demonstrated. Here, we conduct a listening study to comparatively evaluate several music generation systems along six musical dimensions: stylistic success, aesthetic pleasure, repetition or self-reference, melody, harmony, and rhythm. A range of models, both deep learning algorithms and other methods, are used to generate 30-s excerpts in the style of Classical string quartets and classical piano improvisations. Fifty participants with relatively high musical knowledge rate unlabelled samples of computer-generated and human-composed excerpts for the six musical dimensions. We use non-parametric Bayesian hypothesis testing to interpret the results, allowing the possibility of finding meaningful non-differences between systems’ performance. We find that the strongest deep learning method, a reimplemented version of Music Transformer, has equivalent performance to a non-deep learning method, MAIA Markov, demonstrating that to date, deep learning does not outperform other methods for AMG. We also find there still remains a significant gap between any algorithmic method and humancomposed excerpts.  \nKeywords Deep learning · Non-parametric Bayesian hypothesis testing · Markov model · Music generation · Comparative evaluation · Listening study  \nEditor: Tijl De Bie.  \n* Zongyu Yin [zongyu.yin@outlook.com](zongyu.yin@outlook.com)  \nExtended author information available on the last page of the article  \n1 Introduction  \nIn the past decade, breakthroughs in artificial intelligence (AI) and deep learning have been established as such through rigorous, comparative evaluations,1 for example, in computer vision (O’Mahony et al., 2019) and automatic speech recognition (Toshniwal et al., 2018) . In the field of automatic music generation (AMG), however, to our knowledge there has been no comparative evaluation to date between deep learning and other methods (Huang et al., 2018 ; Yang et al., 2017 ; Dong et al., 2018 ; Hadjeres et al., 2017 ; Thickstun et al., 2019 ;","cbCaibL8Gg5UVlOs","https://ap.wps.com/l/cbCaibL8Gg5UVlOs","pdf",1726560,1,39,"English","en",105,"# Introduction\n## Background and motivation\n## Research questions\n# Related work\n## Musical representations and method categories\n# Evaluation approach","[{\"question\":\"What does the paper compare in automatic music generation?\",\"answer\":\"It conducts a comparative listening study across multiple AMG systems, including both deep learning and non-deep learning approaches.\"},{\"question\":\"Which musical dimensions are used to evaluate the generated excerpts?\",\"answer\":\"The study evaluates stylistic success, aesthetic pleasure, repetition or self-reference, melody, harmony, and rhythm.\"},{\"question\":\"What is the main finding about deep learning versus non-deep methods?\",\"answer\":\"The strongest deep learning method (a reimplemented Music Transformer) shows performance equivalent to the non-deep MAIA Markov, indicating deep learning does not yet outperform other methods for AMG.\"}]","Deep learning’s shallow gains - a comparative evaluation of algorithms for automatic music generation | PDF",1785730325,98,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"deep-learnings-shallow-gains-a-comparative-evaluation-of-algorithms-for-automatic-music-generation","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/deep-learnings-shallow-gains-a-comparative-evaluation-of-algorithms-for-automatic-music-generation/120486/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does the paper compare in automatic music generation?","Question",{"text":75,"@type":76},"It conducts a comparative listening study across multiple AMG systems, including both deep learning and non-deep learning approaches.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which musical dimensions are used to evaluate the generated excerpts?",{"text":80,"@type":76},"The study evaluates stylistic success, aesthetic pleasure, repetition or self-reference, melody, harmony, and rhythm.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the main finding about deep learning versus non-deep methods?",{"text":84,"@type":76},"The strongest deep learning method (a reimplemented Music Transformer) shows performance equivalent to the non-deep MAIA Markov, indicating deep learning does not yet outperform other methods for AMG.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]