[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119730-en":3,"doc-seo-119730-105":30,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119730,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1785132997149421697",8,"Research & Report","Steering Latent Audio Models through Interactive Machine Learning","Proof-of-concept mechanism for steering latent audio models through interactive machine learning, mapping a human-performance space to a neural audio model’s high-dimensional latent space. A regressive model is trained from demonstrative actions to create paired examples between performer gestures and latent representations. Implemented during ideation, exploration, and sound/music performance, the method demonstrates efficient, flexible, and immediate control over generative audio processes.","Steering latent audio models through interactive machine learning  \nGabriel Vigliensoni 1 and Rebecca Fiebrink2  \n1, 2 Creative Computing Institute, University of the Arts London, UK  \n1 Centre for Interdisciplinary Research in Music Media and Technology, QC  \n[g.vigliensoni@arts.ac.uk](g.vigliensoni@arts.ac.uk)  \nAbstract  \nIn this paper, we present a proof-of-concept mechanism for steering latent audio models through interactive machine learning. Our approach involves mapping the human-performance space to the highdimensional, computer-generated latent space of a neural audio model by utilizing a regressive model learned from a set of demonstrative actions. By implementing this method in ideation, exploration, and sound and music performance we have observed its efficiency, flexibility, and immediacy of control over generative audio processes.  \nIntroduction  \nRecent advances in neural audio synthesis have made it possible to generate audio signals in real time, enabling the use of applications in musical performance. However, exploring and playing with their high-dimensional spaces remains challenging, as the axes do not necessarily correlate to clear musical labels and may vary from model to model. In this paper, we investigate and propose a useful new approach based on interactive machine learning. This approach allows the performer to map the well-known, low-dimensional, human performance space to the high-dimensional generative audio model’s latent space by providing training examples that pair the two spaces.  \nBackground  \nGenerative AI audio models  \nGenerative AI audio models provide a data-driven approach to sound generation. These systems are designed to autonomously generate audio signals by learning from existing or custom datasets, capturing the underlying patterns and characteristics of the input data. However, historical systems for generative audio modelling and synthesis, such as WaveNet (Oord et al. 2016) and SampleRNN (Mehri et al. 2017)), have been challenging to integrate into creative environments due to their large computational complexity, poor signal quality, short temporal coherency, and lack of interaction means. Newer neural audio synthesis architectures and systems such as DDSP (Engel et al. 2020) and Jukebox (Dhariwal et al. 2020) have introduced advancements that addressed part of the previously mentioned issues. DDSP can model audio signals using small training  \ndatasets and can be steered in real time using pitch and amplitude as generative conditions, but only for monophonic instrument signals. Jukebox can generate a singing voiceoverlaid on top of complex, polyphonic music signal using text, genre, and artist labels as condition factors, but it requires massive computational power and datasets to be trained and lacks real-time control at generation time. The more recent architecture RAVE (Caillon and Esling 2021) addresses all the aforementioned issues in the context of modelling complex, polyphonic audio signals. However, given the potentially large dimensionality of the learned embedding and also the lack of labels for the latent space axes, there is a need to find a better way for real-time interaction and performing with such models.  \nSteering Generative AI  \nReal-time control in neural audio synthesis systems is important as it can enable performers to introduce the longterm temporal coherence often missing in these systems. That is, a generative model producing audio signals with short-term temporal coherence can still be used to generate longer structures if meaningful control is applied during generation. We next describe three main approaches to exerting control on the generative process.  \nTraining data. In creative contexts, the choice of training dataset serves as the primary mechanism through which a human creator specifies what kind of content the machine should generate. This approach is often overlooked due to the extensive data and processing power required by most gener","cbCaigJ9ascrzQyp","https://ap.wps.com/l/cbCaigJ9ascrzQyp","pdf",255886,1,4,"English","en",105,"# Abstract\n# Introduction\n# Background\n## Generative AI audio models\n## Steering Generative AI\n### Training data\n### Conditioning\n### Latent manipulation","[{\"question\":\"How does the proposed method steer latent audio models?\",\"answer\":\"It maps a low-dimensional human-performance space to a model’s high-dimensional latent space using a regressive model trained from demonstrative actions paired across both spaces.\"},{\"question\":\"Why is interactive control important in neural audio synthesis?\",\"answer\":\"It supports longer temporal coherence during generation, enabling generative systems that may have short-term coherence to form longer musical structures when meaningful control is applied.\"},{\"question\":\"What are the main approaches to exerting control in generative audio systems?\",\"answer\":\"The paper outlines training data selection, conditioning during setup or inference, and latent manipulation by overriding latent dimensions with performer input.\"}]","Steering Latent Audio Models through Interactive Machine Learning | PDF",1785725994,10,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":28},"steering-latent-audio-models-through-interactive-machine-learning","",{"@graph":36,"@context":84},[37,53,67],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":21},"https://docshare.wps.com/document/steering-latent-audio-models-through-interactive-machine-learning/119730/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"How does the proposed method steer latent audio models?","Question",{"text":74,"@type":75},"It maps a low-dimensional human-performance space to a model’s high-dimensional latent space using a regressive model trained from demonstrative actions paired across both spaces.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"Why is interactive control important in neural audio synthesis?",{"text":79,"@type":75},"It supports longer temporal coherence during generation, enabling generative systems that may have short-term coherence to form longer musical structures when meaningful control is applied.",{"name":81,"@type":72,"acceptedAnswer":82},"What are the main approaches to exerting control in generative audio systems?",{"text":83,"@type":75},"The paper outlines training data selection, conditioning during setup or inference, and latent manipulation by overriding latent dimensions with performer input.","https://schema.org",{"og:url":52,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,133],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":29,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":29,"slug":132},"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]