[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126890-en":3,"doc-seo-126890-105":30,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},126890,2336474466712,"Maeve","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Hardware-Software Co-Design of an Audio Feature Extraction Pipeline for Machine Learning Applications - Article","Keyword spotting is a critical component of modern speech recognition pipelines, but conventional systems often rely on MFCC audio features that are computationally demanding. This work investigates simplifying MFCC features to better support always-on, resource-constrained embedded keyword spotting. A hardware generator is implemented to produce a matching hardware pipeline for the simplified extraction. Using the Chisel4ml framework, hardware generators are integrated into the Python Keras workflow, enabling model training with the proposed features.","electronics   \nArticle  \nHardware–Software Co-Design of an Audio Feature Extraction Pipeline for Machine Learning Applications  \nJure Vreˇca 1,2, *, Ratko Pilipovi´c 3 and Anton Biasizzo 1  \nCitation: Vreˇca, J.; Pilipovi´c, R.; Biasizzo, A. Hardware–Software Co-Design of an Audio Feature Extraction Pipeline for Machine Learning Applications. Electronics 2024, 13, 875. [https://doi.org/](https://doi.org/)[ ](https://doi.org/)[10.3390/electronics13050875](10.3390/electronics13050875)  \nAcademic Editor: Chunping Li  \nReceived: 31 January 2024  \nRevised: 17 February 2024  \nAccepted: 22 February 2024  \nPublished: 24 February 2024  \nCopyright: © 2024 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license ([https://](https://)[ ](https://)[creativecommons.org/licenses/by/](creativecommons.org/licenses/by/)[ ](creativecommons.org/licenses/by/)[4.0/](4.0/)) .  \n1 Jožef Stefan Institute, 1000 Ljubljana, Slovenia; anton.biasizzo@ijs.si  \n2 Jožef Stefan International Postgraduate School (IPS), 1000 Ljubljana, Slovenia  \n3 Faculty of Computer and Information Science, University of Ljubljana, 1000 Ljubljana, Slovenia; [ratko.pilipovic@fri.uni-lj.si](ratko.pilipovic@fri.uni-lj.si)  \n* [Correspondence: jure.vreca@ijs.si](Correspondence: jure.vreca@ijs.si)  \nAbstract: Keyword spotting is an important part of modern speech recognition pipelines. Typical contemporary keyword-spotting systems are based on Mel-Frequency Cepstral Coefficient (MFCC) audio features, which are relatively complex to compute. Considering the always-on nature of many keyword-spotting systems, it is prudent to optimize this part of the detection pipeline. We explore the simplifications of the MFCC audio features and derive a simplified version that can be more easily used in embedded applications. Additionally, we implement a hardware generator that generates an appropriate hardware pipeline for the simplified audio feature extraction. Using Chisel4ml framework, we integrate hardware generators into Python-based Keras framework, which facilitates the training process of the machine learning models using our simplified audio features.  \nKeywords: FPGA; MFCC; keyword spotting; chisel  \n1. Introduction  \nThe development of deep neural networks has opened up possibilities for applications in diverse areas. One such area is speech recognition, where researchers have been able to show promising results using large transformer-based neural networks [1] . These networks are, however, very computationally expensive and thus are hard to implement on battery-powered devices. This can be overcome using a simpler keyword detection system that merely listens for specific keywords (e.g., Hey Siri) and then wakes up a more powerful system generally implemented in the cloud. In this way, the simpler low-power system acts as an always-on listener, and the more powerful system is used only when needed. Henceforth, we will refer to the first kind of system as the keyword-spotting (KWS) system [2], and the second type as the Large Vocabulary Speech Recognition system.  \nThe keyword-spotting system is an always-on system. This requirement facilitates the need for it to be as energy-efficient as possible. The authors of [3] recognized this problem and explored an integer-only implementation of the MFCC algorithm. They achieved good results. However, they refrained from modifying the MFCC algorithm. Furthermore, they targeted a DSP processing unit. Dedicated hardware circuits could make this process even more energy-efficient. The authors of [4] developed a custom MFCC extraction processing unit. However, they also failed to explore simplifications of the MFCC algorithm and instead focused exclusively on the hardware implementation. The authors of [5] showcased a system that simplifies the MFCC features to a certain degree. While they did obtain satisfactory classifi","cbCaiu7wczKP3mog","https://ap.wps.com/l/cbCaiu7wczKP3mog","pdf",1104699,1,14,"English","en",105,"# Introduction\n## Simplified MFCC for energy-efficient keyword spotting\n# Proposed contributions and paper structure\n## MFCC standard calculation and studied simplifications\n## Hardware modules for simplified MFCC features\n## Chisel4ml framework and Python-compiler integration\n## Accuracy impact of MFCC simplifications in KWS\n## Hardware synthesis results","[{\"question\":\"Why are MFCC features targeted for optimization in keyword spotting systems?\",\"answer\":\"Keyword-spotting systems are always-on and must be energy-efficient. MFCC-based systems compute relatively complex features, making optimization beneficial for embedded deployment.\"},{\"question\":\"What does the paper contribute regarding MFCC audio features?\",\"answer\":\"The paper formulates a simplified version of MFCC features and analyzes how these simplifications affect keyword-spotting (KWS) performance accuracy.\"},{\"question\":\"How are hardware generators used in this work?\",\"answer\":\"A hardware generator is implemented to generate a hardware pipeline that computes the simplified audio feature extraction efficiently, focusing on resource efficiency.\"},{\"question\":\"How does Chisel4ml connect the hardware pipeline with machine learning training?\",\"answer\":\"Chisel4ml integrates hardware generators into a Python-based Keras workflow, allowing training of machine learning models using the simplified audio features.\"}]","Hardware-Software Co-Design of an Audio Feature Extraction Pipeline for Machine Learning Applications - Article | PDF",1785935445,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":28},"hardware-software-co-design-of-an-audio-feature-extraction-pipeline-for-machine-learning-applications-article","",{"@graph":36,"@context":89},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/hardware-software-co-design-of-an-audio-feature-extraction-pipeline-for-machine-learning-applications-article/126890/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"Why are MFCC features targeted for optimization in keyword spotting systems?","Question",{"text":75,"@type":76},"Keyword-spotting systems are always-on and must be energy-efficient. MFCC-based systems compute relatively complex features, making optimization beneficial for embedded deployment.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the paper contribute regarding MFCC audio features?",{"text":80,"@type":76},"The paper formulates a simplified version of MFCC features and analyzes how these simplifications affect keyword-spotting (KWS) performance accuracy.",{"name":82,"@type":73,"acceptedAnswer":83},"How are hardware generators used in this work?",{"text":84,"@type":76},"A hardware generator is implemented to generate a hardware pipeline that computes the simplified audio feature extraction efficiently, focusing on resource efficiency.",{"name":86,"@type":73,"acceptedAnswer":87},"How does Chisel4ml connect the hardware pipeline with machine learning training?",{"text":88,"@type":76},"Chisel4ml integrates hardware generators into a Python-based Keras workflow, allowing training of machine learning models using the simplified audio features.","https://schema.org",{"og:url":52,"og:type":91,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":93,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":96},[97,101,105,109,114,119,124,127,132,135,139],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},5,"Comic",60,"comic",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},6,"Technology",50,"technology",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":125,"slug":126},30,"research-report",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":130,"slug":131},9,"Religion & Spirituality",20,"religion-spirituality",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":130,"slug":134},"World Cup","world-cup",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":136,"slug":138},10,"Lifestyle","lifestyle",{"id":140,"doc_module":4,"doc_module_name":46,"category_name":141,"show_sort_weight":110,"slug":142},19,"General","general"]