[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123116-en":3,"doc-seo-123116-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123116,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Mi-Go - tool which uses YouTube as data source for evaluating general-purpose speech recognition machine learning models","Mi-Go is a tool for evaluating general-purpose automatic speech recognition machine learning models in diverse real-world scenarios. It uses YouTube as a continuously updated data source, covering multiple languages, accents, dialects, speaking styles, and varying audio quality. An experiment evaluated state-of-the-art models using 141 randomly selected YouTube videos. Results highlight YouTube’s value for robust, accurate, and adaptable model assessment under different acoustic conditions. The tool also compares machine transcriptions with human subtitles to help identify possible subtitle misuse, such as SEO manipulation.","Wojnar et al.  \nEURASIP Journal on Audio, Speech, and Music Processing (2024) 2024:24  \n[https://doi.org/10.1186/s13636-024-00343-9](https://doi.org/10.1186/s13636-024-00343-9)  \nEURASIP Journal on Audio, Speech, and Music Processing  \n SOFTWARE Open Access  \nMi-Go: tool which uses YouTube as data  \nsource for evaluating general-purpose speech recognition machine learning models  \nTomasz Wojnar1 , Jarosław Hryszko1* and Adam Roman 1  \nAbstract  \nThis article introduces Mi-Go, a tool aimed at evaluating the performance and adaptability of general-purpose speech recognition machine learning models across diverse real-world scenarios. The tool leverages YouTube as a rich and continuously updated data source, accounting for multiple languages, accents, dialects, speaking styles, and audio quality levels. To demonstrate the effectiveness of the tool, an experiment was conducted, by using Mi-Go to evaluate state-of-the-art automatic speech recognition machine learning models. The evaluation involved a total of 141 randomly selected YouTube videos. The results underscore the utility of YouTube as a valuable data source for evaluation of speech recognition models, ensuring their robustness, accuracy, and adaptability to diverse languages and acoustic conditions. Additionally, by contrasting the machine-generated transcriptions against human-made subtitles, the Mi-Go tool can help pinpoint potential misuse of YouTube subtitles, like search engine optimization.  \n1 Introduction  \nSpeech recognition has become a critical component in numerous applications, ranging from virtual assistantsand transcription services to voice-controlled devices and accessibility tools. The increasing reliance on speech recognition machine learning models necessitates robust and comprehensive evaluation methodologies to ensure their performance, reliability, and adaptability across diverse scenarios.  \nExisting speech recognition models evaluations often rely on curated datasets, such as LibriSpeech [25], CommonVoice [4], and TIMIT [32]. While these datasets provide a controlled environment for evaluation, they may not capture the full spectrum of real-world scenarios, potentially limiting the model’s  \n*Correspondence: Jarosław Hryszko [jaroslaw.hryszko@uj.edu.pl](jaroslaw.hryszko@uj.edu.pl)  \n1 Jagiellonian University, Faculty of Mathematics and Computer Science, Division of Software Engineering, Łojasiewicza 6, Krakow 30-348, Poland  \ngeneralizability. Additionally, these datasets may not be updated frequently, resulting in potential stagnation in performance evaluation.  \nIn this article, we introduce Mi-Go (the name will be explained further), a tool designed to evaluate the prediction performance of general-purpose speech recognition machine learning models. Mi-Go harnesses the power of YouTube as a data source, providing access to a virtually unlimited repository of diverse audio-visual content. YouTube offers a rich and continuously updated collection of spoken language data, encompassing various languages, accents, dialects, speaking styles, and audio quality levels. This makes it an ideal source of data which can be used to evaluate the adaptability and performance of speech recognition models in real-world situations.  \nIn recent years, there has been a growing interest in harnessing the vast amount of data available on platforms such as YouTube for machine learning tasks. Various approaches have been proposed to collect and process data from YouTube, including YouTube- 8M [1], AudioSet [11], and GigaSpeech [6] . However, these methods primarily focus on video and audio  \n© The Author(s) 2024. Open Access This article is licensed under a Creative Commons Attribution 4.0 International License, which permits use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if changes were made. T","cbCaignTx8vRHPa4","https://ap.wps.com/l/cbCaignTx8vRHPa4","pdf",1521676,1,17,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What is Mi-Go used for?\",\"answer\":\"Mi-Go evaluates the prediction performance of general-purpose speech recognition machine learning models across diverse real-world scenarios.\"},{\"question\":\"Why does Mi-Go use YouTube as a data source?\",\"answer\":\"YouTube provides a large, continuously updated repository of spoken language data covering many languages, accents, dialects, speaking styles, and audio quality levels.\"},{\"question\":\"How was Mi-Go validated in the study?\",\"answer\":\"The authors used Mi-Go to evaluate state-of-the-art automatic speech recognition models on 141 randomly selected YouTube videos and compared machine transcriptions with human-made subtitles.\"}]","Mi-Go - tool which uses YouTube as data source for evaluating general-purpose speech recognition machine learning models | PDF",1785814707,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mi-go-tool-which-uses-youtube-as-data-source-for-evaluating-general-purpose-speech-recognition-machine-learning-models","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/mi-go-tool-which-uses-youtube-as-data-source-for-evaluating-general-purpose-speech-recognition-machine-learning-models/123116/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is Mi-Go used for?","Question",{"text":75,"@type":76},"Mi-Go evaluates the prediction performance of general-purpose speech recognition machine learning models across diverse real-world scenarios.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why does Mi-Go use YouTube as a data source?",{"text":80,"@type":76},"YouTube provides a large, continuously updated repository of spoken language data covering many languages, accents, dialects, speaking styles, and audio quality levels.",{"name":82,"@type":73,"acceptedAnswer":83},"How was Mi-Go validated in the study?",{"text":84,"@type":76},"The authors used Mi-Go to evaluate state-of-the-art automatic speech recognition models on 141 randomly selected YouTube videos and compared machine transcriptions with human-made subtitles.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]