[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86324-en":3,"doc-seo-86324-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86324,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Casting Everything to Online API Services A Survey of Integrating Localized Speech Recognition Models in Robotic Systems","Automatic speech recognition (ASR) is a core capability for modern robotic systems, enabling natural spoken human-robot interaction through voice interfaces. The survey reviews integration patterns that range from classic recognition approaches to deep learning models such as OpenAI’s Whisper, alongside major large-scale datasets and open-source toolkits. Deployment strategies are structured by on-device, cloud-based, and hybrid solutions, with examples from real robotic platforms. It concludes with deployment challenges and future directions for robust multilingual and multimodal interaction.","arXiv :2607 . 11792v1 [ cs .RO] 13 Jul 2026  \nCasting Everything to Online API Services? A Survey of Integrating Localized Speech Recognition Models in Robotic Systems  \nSheng Li 1[0000−0001−7636−3797], Jing Li2[0000−0001−8105−5852], Felix Schijve3 , Jun Hu2[0000−0003−2714−6264], and Emilia Barakova2[0000−0001−5688−4878]  \n1 Institute of Science Tokyo, Yokohama, Japan  \n[sheng.li@ieee.org](sheng.li@ieee.org)  \n2 Department of Industrial Design, Eindhoven University of Technology, 5612 AZ  \nEindhoven, The Netherlands  \n{[j.li2](j.li2) ,[j.hu](j.hu) ,[e.i.barakova}@tue.nl](e.i.barakova}@tue.nl)  \n3 Department of Biomedical Engineering, Eindhoven University of Technology, 5612  \nAZ Eindhoven, The Netherlands  \n[felixschijve@hotmail.com](felixschijve@hotmail.com)  \nAbstract. Automatic speech recognition (ASR) has become a critical component of modern robotic systems because it is one of the most natural and intuitive ways for humans to interact with robots. A commonly used method is to directly use API services online. But is that all we can do? This article provides an overview of how ASR technologies are integrated into various intelligent robots and machines. We discuss the evolution of speech recognition from established approaches to state-ofthe-art deep learning models, such as OpenAI’s Whisper. We also list large-scale datasets and open source toolkits that have been widely used in both industry and academia. We structure the survey around ASR model families, deployment strategies in robotics (especially ROS-based, cloud-based, and hybrid solutions), and several real-world robotic platforms. Finally, we outline the challenges of deploying robust speech recognition in robots and discuss future directions, including multimodal interaction in diverse and dynamic environments. This paper can help social robotics researchers better navigate the emerging domain of languagebased natural human-robot interaction.  \nKeywords: Speech recognition · automatic speech recognition (ASR) · human-robot interaction · voice interfaces · robotics.  \n1 Introduction  \nHumans naturally communicate through speech; therefore, giving robots the ability to comprehend spoken language greatly improves human-robot interaction. Automatic Speech Recognition (ASR) has been increasingly integrated into contemporary robotic systems, enabling humans to just use speech to control robots. Applications include industrial robots on factory floors and service  \n2 S. Li et al.  \nrobots in homes and workplaces. ASR makes technology more approachable and user-friendly by enabling robots to receive commands, respond to inquiries, and deliver information.  \nThe applications of early speech recognition systems were limited due to acoustic models that relied heavily on manually constructed labels and vocabularies. In robotics, speech recognition was further limited by the small computational capacity of the onboard robot hardware and their operation in usually noisy natural environments. In addition, failures in robot design, such as the NAO robot that has its microphones near the fans for cooling, introduced additional difficulties in speech recognition. Finally, the early speech models worked mainly with male voices, and interactions with other user groups, such as children and elderly, for example, were challenging [6, 24] . Various dedicated software and hardware solutions have been proposed to alleviate these problems [16, 3] until Deep Learning led to significant advancements in human-robot speech interaction. Under the right circumstances, modern ASR models trained on large audio data sets can attain accuracy comparable to that of humans. For example, with 680,000 hours of training data, OpenAI’s Whisper model [38] exhibits strong multilingual transcription capabilities. These sophisticated models can handle a variety of languages and noisy inputs more effectively than previous systems, and they perform on benchmarks that are close to natural human interaction.[25,","cbCaihmQWEG7vDEc","https://ap.wps.com/l/cbCaihmQWEG7vDEc","pdf",285723,4,1,16,"English","en",105,"# Introduction\n# Method","[{\"question\":\"What motivates integrating ASR into robotic systems?\",\"answer\":\"Speech provides the most natural and intuitive interaction channel for humans. Integrating ASR lets robots receive spoken commands, answer inquiries, and deliver information, improving usability in homes, workplaces, and factories.\"},{\"question\":\"How has speech recognition in robotics evolved according to the survey?\",\"answer\":\"Early approaches were constrained by manual labeling, limited vocabulary, limited onboard compute, and noisy environments. Deep learning improved accuracy and robustness, with models like OpenAI’s Whisper trained on large audio datasets enabling multilingual transcription and better handling of noise.\"},{\"question\":\"What are the main ways to deploy ASR in robotics that the survey organizes around?\",\"answer\":\"The survey structures deployment strategies along three axes: model family, whether processing is onboard, cloud-based, or hybrid. This taxonomy links technical choices to robotic constraints and application requirements.\"}]",1784210478,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"casting-everything-to-online-api-services-a-survey-of-integrating-localized-speech-recognition-models-in-robotic-systems","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/casting-everything-to-online-api-services-a-survey-of-integrating-localized-speech-recognition-models-in-robotic-systems/86324/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What motivates integrating ASR into robotic systems?","Question",{"text":75,"@type":76},"Speech provides the most natural and intuitive interaction channel for humans. Integrating ASR lets robots receive spoken commands, answer inquiries, and deliver information, improving usability in homes, workplaces, and factories.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How has speech recognition in robotics evolved according to the survey?",{"text":80,"@type":76},"Early approaches were constrained by manual labeling, limited vocabulary, limited onboard compute, and noisy environments. Deep learning improved accuracy and robustness, with models like OpenAI’s Whisper trained on large audio datasets enabling multilingual transcription and better handling of noise.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the main ways to deploy ASR in robotics that the survey organizes around?",{"text":84,"@type":76},"The survey structures deployment strategies along three axes: model family, whether processing is onboard, cloud-based, or hybrid. This taxonomy links technical choices to robotic constraints and application requirements.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]