[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86287-en":3,"doc-seo-86287-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86287,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Breaking the 15% Barrier: A Real-World Data-Driven System for Proactive Social Robot Triggered by User Nonverbal Cues","Retail service robots often use cascaded speech pipelines (STT–LLM–TTS), but many customer interactions start through nonverbal actions such as approaching, waving, pointing, or showing items. This work examines a real-world store deployment with a teleoperated humanoid robot and finds that 15.3% of robot utterances are triggered by users’ nonverbal behaviors rather than spoken input. Frequent service-relevant cues are defined and recognized online via a real-time, multi-person, multi-label video model. An LLM-based dialogue framework conditions responses on cue tokens and can use vision-language evidence when items are shown, enabling proactive turn-taking without hand-crafted rules.","Breaking the 15% Barrier: A Real-World Data-Driven System for Proactive Social Robot Triggered by User Nonverbal Cues  \nYuga Yano 1 ,2 , Yuki Okafuji2 ,3 , Ryo Miyoshi2 ,3 , Sanae Yamashita2 ,3 , Yoshiki Ohira2 ,3  \narXiv :2607 . 11633v1 [ cs .RO] 13 Jul 2026  \nAbstract—Service robots in retail stores increasingly rely on cascaded speech pipelines (STT–LLM–TTS), yet many customer–robot interactions are initiated or guided by nonverbal behaviors such as approaching, waving, pointing, or showing items. This paper studies such cues in a real-world store deployment with a teleoperated humanoid robot and shows that anon-negligible portion of robot turns are triggered by nonverbal behaviors rather than spoken input, revealing a limitation of audio-only dialogue systems. In a 6-day in-the-wild deployment, 15.3% of robot utterances were initiated by users’ nonverbal behaviors rather than spoken input. Based on an analysis of observed customer behaviors, we define a set of frequent, service-relevant nonverbal cues and develop a real-time multiperson, multi-label recognizer that runs online from video. We then propose a dialogue framework that conditions LLM-based utterance generation on recognized nonverbal cue tokens, and optionally leverages a vision–language model when items are shown, enabling proactive robot responses without hand-crafted rules. We evaluate the approach offline on nonverbal-triggered turns and demonstrate an online prototype that reacts to users’nonverbal cues in real time.  \nI. INTRODUCTION  \nService robots capable of autonomous customer-service dialogue are being introduced in retail stores to reduce the workload of store staff and improve the customer experience [1]–[3] . Recently, advances in large language models (LLMs) such as ChatGPT [4] and Gemini [5], as well as speech-to-text (STT) and text-to-speech (TTS) technologies, have significantly improved the verbal interaction capabilities of service robots. Meanwhile, many nonverbal behaviors—such as waving and pointing—often trigger interactions [6] . In contrast, widely deployed spoken dialogue systems are often implemented as a cascaded pipeline: the input speech is transcribed by STT, the transcript is fed into an LLM to generate a response, and the response is synthesized by TTS [7] . In such systems based on only audio input, spoken dialogue systems cannot incorporate users’ nonverbal behaviors as inputs, which leads to missed opportunities for engagement.  \nWhile many studies of human activity recognition have focused on developing dedicated models, recent advances in large-scale infrastructure models have made it possible to describe detailed user behavior using vision language models (VLMs) [8]–[10] . However, such VLM-based approaches require a high-performance GPU for real-time recognition, and significant challenges remain before deployment as a  \n1 Kyushu Institute of Technology, Fukuoka, Japan  \n2 CyberAgent, Tokyo, Japan  \n3 The University of Osaka, Osaka, Japan  \nFor correspondence: [yano.yuuga158@mail.kyutech.jp](yano.yuuga158@mail.kyutech.jp)  \nComing / waving  \nShowing items  \nFig. 1. In this study, we (a) collect dialogue data from a remote-controlled robot installed in a physical store,(b) analyze utterances resulting from the user’s nonverbal behavior, and (c) propose an autonomous dialogue system based on that analysis. We build a dialogue robot that can respond to the user’s nonverbal behavior in real time.  \ncustomer-service dialogue system. In this study, real-time recognition refers to a speed of at least 1 fps. This is because humans generally expect a response delay of approximately one second to user actions [11]; therefore, human activity recognition must be completed within this time frame, even when the time required for speech generation and voice synthesis is excluded.  \nTherefore, it is essential to analyze nonverbal behaviors that play important roles in customer-service interactions and to develop a dialogue framework that","cbCaiqPTx0Tnv4lP","https://ap.wps.com/l/cbCaiqPTx0Tnv4lP","pdf",3392379,3,1,"English","en",105,"# Introduction\n## Problem: limitations of audio-only dialogue pipelines\n## Real-time recognition requirement (about 1 fps)\n# Contributions\n## Data collection and quantification of nonverbal-triggered turns\n## Real-time multi-label recognition and proactive dialogue framework\n# System Overview (Fig. 1)\n# Identified nonverbal triggers","[{\"question\":\"What percentage of the robot’s utterances are triggered by users’ nonverbal behaviors in the real-world deployment?\",\"answer\":\"In the 6-day in-the-wild deployment, 15.3% of robot utterances were initiated by users’ nonverbal behaviors rather than spoken input.\"},{\"question\":\"Why do audio-only cascaded speech pipelines limit service-robot interactions?\",\"answer\":\"Because they rely on STT output and cannot incorporate nonverbal behaviors (e.g., waving, pointing, showing items) as inputs, causing missed engagement opportunities.\"},{\"question\":\"How does the proposed system enable proactive robot responses from nonverbal cues?\",\"answer\":\"It defines frequent service-relevant nonverbal cues, recognizes them online with a real-time multi-person, multi-label video recognizer, and conditions LLM-based utterance generation on the recognized cue tokens; it can also leverage a vision-language model when items are shown.\"}]",1784210076,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"breaking-the-15-barrier-a-real-world-data-driven-system-for-proactive-social-robot-triggered-by-user-nonverbal-cues","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":20},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/breaking-the-15-barrier-a-real-world-data-driven-system-for-proactive-social-robot-triggered-by-user-nonverbal-cues/86287/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What percentage of the robot’s utterances are triggered by users’ nonverbal behaviors in the real-world deployment?","Question",{"text":74,"@type":75},"In the 6-day in-the-wild deployment, 15.3% of robot utterances were initiated by users’ nonverbal behaviors rather than spoken input.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"Why do audio-only cascaded speech pipelines limit service-robot interactions?",{"text":79,"@type":75},"Because they rely on STT output and cannot incorporate nonverbal behaviors (e.g., waving, pointing, showing items) as inputs, causing missed engagement opportunities.",{"name":81,"@type":72,"acceptedAnswer":82},"How does the proposed system enable proactive robot responses from nonverbal cues?",{"text":83,"@type":75},"It defines frequent service-relevant nonverbal cues, recognizes them online with a real-time multi-person, multi-label video recognizer, and conditions LLM-based utterance generation on the recognized cue tokens; it can also leverage a vision-language model when items are shown.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]