[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83241-en":3,"doc-seo-83241-105":29,"detail-sidebar-cat-0-en-105":89},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":11},83241,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Immersive Social Interaction with VR and LLM-Assisted Humanoids","The work presents a versatile humanoid teleoperation framework enabling immersive, whole-body control for locomotion, manipulation, and social interaction. High-level locomotion is driven by real-time voice commands translated into robot navigation actions, while manipulation uses VR hand tracking to retarget wrist and finger poses to dexterous humanoid arms via inverse kinematics and PD control. The system records multi-modal data from ego-centric vision, voice/text, body and hand joint angles, and eye movements to support imitation learning and more natural collaboration scenarios.","Immersive Social Interaction with VR and LLM-Assisted Humanoids  \nNiraj Pudasaini 1∗, Geeta Chandra Raju Bethala 1∗, Pranav Doma 1∗, Anthony Tzes 1 , Yi Fang 1  \narXiv :2607 .07430v 1 [ cs .RO] 8 Jul 2026  \nI. INTRODUCTION  \nThe growing demand for human-robot interaction in diverse fields such as elder care, social engagement, and search-and-rescue has driven significant advancements inteleoperation systems for humanoid robots. The aging global population has led to a rise in social isolation and mobility challenges, especially for homebound older adults [1] . This isolation significantly affects the mental and physical well-being of aging people, creating a need for innovative solutions that facilitate remote engagement and interaction. Conversational telepresence and teleoperation robots offera promising tool for these individuals to connect with the external world, participate in social activities, and maintain a sense of independence. Beyond enhancing social interaction for older adults, teleoperation robotic systems can potentially be used for search and rescue operations in hazardous environments [2] where human presence is dangerous. However, existing teleoperated robotic systems often face significant challenges in locomotion control, particularly when navigating complex and unstructured environments [3], [4] . Additionally, traditional teleoperation interfaces for controlling robot locomotion are less user-friendly and cognitively demanding [5], requiring operators to manage multiple joints simultaneously.  \nThis paper introduces a versatile humanoid teleoperation system that leverages voice commands for locomotion and VR hand tracking for manipulation, enabling users to have whole-body control of a humanoid robot for locomotion, manipulation and social interaction tasks. Additionally, Our framework can be used for multi-modal data collection during teleoperation, which can then be utilized to train robots via imitation learning for more complex tasks. Beyond the tasks studied here, the proposed interface can support natural human-humanoid interactions and collaboration scenarios such as cooperative manipulation [6] .  \nII. METHODOLOGY  \nOur method consists of: 1) . voice-controlled locomotion, 2) . teleoperation manipulation, and 3) . social interaction fora humanoid robot, facilitating real-time whole-body control and allowing multi-modal data collection for task learning (see Fig 2) .  \nA. Voice Commands for High-level Locomotion  \nBased on the ego-centric images streamed with a resolution of 640 × 480 on an Apple Vision Pro, the user can  \n1 All authors are affiliated with New York University Abu Dhabi (NYUAD), UAE. {np2289, [gb2643](gb2643}@nyu.edu)[}](gb2643}@nyu.edu)[@nyu.edu](gb2643}@nyu.edu)  \nFig. 1. Demonstration of tasks involving voice-controlled locomotion, teleoperated manipulation, and social interaction. For the manipulation task, the user teleoperates with the hands to pick the bottle and place it in a box. For the social interaction task, the robot takes the blue cube from a person, walks towards another person, and hands the cube. Finally, the robot performs hand-shaking gesture.  \nsend the locomotion command through voice. The voice loco-control module translates the locomotion commands into high-level control commands, such as move(x, y), rotate(angle), stop(), and stand() . Then, the relevant high-level functions are called for robot navigation. The robot’s bi-pedal locomotion is based on a pre-trained deep reinforcement learning model to produce robust locomotion policy, as developed in prior works [7], [8] . Note that during implementation, we faced challenges with the robot’s bi-pedal morphology, which does not inherently stabilize itself.  \nFor the voice loco-control module, we use Deepgram [9] for real-time speech-to-text transcription, GPT-4’s [10] reasoning to parse locomotion commands from text into highlevel control commands, and Silero [11] for text-to-speech synthesis with the LivKit ","cbCaid96BfxxkRJm","https://ap.wps.com/l/cbCaid96BfxxkRJm","pdf",2087945,4,1,3,"English","en",105,"# Introduction\n# Methodology\n## Voice Commands for High-level Locomotion\n## Teleoperation for Manipulation\n## Social Interaction","[{\"question\":\"How does the system control humanoid locomotion using voice commands?\",\"answer\":\"It transcribes speech to text in real time, uses GPT-4 reasoning to parse the text into high-level navigation controls such as move/rotate/stop/stand, and then executes those functions. If the instruction is uncertain, the system requests user confirmation before execution.\"},{\"question\":\"How is VR hand tracking used for teleoperated manipulation?\",\"answer\":\"Apple Vision Pro streams wrist and finger poses, which are transformed into the robot coordinate frame. Inverse kinematics compute joint angles, and a PD controller drives dexterous hands so the robot mimics the operator’s arm and finger motions.\"},{\"question\":\"What role does multi-modal data collection play in the framework?\",\"answer\":\"During teleoperation, the system records ego-centric images, voice/text commands, body and hand joint angles, and eye movements. This multi-modal dataset can be used to train robots via imitation learning for more complex tasks.\"}]",1784186186,{"code":4,"msg":30,"data":31},"ok",{"site_id":25,"language":24,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":84,"head_meta":86,"extra_data":88,"updated_unix":28},"immersive-social-interaction-with-vr-and-llm-assisted-humanoids","",{"@graph":35,"@context":83},[36,51,66],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":22},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":42,"position":20},"https://docshare.wps.com/document/immersive-social-interaction-with-vr-and-llm-assisted-humanoids/83241/",{"url":50,"name":13,"@type":52,"author":53,"headline":13,"publisher":55,"fileFormat":58,"inLanguage":24,"description":14,"dateModified":59,"datePublished":60,"encodingFormat":58,"isAccessibleForFree":61,"interactionStatistic":62},"DigitalDocument",{"name":9,"@type":54},"Person",{"url":40,"name":56,"@type":57},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":63,"interactionType":64,"userInteractionCount":20},"InteractionCounter",{"@type":65},"ViewAction",{"@type":67,"mainEntity":68},"FAQPage",[69,75,79],{"name":70,"@type":71,"acceptedAnswer":72},"How does the system control humanoid locomotion using voice commands?","Question",{"text":73,"@type":74},"It transcribes speech to text in real time, uses GPT-4 reasoning to parse the text into high-level navigation controls such as move/rotate/stop/stand, and then executes those functions. If the instruction is uncertain, the system requests user confirmation before execution.","Answer",{"name":76,"@type":71,"acceptedAnswer":77},"How is VR hand tracking used for teleoperated manipulation?",{"text":78,"@type":74},"Apple Vision Pro streams wrist and finger poses, which are transformed into the robot coordinate frame. Inverse kinematics compute joint angles, and a PD controller drives dexterous hands so the robot mimics the operator’s arm and finger motions.",{"name":80,"@type":71,"acceptedAnswer":81},"What role does multi-modal data collection play in the framework?",{"text":82,"@type":74},"During teleoperation, the system records ego-centric images, voice/text commands, body and hand joint angles, and eye movements. This multi-modal dataset can be used to train robots via imitation learning for more complex tasks.","https://schema.org",{"og:url":50,"og:type":85,"og:title":13,"og:site_name":56,"og:description":14},"article",{"robots":87,"canonical":50},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":90},[91,95,99,103,108,113,118,121,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":92,"show_sort_weight":93,"slug":94},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":96,"show_sort_weight":97,"slug":98},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":100,"show_sort_weight":101,"slug":102},"Exam",70,"exam",{"id":104,"doc_module":4,"doc_module_name":45,"category_name":105,"show_sort_weight":106,"slug":107},5,"Comic",60,"comic",{"id":109,"doc_module":4,"doc_module_name":45,"category_name":110,"show_sort_weight":111,"slug":112},6,"Technology",50,"technology",{"id":114,"doc_module":4,"doc_module_name":45,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":119,"slug":120},30,"research-report",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":104,"slug":136},19,"General","general"]