[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84168-en":3,"doc-seo-84168-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84168,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Ego-Human Motion Prediction with 3D-Aware LLM","Ego-Human Motion Prediction with 3D-Aware LLM focuses on egocentric forecasting of human motion for proactive AR/VR assistance, human–robot collaboration, and embodied AI. It addresses ambiguity in wearable, sparse observations by introducing Ego3DLM, which grounds prediction in 3D spatial and semantic context. The model performs holistic single-pass autoregressive decoding of past/future poses and past/future narrations, enforcing cross-modal and temporal consistency. Training uses scene awareness pretraining, unified instruction tuning, and GRPO reinforcement finetuning to optimize pose–language fidelity.","arXiv :2607 .0700 1v 1 [ cs .CV] 8 Jul 2026  \nEgo-Human Motion Prediction with 3D-Aware LLM  \nYujin Bae⋆ , Jaewoo Jeong⋆ , Hyeonseong Kim⋆ , and Kuk-Jin Yoon  \nVisual Intelligence Lab., KAIST, Korea {yujinbae,jeong207,brian617,[kjyoon}@kaist.ac.kr](kjyoon}@kaist.ac.kr)  \nAbstract. Anticipating human motion from an egocentric perspective is fundamental for proactive assistance in AR/VR, human-robot collaboration, and embodied AI. While recent works incorporate language asa semantic prior to reduce the ill-posed nature of egocentric forecasting, they largely neglect the 3D spatial and semantic context that governshow motion unfolds, and treat pose and language prediction as separate inference streams. We introduce Ego3DLM, built on two core principles:  \naccurate motion forecasting requires explicit spatial and semantic understanding of the 3D environment, and pose and language must be predicted holistically in a single pass, since motion is inherently tied to the semantic interpretation of actions being performed. Given three-point tracking, 3D scene features, and egocentric video, Ego3DLM simultaneously decodes past pose, future pose, past narration, and future narration in a single autoregressive pass, grounding predicted poses and descriptionsin one another to enforce cross-modal and temporal consistency. We adopt a three-stage training scheme: (1) spatial-semantic scene awareness pretraining; (2) holistic instruction tuning over all four outputs in a single pass; and (3) GRPO-based reinforcement finetuning with intraand inter-modal rewards that directly optimize pose-language fidelity.  \nExperiments on the Nymeria benchmark demonstrate that Ego3DLM achieves state-of-the-art performance across future motion prediction, past motion tracking, and motion description, showing that 3D scene grounding and holistic cross-modal prediction yield physically plausible and semantically coherent motion forecasts. The project page is available at [https://jaewoo97.github.io/Ego3DLM/](https://jaewoo97.github.io/Ego3DLM/) .  \nKeywords: Egocentric Vision · Human Motion Forecasting · Multimodal Large Language Models  \n1 Introduction  \nPredicting human motion from an egocentric perspective is fundamental for proactive, real-time assistance in AR/VR, human–robot collaboration, and embodied  \nAI [28,50,51,57,73,76,77,79] . Unlike third-person footage, egocentric observations reflect the wearer’s true field of view and thus encode intent, affordances, and ⋆ Equal contribution.  \n2 Bae et al.  \nPrompt Three-Point Tracking 3D Scene  \n“Perform motion tracking and motion forecasting based on the given 2D video/3D scene and three-point …”  \n2D Video  \n| \u003Cbr>Language Model\u003Cbr> |\n| --- |\n| Tracking (Past)\u003Cbr>\u003Cbr>3D Pose Text\u003Cbr>Forecasting (Future)\u003Cbr>\u003Cbr>3D Pose Text |\n\nFig. 1: Our Ego3DLM incorporates the semantic context of the surrounding 3D environment along with 2D egocentric video and three-point tracking data to generate past and future poses and their corresponding language motion descriptions.  \nnear-field interactions that matter for planning [23, 24, 41] . Yet forecasting from this viewpoint is inherently ill-posed: the camera captures little of the body, only sparse wearable cues are available, and the same partial observation is consistent with many plausible future motions. Resolving this ambiguity requires grounding predictions in semantic context: understanding not merely the kinematic state of the body, but the underlying intent and action semantics that govern how motion will unfold.  \nTowards this end, recent works have sought to reduce forecast ambiguity by incorporating language as a semantic prior on human intention, demonstrating that treating motion as a tokenized sequence enables unified reasoning across vision, language, and pose [32, 38] . Motion-language models such as MotionGPT [38] and EgoLM [32] cast 3D human motion as “motion tokens” for joint modeling with text, while pose-language systems such as ChatPose [22] and Pose","cbCail561fXrUzQk","https://ap.wps.com/l/cbCail561fXrUzQk","pdf",5368219,4,1,36,"English","en",105,"# Introduction\n## Egocentric motion forecasting and ambiguity\n## Prior work with language as semantic priors\n## Prior work with explicit 3D representations\n## Pose-language separation and motivation\n## Core principles and model overview","[{\"question\":\"What problem does Ego3DLM aim to solve in egocentric motion prediction?\",\"answer\":\"It targets the inherent ill-posed ambiguity of egocentric forecasting, where sparse wearable cues correspond to many plausible future motions. Ego3DLM resolves this by grounding predictions in 3D spatial and semantic context.\"},{\"question\":\"How does Ego3DLM connect pose prediction with language?\",\"answer\":\"Ego3DLM predicts past/future poses and past/future narrations holistically in a single autoregressive pass. This design links motion to the semantic interpretation of actions and enforces cross-modal and temporal consistency.\"},{\"question\":\"What training strategy does Ego3DLM use to improve pose–language fidelity?\",\"answer\":\"It follows a three-stage pipeline: spatial-semantic scene awareness pretraining, holistic instruction tuning across all four outputs in one pass, and GRPO-based reinforcement finetuning with intra- and inter-modal rewards that directly optimize pose–language fidelity.\"}]",1784193606,91,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"ego-human-motion-prediction-with-3d-aware-llm","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/ego-human-motion-prediction-with-3d-aware-llm/84168/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Ego3DLM aim to solve in egocentric motion prediction?","Question",{"text":75,"@type":76},"It targets the inherent ill-posed ambiguity of egocentric forecasting, where sparse wearable cues correspond to many plausible future motions. Ego3DLM resolves this by grounding predictions in 3D spatial and semantic context.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Ego3DLM connect pose prediction with language?",{"text":80,"@type":76},"Ego3DLM predicts past/future poses and past/future narrations holistically in a single autoregressive pass. This design links motion to the semantic interpretation of actions and enforces cross-modal and temporal consistency.",{"name":82,"@type":73,"acceptedAnswer":83},"What training strategy does Ego3DLM use to improve pose–language fidelity?",{"text":84,"@type":76},"It follows a three-stage pipeline: spatial-semantic scene awareness pretraining, holistic instruction tuning across all four outputs in one pass, and GRPO-based reinforcement finetuning with intra- and inter-modal rewards that directly optimize pose–language fidelity.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]