[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81981-en":3,"doc-seo-81981-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81981,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","WAM-TTT: Steering World Action Models by Watching Human Play at Test Time","WAM-TTT presents a test-time training approach for steering robot foundation models toward new task variants using raw human videos, without additional robot demonstrations or task-specific fine-tuning. Instead of imitating trajectories, human videos are absorbed into a lightweight adaptive memory inside a frozen World Action Model via self-supervised video prediction. A meta-training stage aligns human demonstrations with robot behaviors using paired human-robot data and a key–value memory reconstruction objective. Deployment adapts memory from unlabeled human videos only, preserving foundation generalization while improving performance across manipulation tasks.","WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time  \narXiv :2607 .06988v2 [ cs .RO] 10 Jul 2026  \nYusen Feng 1 ,2 ,∗ Bingchen Han 1 ,2 ,∗ Jiangran Lyu 1 ,2 ,∗  \nKai Liu2 ,3 Yixin Zheng2 ,3 Yuxuan Wan 1 ,2 Weiheng Liu2 ,3 Sun Han 1 ,2 Ruiqin Li 1 ,2 Yulong Zhang 1 Fangfu Liu4 Xuesong Shi2 Libin Liu 1 ,† Yizhou Wang 1 ,† Zhizheng Zhang2 ,† He Wang 1 ,2 ,†  \n1Peking University 2 Galbot 3 CASIA 4Tsinghua University  \n∗Equal contribution †Corresponding authors  \nOffice  \n•••  \nNo retargeting: Unlabeled Human videos to guide robot actions directly  \nHuman Demonstrations  \nTrain  \nFigure 1: Overview of WAM-TTT. Given unlabeled human demonstrations from diverse environments, WAM-TTT steers a pretrained World Action Model (WAM) without retargeting, robot actions, or human-side annotations. During deployment, human videos are absorbed into lightweight TTT fast weights through self-supervised video prediction, while the pretrained action model remains frozen. The adapted memory then guides robot execution through the WAM’s shared visual-action dynamics, enabling efficient and reusable steering from human demonstrations.  \nAbstract: Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additional robot demonstrations, task-specific fine-tuning, or long-context conditioning. We present WAM-TTT, a test-time training framework for steering world action models from raw human videos. Rather than treating human videos as trajectories to imitate, WAM-TTT absorbs them into a lightweight adaptive memory inside a frozen WAM through self-supervised video prediction. To make this memory useful for control, we introduce a meta-training stage that aligns human demonstrations with robot behaviors using paired human-robot data and a key–value memory reconstruction objective. At test time, only unlabeled human videos are required to adapt the memory, while the pretrained WAM remains frozen. This enables efficient and reusable steering without robot actions, human-side annotations, or task-specific fine-tuning, while preserving the generalization ability of the foundation model. Extensive experiments show that WAM-TTT consistently outperformsin-context human-video conditioning baselines across diverse manipulation tasks and generalization settings.  \nKeywords: World Action Model, Test-time Training, Human Videos  \n1 Introduction  \nRecently, the robotics community has increasingly pursued general-purpose robot foundation models through large-scale pretraining. However, most existing RFMs primarily absorb knowledge into fixed model parameters. Once deployed, their behavior is largely determined by the pretrained weights and a limited conditioning interface, such as language instructions, goal images, or short observation histories[1, 2, 3, 4, 5, 6] . As a result, steering RFMs toward new task variants, object interactions, or user-preferred strategies typically requires collecting additional robot demonstrations or fine-tuning the full model. This limits the flexibility and reusability of RFMs in open-ended deployment settings, where users may wish to quickly specify new behaviors without retraining a robot policy.  \nHuman demonstrations offer a natural and scalable interface for steering RFMs[7, 8, 9, 10, 11, 12]: users can simply show how objects should be handled, without specifying robot actions. Existing methods typically leverage human videos through co-training or fine-tuning with robot data [7, 8, 9, 10, 13, 11, 12], often relying on additional supervision such as hand poses, 3D motion, or retargeted trajectories [2, 14, 3, 15, 4, 6, 16, 17] . Such supervision can be noisy and costly to obtain, while taskspecific fine-tuning may cause catastrophic forgetting and reduce the reusability of the pretrained model. A more direct alternative is to condition robot policies on raw human videos [18], but this requires learning such capabilities during large-scale","cbCaiq0HBferCvXS","https://ap.wps.com/l/cbCaiq0HBferCvXS","pdf",10919587,10,1,28,"English","en",105,"# Introduction\n## Motivation for test-time steering\n## WAM-TTT framework overview\n## Contributions","[{\"question\":\"What problem does WAM-TTT address in robot foundation models?\",\"answer\":\"Steering robot foundation models to new task variants or user-preferred behaviors is difficult because most methods require additional robot demonstrations, task-specific fine-tuning, or long-context conditioning.\"},{\"question\":\"How does WAM-TTT use human videos differently from imitation-based approaches?\",\"answer\":\"WAM-TTT treats human videos as deployment-time memory. The framework absorbs videos into lightweight adaptive memory inside a frozen World Action Model using self-supervised video prediction instead of learning direct trajectory imitation.\"},{\"question\":\"What is updated at test time, and what remains frozen?\",\"answer\":\"At deployment, only unlabeled human videos are used to update the adaptive memory through video prediction, while the pretrained WAM remains frozen. This enables efficient steering without robot actions or human-side annotations.\"}]",1784177404,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"wam-ttt-steering-world-action-models-by-watching-human-play-at-test-time","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/wam-ttt-steering-world-action-models-by-watching-human-play-at-test-time/81981/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does WAM-TTT address in robot foundation models?","Question",{"text":76,"@type":77},"Steering robot foundation models to new task variants or user-preferred behaviors is difficult because most methods require additional robot demonstrations, task-specific fine-tuning, or long-context conditioning.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does WAM-TTT use human videos differently from imitation-based approaches?",{"text":81,"@type":77},"WAM-TTT treats human videos as deployment-time memory. The framework absorbs videos into lightweight adaptive memory inside a frozen World Action Model using self-supervised video prediction instead of learning direct trajectory imitation.",{"name":83,"@type":74,"acceptedAnswer":84},"What is updated at test time, and what remains frozen?",{"text":85,"@type":77},"At deployment, only unlabeled human videos are used to update the adaptive memory through video prediction, while the pretrained WAM remains frozen. This enables efficient steering without robot actions or human-side annotations.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":20,"slug":134},"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]