[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83840-en":3,"doc-seo-83840-105":30,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83840,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Mask2Real-WM Segmentation Masks as a Sim to Real Bridge for Controllable Dexterous World Models","Action-conditioned world models enable robots to forecast the consequences of candidate actions without extra physical interaction, improving planning, policy evaluation, and data augmentation. Mask2Real-WM presents a two-stage model for dexterous manipulation: a dynamics module predicts future segmentation masks from past masks and 23-DoF action sequences, while a rendering module converts predicted masks into photorealistic two-view RGB using a ControlNet-augmented Stable Video Diffusion backbone. Simulation pretraining in segmentation space reduces the sim-to-real gap, enabling fine-tuning with under 2.5 hours of real demonstrations and stronger per-DoF action controllability than monolithic baselines.","Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World  \nModels  \narXiv :2607 .04546v 1 [ cs .RO] 5 Jul 2026  \nRiccardo O. Feingold Davide Liconti Chenyu Yang Robert K. Katzschmann  \nSoft Robotic Lab, Department of Mechanical and Process Engineering ETH Zurich, Switzerland  \nCorrespondence: [rfeingold@ethz.ch](rfeingold@ethz.ch)  \nProject Page: [https://srl-ethz.github.io/Mask2Real-WM/](https://srl-ethz.github.io/Mask2Real-WM/)  \nFigure 1: Mask2Real-WM. A controllable world model that decouples dynamics from rendering: a Dynamics WM predicts future segmentation masks from past masks and past/future actions (6-DoF end-effector pose + 17-DoF hand joints) and is pretrained on >50h of simulation data; a Rendering WM paints photorealistic RGB onto the predicted masks and is trained on \u003C2 .5h of real demonstrations. The combined model supports long-horizon autoregressive rollouts and faithful policy evaluation on dexterous tasks.  \nAbstract: Action-conditioned world models allow robots to predict the future consequences of candidate actions without additional physical interaction, supporting policy evaluation, planning, and data augmentation. We present Mask2Real-WM, a two-stage action-conditioned world model for dexterous manipulation that decouples pixel prediction into a dynamics model and a rendering model. The dynamics model predicts future segmentation masks from past masksand 23-DoF action sequences. The rendering model maps the predicted masks to photorealistic RGB using a ControlNet-augmented Stable Video Diffusion backbone. The smaller sim-to-real gap in segmentation space enables the dynamics model to benefit from large-scale pretraining on over 50 h of synthetic simulation data, followed by fine-tuning on fewer than 2.5 h of real demonstrations. Experiments on a dexterous pick-and-place benchmark show that mask conditioning and simulation pretraining are both required for per-DoF action controllability across all 23 degrees of freedom. In contrast, monolithic baselines capture broad hand and end-effector trajectories but do not reliably reflect fine-grained, per-joint action effects.  \nKeywords: World Models, Dexterous Manipulation, Simulation  \n1 Introduction  \nWorld models are increasingly used in robot learning [1] . In this setting, they commonly take two forms: policy backbones (World Action Models: WAMs) that jointly predict actions and observations, and action-conditioned world models that predict future states given actions. The latter support policy evaluation [2, 3], planning [4], and dreaming [5] .  \nDexterous hands [6, 7, 8, 9, 10] enable more versatile manipulation than parallel-jaw grippers, but action-conditioned world models for dexterous manipulation remain challenging to build. The action space is high-dimensional (23 DoF in our setup), large-scale dexterous datasets are scarce [11, 12], and accurately predicting finger–object contacts remains difficult for existing video models, which often blur these regions.  \nWe address these challenges with Mask2Real-WM, an action-conditioned world model trained on fewer than 2.5 h of real interaction. Mask2Real-WM decouples pixel prediction into two stages: (i) an action-conditioned dynamics model (WM1) that predicts future segmentation masks from past masks and the action sequence, and (ii) a rendering model (WM2) that produces photorealistic twoview RGB conditioned on the predicted masks (see Fig. 2) . Because segmentation masks have a smaller sim-to-real gap than RGB images, WM1 can be pretrained on large amounts of synthetic simulation data and then fine-tuned with minimal real data. WM2 learns appearance from the small real dataset. This use of simulation pretraining is relatively unexplored for image-space world models. Experiments show that the resulting decomposition produces sharper predictions and improves fine-grained controllable video generation than monolithic baselines.  \nContributions.  \n1. We use segmentation spa","cbCairzaoyP5N3Xt","https://ap.wps.com/l/cbCairzaoyP5N3Xt","pdf",21046909,9,1,23,"English","en",105,"# Introduction\n## Related Work\n# Contributions","[{\"question\":\"Why is segmentation space used as the sim-to-real bridge?\",\"answer\":\"Segmentation masks have a smaller sim-to-real gap than RGB images, letting WM1 benefit from large-scale pretraining on synthetic simulation data. WM2 then learns appearance from a small real dataset, enabling effective fine-tuning with fewer than 2.5 hours of demonstrations.\"}]",1784190905,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":28},"mask2real-wm-segmentation-masks-as-a-sim-to-real-bridge-for-controllable-dexterous-world-models","",{"@graph":36,"@context":78},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/mask2real-wm-segmentation-masks-as-a-sim-to-real-bridge-for-controllable-dexterous-world-models/83840/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"Why is segmentation space used as the sim-to-real bridge?","Question",{"text":76,"@type":77},"Segmentation masks have a smaller sim-to-real gap than RGB images, letting WM1 benefit from large-scale pretraining on synthetic simulation data. WM2 then learns appearance from a small real dataset, enabling effective fine-tuning with fewer than 2.5 hours of demonstrations.","Answer","https://schema.org",{"og:url":52,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":85},[86,90,94,98,103,108,113,116,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":99,"doc_module":4,"doc_module_name":46,"category_name":100,"show_sort_weight":101,"slug":102},5,"Comic",60,"comic",{"id":104,"doc_module":4,"doc_module_name":46,"category_name":105,"show_sort_weight":106,"slug":107},6,"Technology",50,"technology",{"id":109,"doc_module":4,"doc_module_name":46,"category_name":110,"show_sort_weight":111,"slug":112},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":114,"slug":115},30,"research-report",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":99,"slug":130},19,"General","general"]