[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83689-en":3,"doc-seo-83689-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83689,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Track the Noise, Move the World: 3D-Grounded Motion-Consistent Noise for Controllable Video Generation","Modern image-and-text-to-video diffusion models generate realistic videos by iteratively denoising an initial Gaussian noise tensor conditioned on reference image and text. Existing methods still struggle to provide precise, unified control over both object motion and camera motion in a single generation. UniCaMo introduces a unified framework that constructs diffusion input noise to enable simultaneous control of object trajectories and camera viewpoints, using 3D-grounded motion-consistent noise space. It applies sparse 3D track-guided noise warping plus globally consistent spherical noise sampling, ensuring geometric and temporal consistency without altering the base model architecture.","arXiv :2607 .02798v 1 [ cs .CV] 2 Jul 2026  \nTrack the Noise, Move the World: 3D-Grounded Motion-Consistent Noise for Controllable Video Generation  \nLong Vu∗ Tan Ngo∗ Animesh Karnewar Amir Habibian  \nBinh-Son Hua† Hung Bui Minh Hoai Nguyen Phong Nguyen-Ha  \nQualcomm AI Research‡  \n{vlong, tanngo, karnewar, ahabibia, hson, hungbui, minhhoai, phongnh}@qti.qualcomm.com  \nAbstract  \nModern image-and-text-to-video diffusion models can synthesize highly realistic videos by iteratively denoising an initial Gaussian noise tensor conditioned on reference image and text inputs. However, existing approaches still lack precise and unified controllability over both object motion and camera motion within a single generation process. We present UniCaMo, a unified framework that enables simultaneous control of object trajectories and camera viewpoints by directly constructing the input noise of the diffusion model. Specifically, UniCaMo buildsa shared 3D-grounded motion-consistent noise space across latent video frames.  \nSparse 3D point tracks are used to warp the Gaussian noise of the reference frame along desired object trajectories, while a virtual spherical noise representation provides globally consistent noise values for newly revealed scene regions under camera motion. By combining local track-guided noise warping with global spherebased noise sampling, UniCaMo maintains geometric and temporal consistency under both object movement and viewpoint changes. Because UniCaMo modifies only the input noise, it requires no auxiliary adapters, control branches, or architectural changes to the underlying video diffusion model. With lightweight LoRA fine-tuning on large pretrained video diffusion models, including Wan 2.1 (14B), UniCaMo achieves state-of-the-art results in both video quality and motion controllability on standard controllable video generation benchmarks.  \n1 Introduction  \nVideo is a highly expressive visual medium, and precise control over generated content enables significant practical value. A scene is defined not only by its objects, but by their motion and the camera’s movement—e.g., a car on a street looks entirely different when tracked, panned, or viewed from above. Beyond storytelling, controllable video generation is crucial for robotics and embodied AI [41, 25], enabling controllable data synthesis during training and predictive imagination at inference time. Motivated by this, we ask: can we build a video generation framework that simultaneously and precisely controls object and camera motion without external adapters? Controllable video generation has gained increasing attention alongside advances in large-scale video diffusion models [40, 32, 20, 47] . Recent works achieve strong control via point tracks, optical flow,  \n∗ Equal contribution.  \n†Binh-Son Hua is affiliated with Trinity College Dublin, Ireland. Work done under consultancy capacity.‡Qualcomm AI Research is an initiative of Qualcomm Technologies, Inc.  \nPreprint.  \nWanMove  \n(2D tracks)  \nOurs  \n(3D tracks)  \nControl Signals Generated 5-second Video  \nText Prompt: A man walks around a woman as the camera arcs to the right.  \nWanMove  \n(2D tracks)  \nOurs  \n(3D tracks)  \n2s 3s 4s 5s  \n2s  \n3s  \n4s  \n5s  \nText Prompt: A women walk to the right behind the man , a man walk to the  \nleft in front of the woman, the camera follows the man.  \n Occlusion-aware  \n Occlusion-aware  \nCamera & Motion Disentanglement  \nCamera & Motion Disentanglement  \nFigure 1: UniCaMo jointly controls object + camera motion with 3D tracks and spherical noise, producing coherent, occlusion-aware videos. WanMove [13] relies on 2D guidance, often causing motion/camera ambiguity and implausible occlusions.  \ncamera trajectories, or auxiliary networks [13, 16, 4, 8, 10, 31] . However, most methods treat object and camera motion separately—either animating objects from a fixed viewpoint or controlling camera motion with unconstrained scene dynamics. When both are specified, this separat","cbCaitTsvLsby2X0","https://ap.wps.com/l/cbCaitTsvLsby2X0","pdf",25957646,5,1,19,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does UniCaMo address in controllable video generation?\",\"answer\":\"UniCaMo targets the lack of precise, unified controllability over both object motion and camera motion within a single generation process, which existing approaches often treat separately and may entangle.\"},{\"question\":\"How does UniCaMo achieve joint control of object and camera motion?\",\"answer\":\"It enables simultaneous control by directly constructing the diffusion model’s input noise, building a shared 3D-grounded motion-consistent noise space across latent video frames.\"},{\"question\":\"What key technical idea keeps motion consistent when the camera viewpoint changes?\",\"answer\":\"UniCaMo combines local track-guided noise warping for desired object trajectories with global spherical noise sampling to maintain globally consistent noise values for newly revealed regions.\"}]",1784189746,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"track-the-noise-move-the-world-3d-grounded-motion-consistent-noise-for-controllable-video-generation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/track-the-noise-move-the-world-3d-grounded-motion-consistent-noise-for-controllable-video-generation/83689/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does UniCaMo address in controllable video generation?","Question",{"text":76,"@type":77},"UniCaMo targets the lack of precise, unified controllability over both object motion and camera motion within a single generation process, which existing approaches often treat separately and may entangle.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does UniCaMo achieve joint control of object and camera motion?",{"text":81,"@type":77},"It enables simultaneous control by directly constructing the diffusion model’s input noise, building a shared 3D-grounded motion-consistent noise space across latent video frames.",{"name":83,"@type":74,"acceptedAnswer":84},"What key technical idea keeps motion consistent when the camera viewpoint changes?",{"text":85,"@type":77},"UniCaMo combines local track-guided noise warping for desired object trajectories with global spherical noise sampling to maintain globally consistent noise values for newly revealed regions.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},"General","general"]