[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83549-en":3,"doc-seo-83549-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83549,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","TrajLoc: Trajectory-Attention Localization for Multi-Object Motion Control","Multi-object motion control in image-to-video generation requires preserving object identities while following distinct target trajectories, which becomes harder as object count increases and paths intersect or occlude. Existing methods mix multiple trajectories into dense conditioning signals, weakening object-level correspondence in crowded scenes. TrajLoc replaces cross-attention weights for each object token with Gaussian heatmaps centered on per-frame target locations, using per-object tokens to encode trajectory, depth, and first-frame appearance identity. Evaluations across six datasets and two backbones show improved visual fidelity and tighter trajectory adherence.","arXiv :2607 .0086 1v 1 [ cs .CV] 1 Jul 2026  \nTrajLoc: Trajectory-Attention Localization for Multi-Object Motion Control  \nOmer Sela 1 ,2 Inbar Huberman-Spiegelglas 1 Michael Rotman 1 Sagie Benaim 1 Avi Ben-Cohen 1  \n1Amazon Prime Video 2Tel Aviv University  \n􀂀 [sela-omer.github.io/traj](sela-omer.github.io/traj)-loc  \nAbstract  \nControlling the motion of multiple objects in image-to-video (I2V) generation requires preserving object identities while enforcing adherence to distinct target trajectories. This becomes particularly challenging as the number of objects increases and their paths intersect or occlude one another. Existing approaches entangle multiple trajectories within a shared, dense conditioning signal, making object-level correspondence difficult to preserve in crowded scenes. We depart from this paradigm and enforce a strict, per object spatial constraint that isolates instances independently. Our method, TrajLoc, achieves this directly within the attention layers by substituting the cross-attention weights of each object token with a Gaussian heatmap centered on its target location at every frame. The same per object token interface carries trajectory and depth through a learned embedding and preserves identity by encoding first frame appearance in place of an object token. Evaluations across six datasets, featuring up to 20 simultaneously controlled objects and out of distribution real world scenes, demonstrate that our method consistently improves both visual fidelity and trajectory adherence. Applied to two architecturally distinct backbones (CogVideoX 5B and WaN 2.1 14B), our approach achieves average gains of +4.3 dB PSNR and a 51% reduction in trajectory end point error compared to the strongest baselines.  \n1 Introduction  \nModern video diffusion models generate highly realistic videos [37, 26, 43], yet controlling the motion of multiple interacting objects, each following a distinct trajectory while preserving identity, remains a significant challenge. Generating videos with precise object control is essential for applications such as simulation and synthetic data generation for autonomous driving [13, 32, 11], robot learning [6, 19], and professional video editing [35, 31], where multiple objects must follow prescribed motion patterns while maintaining visual consistency. As the number of controlled objects grows and trajectories begin to interact, preserving object-level correspondence becomes increasingly difficult, often leading to identity confusion and motion drift.  \nWhile recent work has made progress on trajectory-conditioned generation [40, 4, 28, 20, 34, 33], existing methods struggle to scale to crowded scenes with many moving objects and are therefore typically evaluated on only one to three objects with limited interactions. This limitation is tied to how trajectory control is represented. The user input is inherently sparse: each object trajectory is specified by only a 2D coordinate per frame, requiring at most 2T scalar values for a video of T frames. Existing methods inflate this into dense tensors of size up to T×H×W×C, whether through rasterized motion patches [40], latent feature propagation [4, 28], rendered trajectory videos [20], or dense conditioning volumes [10, 29] . Because multiple trajectories are fused within the same dense signal, object-level correspondence degrades under occlusions and crossings. The dense  \nPreprint.  \nFigure 1: TrajLoc. Given a first frame and a set of target trajectories (left column, with colored polylines), the goal is to generate a video that moves each object along its prescribed path while preserving its visual identity. Top: multiple pedestrians on a synthetic urban scene. Bottom: sheep in a natural outdoor scene. The remaining columns show three uniformly spaced generated frames with the ground-truth position of each object marked as a colored dot. In both scenes, our method places each object on its target dot, showing accurate trajectory adhe","cbCaijUMigs71Fgl","https://ap.wps.com/l/cbCaijUMigs71Fgl","pdf",18893452,4,1,26,"English","en",105,"# Abstract\n# Introduction\n## Challenges in multi-object trajectory control\n## Limitations of dense conditioning representations\n## Core insight and TrajLoc approach\n## Evaluation setup and datasets","[{\"question\":\"What problem does TrajLoc target in multi-object image-to-video generation?\",\"answer\":\"It targets maintaining object identity while generating video motion that follows separate prescribed trajectories for multiple objects, especially when many objects intersect or occlude.\"},{\"question\":\"How does TrajLoc enforce per-object trajectory constraints inside the model?\",\"answer\":\"It modifies cross-attention by replacing each object token’s cross-attention weights with Gaussian heatmaps centered on the object’s target location at every frame, isolating each object independently.\"},{\"question\":\"How does TrajLoc preserve visual identity and depth information beyond 2D trajectory location?\",\"answer\":\"It injects learned per-object representations into the conditioning space using dedicated tokens: a trajectory token encoding motion and depth, and an appearance token encoding the first-frame visual identity.\"}]",1784188758,66,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"trajloc-trajectory-attention-localization-for-multi-object-motion-control","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/trajloc-trajectory-attention-localization-for-multi-object-motion-control/83549/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does TrajLoc target in multi-object image-to-video generation?","Question",{"text":75,"@type":76},"It targets maintaining object identity while generating video motion that follows separate prescribed trajectories for multiple objects, especially when many objects intersect or occlude.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does TrajLoc enforce per-object trajectory constraints inside the model?",{"text":80,"@type":76},"It modifies cross-attention by replacing each object token’s cross-attention weights with Gaussian heatmaps centered on the object’s target location at every frame, isolating each object independently.",{"name":82,"@type":73,"acceptedAnswer":83},"How does TrajLoc preserve visual identity and depth information beyond 2D trajectory location?",{"text":84,"@type":76},"It injects learned per-object representations into the conditioning space using dedicated tokens: a trajectory token encoding motion and depth, and an appearance token encoding the first-frame visual identity.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]