[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82181-en":3,"doc-seo-82181-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82181,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Adaptive Latent Trajectory Anchoring for Action Segmentation Dataset Condensation","Dataset condensation for action segmentation synthesizes compact, informative representations of long, untrimmed video datasets. Existing methods based on variational autoencoders and iterative latent optimization are computationally expensive and can yield over-smoothed reconstructions with rigid temporal constraints. This paper introduces a deterministic latent mapping approach using denoising diffusion implicit models. Action segments are modeled as continuous trajectories anchored by sparse latent points on the noise manifold, with an adaptive allocation that reallocates anchoring budget based on segment-wise reconstruction difficulty, achieving strong segmentation performance.","arXiv :2607 .09081v1 [ cs .CV] 10 Jul 2026  \nAdaptive Latent Trajectory Anchoring for Action Segmentation Dataset Condensation  \nArthème Gauthier-Villars 1 ,2 ∗, Guodong Ding 1∗ ,†, and Angela Yao 1  \n1 National University of Singapore, Singapore  \n2 ETH Zurich, Switzerland  \n[agauthier@ethz.ch {dinggd](agauthier@ethz.ch {dinggd) ,[ayao}@comp.nus.edu.sg](ayao}@comp.nus.edu.sg)  \nAbstract. Dataset condensation for action segmentation synthesizes compact, informative representations of long, untrimmed video datasets.  \nThe existing approach relies on Variational Autoencoders and an iterative latent optimization; it is computationally expensive and suffers from over-smoothed reconstructions and rigid temporal constraints. This paper proposes to shift the condensation paradigm from optimizationbased inversion to deterministic latent mapping. By leveraging Denoising Diffusion Implicit Models, we represent action segments as continuous trajectories anchored by sparse latent points in the noise manifold.  \nTo maximize representational efficiency, we introduce an adaptive allocation mechanism that dynamically redistributes the anchoring budget based on segment-wise reconstruction difficulty. Extensive experiments demonstrate that our framework significantly outperforms state-of-theart methods in segmentation performance across common datasets. Notably, our approach achieves performance parity with real data training while maintaining a condensation ratio of 2 .4% on Breakfast dataset.  \nKeywords: Dataset Condensation · Temporal Action Segmentation · Diffusion Models  \n1 Introduction  \nTemporal action segmentation (TAS) [7] targets the task of predicting a semantic label for every frame in a long, untrimmed video. Despite the success of modern TAS architectures [10, 17, 31], their performance remains heavily dependent on the availability of massive, densely-labeled datasets, which introduces significant bottlenecks in terms of storage and training efficiency. To address this, dataset condensation [29] has emerged as a promising direction, seeking to synthesize a compact, highly informative representation of the original data.  \nA recent pioneering study formulates this problem as generative network inversion [5] . A conditional Variational Autoencoder (cVAE) [14] is first trained  \n∗ Equal contribution.  \n† Corresponding author and project lead.  \n2 A. Gauthier-Villars, G. Ding and A. Yao  \n Original  \n Reconstructed  \n Interpolated  \n(a) Direct Optimization [5] (b) Latent Anchoring (Ours)  \nFig. 1: Comparison of TAS condensation paradigms. (a) Direct Optimization: An existing method uses iterative optimizations (Ok , Ok−1) to find optimal latent codes for condensation and rely on these fixed codes for reconstruction, which can cause low reconstruction fidelity. (b) Latent Anchoring: Our method adaptively selects latent anchors (A 1 , A2 ) per segment for condensation and uses latent trajectory interpolation to enable reconstruction of fine grained action dynamics.  \nto model action priors in the feature space and each action segment is then reconstructed by optimizing latent variables to minimize reconstruction error. While effective, this strategy inherits structural constraints: reconstruction fidelity depends on the expressiveness of a variational latent space trained under a unimodal Gaussian prior, and condensation requires iterative per-segment optimization. As a result, fine temporal variations may be over-smoothed. More importantly, the same latent code is reused across consecutive frames during reconstruction, imposing a block-wise temporal structure that may even obscure the continuous evolution of action dynamics.  \nIn this work, we depart from the inversion perspective and revisit TAS condensation through the lens of generative dynamics. Our central observation is that action segments are not arbitrary collections of feature vectors; they exhibit structured, smooth trajectories in representation space. A desirable condensa","cbCaivVZk601pIzM","https://ap.wps.com/l/cbCaivVZk601pIzM","pdf",743856,1,16,"English","en",105,"# Introduction\n## Temporal action segmentation and dataset condensation\n## Limitations of inversion-based condensation\n## Deterministic diffusion and latent mapping approach\n## Diffusion latent trajectory anchoring and reconstruction\n## Adaptive anchoring by reconstruction difficulty","[{\"question\":\"How does the proposed method condense action segments during training and reconstruction?\",\"answer\":\"Each segment is encoded into a diffusion-induced latent trajectory via a deterministic DDIM reverse process and stored as sparse latent anchors. Reconstruction recovers the full feature sequence by interpolating between these anchors, with adaptive anchor density allocated according to segment-wise reconstruction difficulty.\"}]",1784178631,40,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"adaptive-latent-trajectory-anchoring-for-action-segmentation-dataset-condensation","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/adaptive-latent-trajectory-anchoring-for-action-segmentation-dataset-condensation/82181/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How does the proposed method condense action segments during training and reconstruction?","Question",{"text":75,"@type":76},"Each segment is encoded into a diffusion-induced latent trajectory via a deterministic DDIM reverse process and stored as sparse latent anchors. Reconstruction recovers the full feature sequence by interpolating between these anchors, with adaptive anchor density allocated according to segment-wise reconstruction difficulty.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,111,114,119,122,126],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":28,"slug":110},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":45,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]