[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84801-en":3,"doc-seo-84801-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84801,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Learning Probabilistic Embeddings for Unsupervised Action Segmentation","This paper addresses unsupervised temporal action segmentation in long, untrimmed videos, where prior joint representation learning and clustering methods use optimal transport (OT) to generate pseudo labels. Those methods learn deterministic frame embeddings and may get stuck in local optima, overfitting to incorrect pseudo labels. The work proposes probabilistic frame embeddings modeled as Gaussian distributions, sampling before OT-based pseudo-label estimation. Experiments on four benchmarks show results comparable to and sometimes better than state of the art, with MoF gains up to 20.7% and F1 improvements up to 19.0%.","arXiv :2607 .05263v 1 [ cs .CV] 6 Jul 2026  \nLearning Probabilistic Embeddings for Unsupervised Action Segmentation  \nShuai Li 1 , Duc Manh Vu 1 , and Juergen Gall 1 ,2  \n1 University of Bonn, Germany  \n2 Lamarr Institute for Machine Learning and Artificial Intelligence, Germany  \n{lishuai,[gall}@iai.uni-bonn.de](gall}@iai.uni-bonn.de)  \nAbstract. This paper concerns the problem of unsupervised temporal action segmentation for long, untrimmed videos. Recent successful approaches follow a joint representation learning and clustering paradigm, where optimal transport (OT) is adopted to produce pseudo labels for learning frame representations. These approaches alternate between estimating pseudo labels using OT and optimizing the parameters with gradient descent during training, where OT is used for obtaining the final temporal action segmentation. A major limitation of these works is that they learn a deterministic embedding for frame representations. The iterative procedure between learning deterministic embeddings based on pseudo labels and estimating pseudo labels from the learned embedding can thus get quickly stuck in a local optimum. As an alternative, we thus propose to learn a probabilistic embedding for frame representations. The embeddings are modeled by Gaussian distributions and we sample from the distributions before estimating the pseudo labels. We evaluate our approach on several challenging temporal action segmentation datasets and achieve results comparable to, and in some cases, better than the state of the art. Compared to baselines with deterministic embeddings, our approach improves MoF up to 20 .7% and F1-score up to 19 .0% . Our code is available at [https://github.com/derkbreeze/PEOT](https://github.com/derkbreeze/PEOT).  \nKeywords: Unsupervised Temporal Action Segmentation · Probabilistic Embedding · Optimal Transport  \n1 Introduction  \nUnsupervised temporal action segmentation is highly relevant for many applications, such as monitoring and optimizing workflows in manufacturing, phenotyping human or animal behavior, as well as human-robot interaction and collaboration [23, 39] . The task, however, is challenging, as videos can contain different numbers of actions and actions can happen in a different order within the video. Furthermore, the same action can occur multiple times within a video. To tackle this problem, approaches based on joint representation learning and clustering [17, 32] have recently become popular. They iterate during training between estimating pseudo-labels, using optimal transport (OT), and updating the frame embeddings by using the estimated pseudo-labels as target for the cross-entropy  \n2 S. Li et al.  \nFig. 1: Most previous works learn deterministic embeddings as frame embeddings; we propose to learn probabilistic embeddings such that frame representations are samples from Gaussian distributions that explicitly model embedding uncertainty. Our representations lead to more accurate segmentations compared to the baseline [38] . Different colors indicate different actions.  \nloss. Following this paradigm, Xu and Gould [38] recently proposed ASOT. The core idea is to use a combination of Kantorovic and Gromov-Wasserstein optimal transport, which offers temporal consistency. They also relax the balanced action assignment assumption in prior works [17, 32] via an unbalanced OT formulation. Similar to the idea of [32], CLOT [3] extends ASOT by building a three-level OT that introduces feedback between frame embeddings and action embeddings, which improves the detection of short segments.  \nWe notice that all previous unsupervised works [3,16–18,30,32,34] learn a deterministic embedding as frame representations before computing pseudo-labels. This, however, has the disadvantage that the optimization using optimal transport can get very quickly stuck in a local optimum such that the deterministic embedding overfits to the wrong pseudo-labels. In this work, we thus propose to learn prob","cbCaidQP3RD87nyV","https://ap.wps.com/l/cbCaidQP3RD87nyV","pdf",6266475,1,21,"English","en",105,"# Introduction\n## Related Work\n## Method\n## Experiments and Results","[{\"question\":\"What limitation of deterministic embeddings motivates the proposed method?\",\"answer\":\"Previous unsupervised approaches learn deterministic frame embeddings, and the alternating optimization with OT-derived pseudo labels can quickly get stuck in a local optimum, causing overfitting to wrong pseudo labels.\"},{\"question\":\"How do probabilistic embeddings improve pseudo-label estimation?\",\"answer\":\"The method models each frame embedding as a Gaussian distribution and samples features from it, then applies optimal transport (OT) on the sampled features to estimate pseudo labels.\"},{\"question\":\"How is the approach evaluated and how does it perform?\",\"answer\":\"The method is evaluated on four unsupervised temporal action segmentation benchmarks (Breakfast, Youtube Instructional, 50Salads, Desktop Assembly). It achieves results comparable to and sometimes better than state of the art, improving MoF up to 20.7% and F1-score up to 19.0% versus deterministic embedding baselines.\"}]",1784198333,53,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"learning-probabilistic-embeddings-for-unsupervised-action-segmentation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/learning-probabilistic-embeddings-for-unsupervised-action-segmentation/84801/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What limitation of deterministic embeddings motivates the proposed method?","Question",{"text":75,"@type":76},"Previous unsupervised approaches learn deterministic frame embeddings, and the alternating optimization with OT-derived pseudo labels can quickly get stuck in a local optimum, causing overfitting to wrong pseudo labels.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do probabilistic embeddings improve pseudo-label estimation?",{"text":80,"@type":76},"The method models each frame embedding as a Gaussian distribution and samples features from it, then applies optimal transport (OT) on the sampled features to estimate pseudo labels.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the approach evaluated and how does it perform?",{"text":84,"@type":76},"The method is evaluated on four unsupervised temporal action segmentation benchmarks (Breakfast, Youtube Instructional, 50Salads, Desktop Assembly). It achieves results comparable to and sometimes better than state of the art, improving MoF up to 20.7% and F1-score up to 19.0% versus deterministic embedding baselines.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]