[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83467-en":3,"doc-seo-83467-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83467,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","OnPoint: Offline-to-Online Multi-Level Distillation for Point-Supervised Online Temporal Action Localization","Temporal Action Localization (TAL) is often constrained by segment annotations or offline access to full videos, which limits scalability and online deployment. The work introduces Point-Supervised Online TAL (POTAL), localizing actions in streaming videos using only one temporal point per instance. It presents OnPoint, an offline-to-online multi-level distillation framework transferring knowledge from a point-supervised offline teacher through pseudo-segment distillation, class-activation sequence distillation, and anticipatory window-level distillation, with robustness via point-label integration and actionness-guided attention calibration. Experiments on five datasets show consistent gains over strong baselines.","arXiv :2607 .00289v 1 [ cs .CV] 1 Jul 2026  \n OnPoint: Offline-to-Online Multi-Level Distillation for Point-Supervised Online Temporal Action Localization  \nSakib Reza 1 ,3⋆, Gauri Jagatap2, Mohsen Moghaddam3, Octavia Camps 1, and Andrea Fanelli2   \n1 Northeastern University, Boston, MA 02115, USA {reza.s, [o.camps}@northeastern.edu](o.camps}@northeastern.edu)  \n2 Dolby Laboratories, Inc. , San Francisco, CA 94103, USA {gauri.jagatap, [andrea.fanelli}@dolby.com](andrea.fanelli}@dolby.com)  \n3 Georgia Institute of Technology, Atlanta, GA 30332, USA {sreza32, [mohsen.moghaddam}@gatech.edu](mohsen.moghaddam}@gatech.edu)  \nAbstract. Temporal Action Localization (TAL) typically relies on segment annotations or offline access to full videos, limiting scalability and online use. We introduce Point-Supervised Online TAL (POTAL), which localizes actions in streaming videos using only one temporal point per instance. To solve POTAL, we propose OnPoint, an offline-to-online multi-level distillation framework that transfers knowledge from a pointsupervised offline teacher to an online student via (i) pseudo-segment instance distillation, (ii) class-activation sequence distillation, and (iii) anticipatory window-level distillation. We further improve robustness by incorporating the original point labels into student training and by refining anchor decoding with actionness-guided attention calibration. Experiments on five datasets show OnPoint consistently outperforms strong baselines, establishing a solid foundation for POTAL. †  \nKeywords: Video Understanding · Temporal Action Localization  \n1 Introduction  \nWith the rapid growth of video platforms, Temporal Action Localization (TAL) [42, 45], which identifies action boundaries and class labels in untrimmed videos, has become a central task in computer vision. However, many real-world systems must localize actions directly from streaming video while minimizing annotation cost. Examples include augmented reality (AR) task assistance for surgical training or industrial maintenance [2,51], as well as live sports analytics [37], surveillance [41], and robotics [10] . In such settings, models must make frame-by-frame predictions without future context while operating under limited supervision, as dense action boundary annotation is often costly and infeasible at scale.  \n⋆ Work primarily done during an internship at Dolby Laboratories.† Project Page: [https://sakibreza.github.io/OnPoint/](https://sakibreza.github.io/OnPoint/)  \n2 S. Reza et al.  \nA key challenge arises from the inference regime. Most TAL methods [8, 20, 21,44,46–48,50,52] assume offline access to complete videos, allowing models to leverage future context during training and inference. In contrast, real deployments require online inference, where actions must be localized in streaming video as frames arrive. Recent Online Temporal Action Localization (OnTAL) approaches [12, 34] address this setting but rely on full segment-level supervision during training, limiting scalability in continuously recorded environments where annotation is scarce.  \nA second challenge concerns the supervision cost. Training state-ofthe-art TAL or OnTAL models typically requires dense temporal annotations specifying action start and end boundaries, which are expensive to obtain. Recent large multimodal foundation models [4, 28, 32] show promising zero-shot capabilities but remain unreliable for precise temporal action instance detection and boundary localization (Fig. 1) . Consequently, fully automatic label generation remains impractical for many real-world deployments.  \nA practical alternative is point supervision, where each action instance is annotated with a single timestamp. Prior work [27] and our supplementary study (Supp. Sec. C) show that point labeling can reduce annotation effort by up to 6 × while maintaining reliability. This paradigm has been explored in Point-Supervised TAL (PSTAL) [27, 53], which learns action boundaries from spars","cbCaikSlw0JyjTJb","https://ap.wps.com/l/cbCaikSlw0JyjTJb","pdf",6061210,4,1,34,"English","en",105,"# Introduction\n## Temporal action localization and online constraints\n## Supervision cost and point supervision\n## POTAL task and OnPoint approach","[{\"question\":\"What limitation motivates Point-Supervised Online Temporal Action Localization (POTAL)?\",\"answer\":\"Most TAL methods assume offline access to complete videos and dense segment supervision, which hinders scalable, streaming deployment. POTAL targets streaming inference while reducing annotation cost by using only one temporal point per action instance.\"},{\"question\":\"How does OnPoint transfer knowledge from the offline teacher to the online student?\",\"answer\":\"OnPoint uses offline-to-online multi-level distillation, including pseudo-segment instance distillation, class-activation sequence distillation, and anticipatory window-level distillation, so the online model can learn structured supervision under streaming constraints.\"},{\"question\":\"What techniques improve robustness and decoding quality in the online student?\",\"answer\":\"The framework incorporates the original point labels into student training and refines anchor decoding using actionness-guided attention calibration, improving boundary localization reliability in the online setting.\"}]",1784188171,86,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"onpoint-offline-to-online-multi-level-distillation-for-point-supervised-online-temporal-action-localization","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/onpoint-offline-to-online-multi-level-distillation-for-point-supervised-online-temporal-action-localization/83467/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What limitation motivates Point-Supervised Online Temporal Action Localization (POTAL)?","Question",{"text":75,"@type":76},"Most TAL methods assume offline access to complete videos and dense segment supervision, which hinders scalable, streaming deployment. POTAL targets streaming inference while reducing annotation cost by using only one temporal point per action instance.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does OnPoint transfer knowledge from the offline teacher to the online student?",{"text":80,"@type":76},"OnPoint uses offline-to-online multi-level distillation, including pseudo-segment instance distillation, class-activation sequence distillation, and anticipatory window-level distillation, so the online model can learn structured supervision under streaming constraints.",{"name":82,"@type":73,"acceptedAnswer":83},"What techniques improve robustness and decoding quality in the online student?",{"text":84,"@type":76},"The framework incorporates the original point labels into student training and refines anchor decoding using actionness-guided attention calibration, improving boundary localization reliability in the online setting.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]