[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83204-en":3,"doc-seo-83204-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83204,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Attention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation","Video object segmentation (VOS) requires accurate delineation and consistent object tracking across frames. Supervised approaches depend on densely annotated datasets that are costly and limited in domain coverage, while existing self-supervised methods often struggle to balance spatial accuracy and temporal coherence in unconstrained multi-object settings. The proposed CTC2 framework learns mid-level, part-aware representations using attention-guided token selection and lightweight cross-temporal clustering, aligning soft part assignments with a saliency-weighted symmetric consistency objective. Using a frozen transformer backbone and lightweight modules, CTC2 scales efficiently without motion cues or domain-specific adaptation, maintains real-time throughput, shows stable cross-dataset behavior, and can extend to semi-supervised learning with a first-frame mask.","arXiv :2607 .07230v 1 [ cs .CV] 8 Jul 2026  \nAttention-Guided Cross-Temporal Clustering for Self-Supervised Video Object Segmentation  \nWaqas Arshid 1 , Mohammad Awrangjeb 1*, Alan Wee-Chung Liew 1 , Yongsheng Gao2  \n1 School of Information and Communication Technology, Griffith University, Brisbane, QLD, Australia.  \n2 School of Engineering and Built Environment – Electrical and Electronic Engineering, Griffith University, Brisbane, QLD, Australia.  \n*Corresponding author(s) . E-mail(s):  \n[mohammad.awrangjeb@griffith.edu.au](mohammad.awrangjeb@griffith.edu.au) ; [Contributing authors: waqas.arshid@griffith.edu.au](Contributing authors: waqas.arshid@griffith.edu.au); [a.liew@griffith.edu.au](a.liew@griffith.edu.au) ; [yongsheng.gao@griffith.edu.au](yongsheng.gao@griffith.edu.au) ;  \nAbstract  \nVideo object segmentation (VOS) is a fundamental task in video understanding, requiring accurate delineation and consistent tracking of objects across frames. While supervised methods achieve strong performance, they depend on densely annotated datasets that are costly to obtain and limited in domain coverage. Self-supervised learning offers a promising alternative by removing the need for manual labels; however, existing approaches often struggle to jointly maintain spatial accuracy and temporal coherence, particularly in unconstrained multi-object scenarios. Many rely on optical flow, synthetic motion cues, or task-specific pretraining, limiting scalability and generalisation. We propose a self-supervised framework, Cross-Temporal Consistency and Clustering (CTC2 ), that learns mid-level, part-aware representations by combining attention-guided token selection with lightweight temporal clustering. Instead of operating atthe pixel or whole-object level, the method aligns soft part assignments across time using a saliency-weighted symmetric consistency objective. The framework leverages a frozen transformer backbone with lightweight modules for adaptive token selection and multi-offset temporal alignment, enabling efficient scaling across resolutions and motion patterns. CTC2 achieves competitive performance among recent self-supervised methods while maintaining real-time throughput, without relying on motion cues or domain-specific adaptation. It further demonstrates stable behaviour under cross-dataset evaluation and can be extended  \n1  \nto a semi-supervised setting using a first-frame mask. These results suggest that attention-guided token selection combined with temporal clustering offers a practical and scalable direction for label-free video segmentation.  \nKeywords: video object segmentation, self-supervised learning, unsupervised  \nrepresentation learning, vision transformers, saliency-guided attention, part-level  \nrepresentation, temporal consistency  \nAccepted for publication in Machine Intelligence Research. DOI: 10.1007/s11633-026-1648-7 .  \n1 Introduction  \nUnderstanding how objects evolve over time—how they move, deform, interact, or become occluded—is central to visual intelligence. Video Object Segmentation (VOS) supports applications such as autonomous navigation, intelligent video editing, augmented reality, and human–robot interaction [1, 2] . While supervised methods achieve strong performance using dense per-frame annotations [3, 4], they scale poorly: pixelaccurate labels remain costly to obtain at scale, and large annotated datasets are not always available for new domains or rare categories. This bottleneck limits deployment in settings with privacy constraints, long-tail categories, or rapidly changing domains.  \nTo mitigate annotation dependence, recent work has increasingly explored selfsupervised VOS, which aims to learn spatiotemporal representations directly from unlabeled videos [5–7] . Many methods replace human labels with intrinsic signals from temporal coherence, motion consistency, or appearance regularities. However, existing approaches vary widely in the assumptions they require, ranging from moti","cbCaiifyq7v8yGON","https://ap.wps.com/l/cbCaiifyq7v8yGON","pdf",1446125,1,34,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does the paper address in video object segmentation?\",\"answer\":\"The work targets label-efficient VOS by reducing reliance on dense per-frame annotations while preserving accurate object boundaries and consistent tracking across time, especially for unconstrained multi-object videos.\"},{\"question\":\"How does the proposed CTC2 framework achieve temporal consistency without motion cues?\",\"answer\":\"CTC2 uses attention-guided token selection to align soft part assignments across frames and applies a saliency-weighted symmetric consistency objective. Lightweight temporal clustering provides alignment while avoiding optical-flow-based assumptions.\"},{\"question\":\"What is the role of part-level representations in the method?\",\"answer\":\"Instead of modeling whole objects or individual pixels, CTC2 learns mid-level semantic parts that tend to persist through occlusion and compose into objects, improving stability under viewpoint change and intra-class variation.\"}]",1784185934,86,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"attention-guided-cross-temporal-clustering-for-self-supervised-video-object-segmentation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/attention-guided-cross-temporal-clustering-for-self-supervised-video-object-segmentation/83204/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in video object segmentation?","Question",{"text":75,"@type":76},"The work targets label-efficient VOS by reducing reliance on dense per-frame annotations while preserving accurate object boundaries and consistent tracking across time, especially for unconstrained multi-object videos.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed CTC2 framework achieve temporal consistency without motion cues?",{"text":80,"@type":76},"CTC2 uses attention-guided token selection to align soft part assignments across frames and applies a saliency-weighted symmetric consistency objective. Lightweight temporal clustering provides alignment while avoiding optical-flow-based assumptions.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the role of part-level representations in the method?",{"text":84,"@type":76},"Instead of modeling whole objects or individual pixels, CTC2 learns mid-level semantic parts that tend to persist through occlusion and compose into objects, improving stability under viewpoint change and intra-class variation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]