[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85902-en":3,"doc-seo-85902-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85902,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Self-Supervised Automatic Matting","High-quality alpha mattes are expensive to annotate, creating a major data bottleneck for deep image matting. Existing approaches reduce cost using trimaps or masks, yet still depend on costly pixel-level supervision, limiting scalability and generalization. This work introduces SSMatte, a self-supervised framework that trains automatic matting using only RGB images without any manual annotation. It decomposes the task into semantic anchoring and detail matting, enabling fully annotation-free performance.","arXiv :2607 . 10395v1 [ cs .CV] 11 Jul 2026  \nSelf-Supervised Automatic Matting  \nXiaonan Hu 1 , Zhiyuan Lu2 , Jingdong Zhao3 , and Hao Lu 1 ,∗  \n1 Huazhong University of Science and Technology, China  \n2 Beijing Normal University, China  \n3 Waseda University, Japan  \n{xiaonah, [hlu}@hust.edu.cn](hlu}@hust.edu.cn)  \nAbstract. High-quality alpha mattes are notoriously expensive to annotate, creating a fundamental data bottleneck for deep image matting.  \nWhile prior work attempts to reduce annotation cost using coarser labels like trimaps or masks, they remain reliant on costly per-pixel supervision, limiting scalability and generalization. In this work, we push the boundary further and ask: can we train an automatic matting model using only RGB images, with no manual annotation at all? We answer this by presenting SSMatte, a self-supervised framework that for the first time achieves performance on par with fully-supervised automatic matting. Our key insight is to decompose the problem into semantic anchoring and detail matting. SSMatte first generates a semantic matting prompt from frozen self-supervised ViT features by propagating classtoken seeds via a novel, training-efficient semantic anchoring loss based on a generalized Rayleigh quotient. This prompt then anchors a detail matting network, which is optimized via a fixed-point-based loss that enforces alpha-RGB consistency. Extensive experiments show SSMatte outperforms prior weakly-supervised methods, matches the performance of fully-supervised models on portrait benchmarks, and demonstrates favorable scaling and generalization behaviors with additional data. Our work pushes automatic matting to an fresh, fully annotation-free paradigm.  \nCode will be available.  \nKeywords: Image Matting · Self-Supervised Learning  \n1 Introduction  \nImage matting aims to estimate the per-pixel foreground opacity α, a.k.a. alpha matte, from an image I, following the matting equation  \nI = αF + (1 − α)B , (1)  \nwhere F and B are foreground and background, respectively. It is crucial for applications requiring high-fidelity foreground extraction. While deep learning has  \nrevolutionized the field [41], its success heavily relies on large-scale, high-quality annotations of alpha mattes, which are extremely costly and time-consuming ⋆ Corresponding author.  \n2 F. Author et al.  \nFig. 1: Comparison between prior training paradigms of deep matting and our self-supervised matting paradigm. Prior paradigms require large-scale dense annotations, while we can train an automatic matting model with only RGB images.  \nto obtain. The nature of alpha matte, however, makes it almost impossible to annotate/generate at scale, particularly in complex real-world scenarios. This annotation bottleneck severely limits the scalability and generalization of deep matting models.  \nPrior effort to alleviate this bottleneck follows two main strands. The first uses synthetic data [30,41,46] or collects subject-specific datasets [19,20,36], yet the domain gap and limited object categories remain issues. The second strand seeks label-efficient supervision, such as training in a semi-supervised manner that uses abundant coarse masks and a few fine alphas [15, 17, 22, 25, 29, 40, 45, 47] . In particular, a notable stride is Alpha-Free Matting (AFM) [26], which eliminates alpha supervision by using only trimaps and a color affinity loss.  \nHowever, a critical limitation persists: existing methods still require per-pixel annotations, be it trimaps, masks, or scribbles. Acquiring such annotations, even if coarser than alpha, remains labor-intensive and impedes large-scale data collection. This leads us to a fundamental, unexplored question: Is it possible to train a competitive automatic matting model with no manual annotation whatsoever, using only RGB images?  \nWe posit that the answer lies in the synergy between modern self-supervised visual representations and the intrinsic image-matte relationship. Recent selfsupervised","cbCaidkH422W5yy7","https://ap.wps.com/l/cbCaidkH422W5yy7","pdf",15588683,5,1,18,"English","en",105,"# Introduction\n## Problem: annotation bottleneck in alpha matting\n## Prior work and remaining limitations\n## Proposed solution: SSMatte\n# Method overview\n## Semantic anchoring from self-supervised ViT features\n## Detail matting via fixed-point consistency loss","[{\"question\":\"Why is data annotation a bottleneck for deep image matting?\",\"answer\":\"High-quality alpha mattes require expensive and time-consuming per-pixel annotations, and their complexity makes large-scale generation difficult. This restricts scalability and generalization of deep matting models.\"},{\"question\":\"What is SSMatte and what training supervision does it use?\",\"answer\":\"SSMatte is a self-supervised framework that trains an automatic matting model using only RGB images with no manual annotation at all. It is designed to match fully-supervised automatic matting performance.\"},{\"question\":\"How does SSMatte connect self-supervised representations to matting outputs?\",\"answer\":\"SSMatte decomposes the task into semantic anchoring and detail matting. It derives a semantic matting prompt from frozen self-supervised ViT features and then enforces alpha-RGB consistency through a fixed-point-based target loss.\"}]",1784207058,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"self-supervised-automatic-matting","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/self-supervised-automatic-matting/85902/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is data annotation a bottleneck for deep image matting?","Question",{"text":76,"@type":77},"High-quality alpha mattes require expensive and time-consuming per-pixel annotations, and their complexity makes large-scale generation difficult. This restricts scalability and generalization of deep matting models.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is SSMatte and what training supervision does it use?",{"text":81,"@type":77},"SSMatte is a self-supervised framework that trains an automatic matting model using only RGB images with no manual annotation at all. It is designed to match fully-supervised automatic matting performance.",{"name":83,"@type":74,"acceptedAnswer":84},"How does SSMatte connect self-supervised representations to matting outputs?",{"text":85,"@type":77},"SSMatte decomposes the task into semantic anchoring and detail matting. It derives a semantic matting prompt from frozen self-supervised ViT features and then enforces alpha-RGB consistency through a fixed-point-based target loss.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]