[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83521-en":3,"doc-seo-83521-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83521,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","LUMA Benchmarking Segmentation via a Lightweight Universal Mask Adapter","Comparing transformer backbones for image segmentation is confounded because each backbone is coupled with different decoders, training recipes, and pretraining, so reported gains rarely isolate the backbone itself. LUMA introduces a Lightweight Universal Mask Adapter, a lightweight, backbone-agnostic mask-transformer head that treats the backbone as a black-box feature extractor. LUMA matches the accuracy of EoMT at lower cost and attaches unchanged to diverse backbone types. With one fixed head, experiments benchmark 20 backbones, 11 pretraining schemes, and multiple resolutions on ADE20K and Cityscapes, showing that efficient token mixers fail to deliver true efficiency and that pretraining objective, not architecture, most strongly governs quality.","LUMA: Benchmarking Segmentation via a Lightweight Universal Mask Adapter  \nTobias Christian Nauen 1 ,2 , Anosh Billimoria 1 , Federico Raue2 , Stanislav Frolov2 , Brian B. Moser2 , Andreas Dengel 1 ,2  \n1 RPTU University Kaiserslautern-Landau, Kaiserslautern, Germany  \n2 German Research Center for Artificial Intelligence (DFKI), Kaiserslautern, Germany  \nfirst   [second.last@dfki.de / first.last@dfki.de](second.last@dfki.de / first.last@dfki.de)  \narXiv :2607 .00687v 1 [ cs .CV] 1 Jul 2026  \nAbstract  \nComparing transformer backbones for image segmentation is confounded: each is paired with a different decoder, recipe, and pretraining, so reported differences rarely reflect the backbone itself. We introduce the Lightweight Universal Mask Adapter (LUMA), a lightweight, backbone-agnostic mask-transformer head that treats any backbone as a blackbox feature extractor, letting a set of queries read from its features through cheap cross-attention. LUMA matches the accuracy of EoMT, the state-of-the-art efficient ViT-segmenter, at lower cost, while attaching unchanged to isotropic, hierarchical, convolutional, and mixture-of-experts backbones alike. Holding this head fixed, we benchmark 20 backbones, 11 pretraining schemes and a range of resolutionson ADE20K and Cityscapes under one modern recipe. We find that “efficient” token mixers fail to deliver efficiency even at the high resolutions that motivate them, with plain ViT holding the throughput Pareto–front at every resolution. Additionally, thepretraining objective, not the architecture, the lever the field has tuned hardest, governs segmentation quality.  \n1. Introduction  \nModern image segmentation is built on the mask transformer framework [11, 12], where a set of learnable queries is matched against the features of a pretrained backbone and decoded into a mask and class per segment. Many efforts have since optimized the mask-based segmentationdecoder [26, 31, 34, 65] . Recently, EoMT [29] showed that this task-specific decoder part of the pipeline is largely unnecessary for plain ViT [17] . With a strong pretraining, the ViT backbone’s own self-attention can process the queries directly. Thus, in modern pipelines, most of the compute sits in the ViT backbone, for which decade of work has now gone into optimizing: efficient and linear attention [40, 55, 59], hierarchical and windowed token mixers [30, 37], convolutional hybrids [7, 32, 57, 68], and state-space models [36];  \nalmost all of these justified by efficiency at high resolution. This invites the question: Is the common backbone architecture of ViT itself necessary, or do pretraining and scale dominate dense transfer, and we can thus reap the efficiency gains of efficient backbones?  \nAnswering this requires comparing backbones on equal footing: One segmentation head, that attaches to any backbone, adds only little compute, and reaches state-of-the-art accuracy. Trained with one training recipe, with only the backbone varying, so that measured differences reflect the architecture rather than the surrounding apparatus. Existing comparisons do not meet this bar, as each backbone is typically paired with its own decoder, recipe, pretraining, and resolution [21, 33] and the heavy task-specific decoders that dominate the leaderboard confound the backbone with the capacity of the head itself [10, 12] . EoMT, sporting only avery lightweight mask-head, is the natural fair instrument, but its efficiency comes from concatenating the queries into the token sequence and processing them with the backbone’sown attention, which presumes global attention, reach-in access to the block internals, and a constant token width. It is therefore a plain-ViT method by construction, and cannot be attached to the very backbones we wish to study.  \nWe close this gap with the Lightweight Universal Mask Adapter (LUMA), a lightweight, efficient segmentation head that treats the backbone as a black-box feature extractor. Rather than inserting queries ","cbCaikvJOtuyC75B","https://ap.wps.com/l/cbCaikvJOtuyC75B","pdf",982445,5,1,17,"English","en",105,"# Introduction\n## Fair comparison of segmentation backbones\n## LUMA design and head decoupling\n## Controlled benchmarking study\n## Findings on efficient token mixers and pretraining","[{\"question\":\"What problem does LUMA address in benchmarking transformer backbones for segmentation?\",\"answer\":\"Reported backbone differences are often confounded by using different decoders, training recipes, and pretraining for each backbone. LUMA introduces a single universal head so measured differences reflect the backbone architecture itself.\"},{\"question\":\"How does LUMA integrate with a backbone while remaining backbone-agnostic?\",\"answer\":\"LUMA keeps learnable queries in a separate side stream and performs lightweight cross-attention to patched features tapped from each block. The backbone runs unchanged, so the same head can attach to multiple backbone families.\"},{\"question\":\"What do the results suggest about efficient token mixers and segmentation quality?\",\"answer\":\"The study finds that “efficient” token mixers do not deliver meaningful throughput efficiency in the high-resolution regimes relevant to segmentation. Accuracy differences largely vanish when using the same head, and the pretraining objective—more than architecture—most strongly governs segmentation quality.\"}]",1784188599,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"luma-benchmarking-segmentation-via-a-lightweight-universal-mask-adapter","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/luma-benchmarking-segmentation-via-a-lightweight-universal-mask-adapter/83521/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does LUMA address in benchmarking transformer backbones for segmentation?","Question",{"text":76,"@type":77},"Reported backbone differences are often confounded by using different decoders, training recipes, and pretraining for each backbone. LUMA introduces a single universal head so measured differences reflect the backbone architecture itself.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does LUMA integrate with a backbone while remaining backbone-agnostic?",{"text":81,"@type":77},"LUMA keeps learnable queries in a separate side stream and performs lightweight cross-attention to patched features tapped from each block. The backbone runs unchanged, so the same head can attach to multiple backbone families.",{"name":83,"@type":74,"acceptedAnswer":84},"What do the results suggest about efficient token mixers and segmentation quality?",{"text":85,"@type":77},"The study finds that “efficient” token mixers do not deliver meaningful throughput efficiency in the high-resolution regimes relevant to segmentation. Accuracy differences largely vanish when using the same head, and the pretraining objective—more than architecture—most strongly governs segmentation quality.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]