[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82182-en":3,"doc-seo-82182-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82182,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Subtoken Vision Transformer for Fine-grained Recognition","Subtoken Vision Transformer (SubViT) introduces selective image tokenization for fine-grained visual recognition by allocating extra representational capacity only to discriminative patches. While standard Vision Transformers compress each fixed patch into a single token, SubViT represents selected patches with multiple subtokens yet preserves the original token sequence for global context. A two-stage training strategy fine-tunes using attention-derived subdivision patterns and distills informative attention maps into a lightweight router for deterministic token-importance scoring. Evaluations on GCD and multiple benchmarks show higher novel-category accuracy with minimal added latency.","arXiv :2607 .09086v1 [ cs .CV] 10 Jul 2026  \nSubtoken Vision Transformer for Fine-grained  \nRecognition  \nJie Zhu 1 , Ivy Zhang2 , Minchul Kim 1 , and Xiaoming Liu 1 ,3  \n1 Michigan State University 2 Cranbrook Kingswood School  \n3 University of North Carolina at Chapel Hill  \n[zhujie4@msu.edu](zhujie4@msu.edu) [izhang27@cranbrook.edu](izhang27@cranbrook.edu) [kimminc2@msu.edu](kimminc2@msu.edu)[ ](kimminc2@msu.edu)[liuxm@cs.unc.edu](liuxm@cs.unc.edu)  \nAbstract. We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition.  \nStandard Vision Transformers compress each fixed-size patch into a single token, although fine-grained distinctions often depend on localized variations within only a few patches. SubViT addresses this mismatch by representing discriminative patches with multiple subtokens while retaining the original token sequence for global context, thereby allocating additional capacity where it is most needed. Since attention heads encode complementary semantics and extracting attention maps at inference requires an extra backbone forward, we adopt a two-stage training strategy. Stage 1 fine-tunes the ViT using subdivision regions sampled from random attention heads, exposing the model to diverse subdivision patterns. Stage 2 identifies informative attention maps through featuredegradation distances and distills them into a lightweight single-map router, which directly predicts deterministic token-importance scores without a separate attention forward. We evaluate SubViT on Generalized Category Discovery (GCD), a challenging task requiring both finegrained discrimination and generalization to unlabeled novel categories.  \nAcross CUB, FGVC-Aircraft, and Stanford Cars, SubViT improves the average novel-category accuracy of DINOv2 from 81.3% to 84.7%, with only 0.50 ms additional latency and 3.4% more FLOPs, while reducing latency by 73.8% relative to Retina Patch. Results on CIFAR-10 and ImageNet-100 demonstrate its broader applicability.  \n1 Introduction  \nVision Transformers (ViTs) [1] represent an image as a sequence of patch tokens and use self-attention to model their global relationships. Standard ViTs uniformly partition the image into fixed-size patches and project every patch into a single token. This design is effective for capturing global structure, but it imposes the same representational granularity on all regions: spatial variations within each patch are compressed into one embedding regardless of their semantic importance. Such uniform allocation is poorly matched to fine-grained recognition [2–6], where the distinction between visually similar subcategories may depend on a small texture, contour, or object part. Uniformly refining the  \n2 Jie Zhu, Ivy Zhang, Minchul Kim, and Xiaoming Liu  \n(c) MsViT (d) SubViT (Ours)  \nFig. 1: Motivation of SubViT. Instead of densely tokenizing multiple image scales, SubViT identifies discriminative patches using attention supervision during training and a learned router at inference. Attention-based Token Subdivision (ATS) allocates additional subtokens only to selected patches while retaining the original tokens for global context.  \nentire image could preserve more local structure, but would substantially increase the token sequence and computation. The central challenge is therefore to increase token-level representational capacity only where fine-grained evidence is likely to occur, while retaining the global context provided by the original token sequence.  \nRecent work has explored spatially adaptive tokenization. Retina Patch in SapiensID [7] uses a keypoint predictor to construct region-level crops and concatenates tokens from multiple resized views, introducing external localization priors and a relatively dense token sequence. MSViT [8] learns to choose between coarse and fine tokenization scales across image regions. These approaches demonstrate the value of non-uniform visual processing, but lea","cbCaiqCgJiSgXjEM","https://ap.wps.com/l/cbCaiqCgJiSgXjEM","pdf",1754693,1,18,"English","en",105,"# Introduction\n## Problem: mismatch between uniform tokenization and fine-grained evidence\n## Prior adaptive tokenization approaches and unresolved trade-offs\n## Proposed method: SubViT and Attention-based Token Subdivision (ATS)","[{\"question\":\"What is the core idea behind SubViT?\",\"answer\":\"SubViT selectively subdivides discriminative patches into multiple subtokens while keeping the original token sequence to preserve global context.\"},{\"question\":\"How does SubViT decide which patches to subdivide?\",\"answer\":\"It uses an attention-based token subdivision mechanism: during training it explores diverse subdivision patterns from sampled attention heads, then distills informative attention into a lightweight router that outputs deterministic token-importance scores at inference.\"},{\"question\":\"What tasks and datasets are used to evaluate SubViT?\",\"answer\":\"SubViT is evaluated on Generalized Category Discovery (GCD) and benchmark datasets including CUB, FGVC-Aircraft, Stanford Cars, CIFAR-10, and ImageNet-100.\"}]",1784178649,45,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"subtoken-vision-transformer-for-fine-grained-recognition","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/subtoken-vision-transformer-for-fine-grained-recognition/82182/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the core idea behind SubViT?","Question",{"text":75,"@type":76},"SubViT selectively subdivides discriminative patches into multiple subtokens while keeping the original token sequence to preserve global context.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does SubViT decide which patches to subdivide?",{"text":80,"@type":76},"It uses an attention-based token subdivision mechanism: during training it explores diverse subdivision patterns from sampled attention heads, then distills informative attention into a lightweight router that outputs deterministic token-importance scores at inference.",{"name":82,"@type":73,"acceptedAnswer":83},"What tasks and datasets are used to evaluate SubViT?",{"text":84,"@type":76},"SubViT is evaluated on Generalized Category Discovery (GCD) and benchmark datasets including CUB, FGVC-Aircraft, Stanford Cars, CIFAR-10, and ImageNet-100.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]