[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84892-en":3,"doc-seo-84892-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84892,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Structured-Condensed Prompt Tuning in Vision Language Models for Fine-grained Image Recognition","Fine-grained image recognition is hindered by the high cost of expert-level annotation and by limited ability of vision-language models (VLMs) to capture subtle inter-class differences. Prompt tuning helps adapt VLMs, yet existing approaches often treat class labels as independent symbols, ignoring semantic hierarchies and correlations that are essential for discriminating visually similar categories. Structured-Condensed Prompt Tuning (SCPT) introduces Semantic Relation Encoding (SRE) and a Semantic Condensation loss (ScLoss) to model label topology and reduce redundant supervision.","arXiv :2607 .06 185v 1 [ cs .CV] 7 Jul 2026  \nStructured-Condensed Prompt Tuning in Vision-Language Models  \nfor Fine-grained Image Recognition  \nXinda Liu 1 Qinyu Zhang 1 Weiqing Min2  \nGuohua Geng 1 Shuqiang Jiang2  \n1 School of Information Science and Technology, Northwest University, Xi’an, Shaanxi, China  \n2 Laboratory of Intelligent Information Processing, Institute of Computing Technology, Beijing, China  \n[liuxinda@nwu.edu.cn](liuxinda@nwu.edu.cn) , [zhangqinyu@stu.nwu.edu.cn](zhangqinyu@stu.nwu.edu.cn) , [minweiqing@ict.ac.cn](minweiqing@ict.ac.cn)  \n[ghgeng@nwu.edu.cn](ghgeng@nwu.edu.cn) , [sqjiang@ict.ac.cn](sqjiang@ict.ac.cn)  \nAbstract  \nFine-grained image recognition poses a significant challenge due to the substantial expertise and effort required for manual annotation. Vision-language models (VLMs) like CLIP provide a compelling zero-shot alternative, reducing reliance on extensive labeled data. However, their ability to capture subtle distinctions remains limited, leading to subpar recognition performance. While prompt tuning has proven effective for adapting VLMs, most existing methods treat class labels as isolated, discrete entities, overlooking the rich semantic relationships between them. This oversimplified assumption limits the model’s ability to capture hierarchical dependencies and inter-class correlations—both critical for distinguishing visually similar categories. The problem is especially acute in fine-grained classification, where accurate recognition depends on understanding complex label semantics. To address this, we propose Structured-Condensed Prompt Tuning (SCPT), which enhances semantic structure modeling in prompt learning. Specifically, we introduce Semantic Relation Encoding (SRE) to explicitly model inter-class semantic topology and encode structured label relationships. In parallel, we design a Semantic Condensation loss (ScLoss) to suppress redundant supervision and extract discriminative components from the global semantic space. Together, these components significantly improve semantic alignment and fine-grained discrimination. Extensive experiments on 14 fine-grained benchmarks show that SCPT effectively mitigates semantic ambiguity and achieves state-of-the-art performance in both few-shot and base-to-novel generalization settings.  \nKeywords: Prompt tuning; Vision-language models; Semantic relation; Few-shot learning  \n1 Introduction  \nFine-grained image recognition (FGIR) plays a pivotal role in scenarios demanding precise differentiation between visually similar subcategories, rich image captioning [1–3], image generation [4 , 5], food recognition [6–8], and food recommendation [9 , 10] . The essential challenge stems from the requirement for expert-level annotation precision, which necessitates a nuanced understanding of subtle visual differences, often demanding domain-specific knowledge [11] . The exorbitant cost associated with acquiring such human-expert annotations has emerged as a critical bottleneck, that fundamentally constrains the development of FGIR systems in novel domains.  \nVision-language models (VLMs), such as CLIP, offer a promising solution by leveraging large-scale web data for zero-shot learning, thereby mitigating the reliance on extensive manual annotations [12] . CLIP employs contrastive learning to align text and image representations, enabling category  \nMiso Soup  \nSpaghetti Bolognese  \n...  \nChocolate Cake  \n(a) Hand-crafted Prompt  \nMiso Soup  \nSpaghetti Bolognese  \n...  \nChocolate Cake  \n(b) CoOp-style Prompt  \n~~ ~~ Class  \n~~ ~~ Class  \n(c) Structured-Condensed Prompt  \nFigure 1: From static to semantically structured prompts: a paradigm shift in prompting mechanisms.  \nclassification without task-specific training. While VLMs demonstrate strong generalization in broad categories, this advantage diminishes notably in fine-grained recognition contexts, predominantly stemming from the model’s constrained capacity to discern nuanced inter-class var","cbCaidZaFrMPtLPi","https://ap.wps.com/l/cbCaidZaFrMPtLPi","pdf",3888803,2,1,25,"English","en",105,"# Introduction\n## Problem: Fine-grained recognition and annotation cost\n## Vision-language models and zero-shot limitations\n## Prompt tuning and its discrete-label assumption\n## Proposed method: Structured-Condensed Prompt Tuning (SCPT)","[{\"question\":\"Why do vision-language models underperform on fine-grained image recognition?\",\"answer\":\"They struggle to model subtle distinctions between visually similar categories, which requires capturing nuanced inter-class variations and semantic structure.\"},{\"question\":\"What limitation is common in existing prompt tuning methods for fine-grained tasks?\",\"answer\":\"They typically treat class labels as independent discrete entities, overlooking hierarchical dependencies and correlations among labels.\"},{\"question\":\"How does Structured-Condensed Prompt Tuning (SCPT) address this limitation?\",\"answer\":\"SCPT uses Semantic Relation Encoding (SRE) to preserve structured semantic topology among labels and a Semantic Condensation loss (ScLoss) to suppress redundant supervision while extracting discriminative components.\"}]",1784199116,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"structured-condensed-prompt-tuning-in-vision-language-models-for-fine-grained-image-recognition","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/structured-condensed-prompt-tuning-in-vision-language-models-for-fine-grained-image-recognition/84892/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do vision-language models underperform on fine-grained image recognition?","Question",{"text":75,"@type":76},"They struggle to model subtle distinctions between visually similar categories, which requires capturing nuanced inter-class variations and semantic structure.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What limitation is common in existing prompt tuning methods for fine-grained tasks?",{"text":80,"@type":76},"They typically treat class labels as independent discrete entities, overlooking hierarchical dependencies and correlations among labels.",{"name":82,"@type":73,"acceptedAnswer":83},"How does Structured-Condensed Prompt Tuning (SCPT) address this limitation?",{"text":84,"@type":76},"SCPT uses Semantic Relation Encoding (SRE) to preserve structured semantic topology among labels and a Semantic Condensation loss (ScLoss) to suppress redundant supervision while extracting discriminative components.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]