[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84588-en":3,"doc-seo-84588-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84588,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","AdaBoosting Text Prompts for Vision-Language Models","The classification accuracy of pretrained Vision-Language Models depends on the quality of text prompts. Handcrafted templates and LLM-generated descriptions improve interpretability and reuse across heterogeneous VLMs, yet existing few-shot prompt methods do not explicitly target misclassified samples, yielding limited gains as shot counts increase. Text Prompt Boosting (TPB) treats each prompt-based classifier as a weak learner and sequentially ensembles them by reweighting hard errors. Experiments across eleven benchmarks show accuracy improvements on the source model and robust transfer that preserves shot-driven gains.","arXiv :2607 .00684v 1 [ cs .LG] 1 Jul 2026  \nAdaBoosting Text Prompts for Vision-Language Models  \nSeokhee Jin 1 ⋆, Changhwan Sung2⋆, Sunung Mun2, Hoyoung Kim3, and  \nJungseul Ok2†  \n1 KT Corporation, Republic of Korea  \n2 Pohang University of Science and Technology (POSTECH), Republic of Korea  \n3 National AI Research Lab, Republic of Korea  \n[seokhee.jin@kt.com](seokhee.jin@kt.com) , {changhwan.sung, mtablo, [hoyoung.kim](hoyoung.kim) , [jungseul.ok}@postech.ac.kr](jungseul.ok}@postech.ac.kr)  \n[https://sung0503.github.io/TPB](https://sung0503.github.io/TPB)  \nAbstract. The classification accuracy of pretrained Vision-Language Models (VLMs) relies on the quality of the text prompts. Handcrafted templates and Large Language Model (LLM)-generated descriptions not only make predictions more interpretable, but also enable reuse of the same prompts across heterogeneous VLMs. Recent works construct taskadapted text prompts with a small number of labeled images. However, existing few-shot text prompting methods do not explicitly focus on misclassified examples during prompt construction, leading to only marginal improvements even as more shots become available. To fully exploit few-shot supervision, we propose Text Prompt Boosting (TPB), an AdaBoost-inspired framework that treats each text-prompt-based classifier as a weak learner and sequentially aggregates them into a strong ensemble by explicitly targeting hard, misclassified examples. Extensive experiments show that TPB preserves task-intrinsic, model-agnostic cuesin text space, enabling robust cross-model transfer. Across eleven classification benchmarks, TPB improves accuracy on the source model and preserves shot-driven gains when transferred to larger, more capable VLMs, where existing methods struggle to sustain such improvements.  \nKeywords: Vision-Language Models · Few-Shot Adaptation · CrossModel Transferability  \n1 Introduction  \nVision-Language Models (VLMs) [4, 5, 9, 17, 21, 34, 41, 42] such as CLIP [34] enable zero-shot image classification using natural-language text prompts that describe each class. In practice, the classification accuracy is sensitive to the quality of the prompts. To avoid labor-intensive manual prompt engineering,  \nprior studies [25, 32, 35, 36, 45] have explored LLM-based automated prompt ⋆ These authors contributed equally to this work.  \n†Corresponding author.  \n2 S. Jin et al.  \nAvg . Acc . (%)  \n84  \n80  \n76  \n72  \n68  \n64  \n\n| \u003Cbr>ViT-B/32 → ViT-L |  |  |  |\n| --- | --- | --- | --- |\n|  |  |  |  |\n|  | \u003Cbr> |  |  |\n\n1 2 4 8 16 Shot  \n(b) Shot scalability  \nFig. 1: Concept of TPB and its transfer-robust shot scalability. (a) While soft prompting (e.g ., CoOp) optimizes continuous prompt vectors via backpropagation and text-based baselines (e.g ., ProAPO) rely on a single prompt set, TPB iteratively buildsa collection of prompt banks to form a strong classifier. (b) Performance comparison upon cross-model transfer from ViT-B/32 to five ViT-L-scale VLMs, where dashed and solid lines indicate source and transfer performance, respectively. While CoOp cannot be directly transferred and ProAPO fails to preserve shot-driven gains, TPB maintains these shot-driven gains while demonstrating superior scalability on the source model.  \ngeneration. A common practice is to query the Large Language Model with a generic instruction (e.g ., “ What does a {class} look like?”) [32], resulting in prompts based solely on the LLM’s prior knowledge. However, such prompts often lack proper grounding in the target visual distribution and suffer from inaccuracies due to class-name ambiguity or hallucinations.  \nRecent studies [22, 33] address this lack of visual grounding by leveraging a small number of labeled images to facilitate the discovery of task-optimized prompts. To exploit these few-shot examples, these methods typically optimize a single aggregate metric (e.g ., overall validation accuracy) as their primary objective for prompt generation. However, the ","cbCaieIgPmpGW1Kv","https://ap.wps.com/l/cbCaieIgPmpGW1Kv","pdf",3860223,3,1,33,"English","en",105,"# Introduction\n## Motivation: prompt quality and grounding\n## Limitation of existing few-shot prompting methods\n## Text Prompt Boosting (TPB) approach\n## Experimental evaluation and contributions","[{\"question\":\"Why does a Vision-Language Model’s classification accuracy depend on text prompts?\",\"answer\":\"In prompt-based VLM classification, each class is described through natural-language text, so prompt quality directly affects how well the model matches the target visual distribution and avoids ambiguity or inaccuracies.\"},{\"question\":\"What problem do existing few-shot text prompting methods have?\",\"answer\":\"They mainly optimize a single aggregate metric such as validation accuracy, which can saturate early on easy samples. This leads to overfitting to the small labeled set and only marginal improvements when more shots are added.\"},{\"question\":\"How does Text Prompt Boosting (TPB) improve prompt learning in few-shot settings?\",\"answer\":\"TPB builds an ensemble of text prompts in an AdaBoost-like manner by reweighting misclassified examples at each boosting round, forcing the method to select new prompts that specifically address hard, previously failed cases. This increases coverage of visual diversity and supports robust cross-model transfer.\"}]",1784196962,83,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"adaboosting-text-prompts-for-vision-language-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/adaboosting-text-prompts-for-vision-language-models/84588/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does a Vision-Language Model’s classification accuracy depend on text prompts?","Question",{"text":75,"@type":76},"In prompt-based VLM classification, each class is described through natural-language text, so prompt quality directly affects how well the model matches the target visual distribution and avoids ambiguity or inaccuracies.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What problem do existing few-shot text prompting methods have?",{"text":80,"@type":76},"They mainly optimize a single aggregate metric such as validation accuracy, which can saturate early on easy samples. This leads to overfitting to the small labeled set and only marginal improvements when more shots are added.",{"name":82,"@type":73,"acceptedAnswer":83},"How does Text Prompt Boosting (TPB) improve prompt learning in few-shot settings?",{"text":84,"@type":76},"TPB builds an ensemble of text prompts in an AdaBoost-like manner by reweighting misclassified examples at each boosting round, forcing the method to select new prompts that specifically address hard, previously failed cases. This increases coverage of visual diversity and supports robust cross-model transfer.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]