[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128044-en":3,"doc-seo-128044-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128044,962084931830,"Theodore","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Dude - Dual Distribution-Aware Context Prompt Learning For Large Vision-Language Model - Abstract","Prompt learning methods customize large vision-language models to new domains using pre-trained contextual knowledge and limited training data. Existing approaches often optimize unified prompt inputs, which can hinder fine-grained classification when discriminative attributes are insufficient. The work introduces a dual context framework using domain-shared and class-specific contexts, where class context is generated by large language models. It further develops unbalanced optimal transport to align constructed prompts with visual tokens under partial matching, supports training via image augmentation, and demonstrates improved performance on few-shot classification and adapter settings across extensive experiments.","Dude: Dual Distribution-Aware Context Prompt Learning For  \nLarge Vision-Language Model  \nDuy M. H. Nguyen∗ 1 ,2 ,3 , An T. Le∗ 4, Trung Q. Nguyen2 ,5 , Nghiem T. Diep2 , Tai Nguyen2 , Duy Duong-Tran6 ,8 , Jan Peters2 ,4 ,9 , Li Shen8 ,  \nMathias Niepert 1 ,3 , Daniel Sonntag2 ,7  \n1 University of Stuttgart, 2 German Research Center for Artificial Intelligence (DFKI),  \n3 Max Planck Research School for Intelligent Systems, 4 Technical University of Darmstadt  \n5 Technical University of Munich, 6 United States Naval Academy, 7 Oldenburg University  \n8 University of Pennsylvania, 9 Hessian. AI. ∗ Co-equal contribution.  \nEditors: Vu Nguyen and Hsuan-Tien Lin  \nAbstract  \nPrompt learning methods are gaining increasing attention due to their ability to customize large vision-language models to new domains using pre-trained contextual knowledge and minimal training data. However, existing works typically rely on optimizing unified prompt inputs, often struggling with fine-grained classification tasks due to insufficient discriminative attributes. To tackle this, we consider a new framework based on a dual context of both domain-shared and class-specific contexts, where the latter is generated by Large Language Models (LLMs) such as GPTs. Such dual prompt methods enhance the model’s feature representation by joining implicit and explicit factors encoded in LLM knowledge.  \nMoreover, we formulate the Unbalanced Optimal Transport (UOT) theory to quantify the relationships between constructed prompts and visual tokens. Through partial matching, UOT can properly align discrete sets of visual tokens and prompt embeddings under different mass distributions, which is particularly valuable for handling irrelevant or noisy elements, ensuring that the preservation of mass does not restrict transport solutions. Furthermore, UOT’s characteristics integrate seamlessly with image augmentation, expanding the training sample pool while maintaining a reasonable distance between perturbed images and prompt inputs. Extensive experiments across few-shot classification and adapter settings substantiate the superiority of our model over current state-of-the-art baselines.  \nKeywords: prompt learning, adapter learning, unbalanced optimal transport, large visionlanguage models.  \n1. Introduction  \nRecent advancements in vision-language models (VLMs), exemplified by CLIP (Radford et al. , 2021), ALIGN (Jia et al. , 2021), or Flava (Singh et al. , 2022), have demonstrated remarkable capabilities in learning comprehensive visual and textual concepts in classification, generation, or recognition. During pre-training, these models leverage web-scale imagetext pairs to establish aligned representations of images and text through contrastive loss. For instance, through prompts like “A picture of a {label }”, VLMs seamlessly transfer their knowledge into downstream applications, employing zero-shot learning by comparing task-specific descriptions with encoded images and texts (Figure 1 (a)) . Such approaches eliminate the need for extensive fine-tuning, underscoring their adaptability and efficiency in various practical scenarios.  \n© 2024 D.M.H. Nguyen∗ 1 ,2 ,3 et al.  \nNguyen∗ 1 ,2 ,3 et al.  \nFigure 1: (a) Zero-shot learning; (b) Shared classes prompt learning; (c) Our method with dual prompts and Unbalanced Optimal Transport (UOT) as the distance between visual tokens and prompt sets.  \nHowever, the effectiveness of these zero-shot capabilities is highly dependent on the quality of the information embedded in the manually created prompts (Zhou et al. , 2022a) . While significant improvement can be achieved through prompt engineering (Gu et al. , 2023), it is time-consuming, requires domain expertise, and has unpredictable performance under domain shifts. As a direct consequence, data-driven approaches (e.g., prompt learning) are introduced to leverage the rich context of additional information for classification. Early efforts considered a single learnable pro","cbCaitqXfxWXvHr8","https://ap.wps.com/l/cbCaitqXfxWXvHr8","pdf",25351609,3,1,16,"English","en",105,"# Introduction\n## Vision-language model capabilities and prompt-based transfer\n## Limitations of unified prompts and cosine affinity\n## Dual context prompt learning framework\n## Unbalanced optimal transport for prompt-token alignment\n## Experimental validation and results","[{\"question\":\"What problem does dual context prompt learning address in large vision-language models?\",\"answer\":\"It targets fine-grained classification cases where unified, shared prompts lack subtle class-specific discriminative attributes.\"},{\"question\":\"How are class-specific contexts generated in the proposed framework?\",\"answer\":\"Class-specific contexts are generated by large language models such as GPTs.\"},{\"question\":\"What role does unbalanced optimal transport (UOT) play in the method?\",\"answer\":\"UOT quantifies relationships between constructed prompts and visual tokens, enabling partial matching and alignment under differing mass distributions while handling irrelevant or noisy elements.\"}]","Dude - Dual Distribution-Aware Context Prompt Learning For Large Vision-Language Model - Abstract | PDF",1785944411,40,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"dude-dual-distribution-aware-context-prompt-learning-for-large-vision-language-model-abstract","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":20},"https://docshare.wps.com/document/research-report/",{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/dude-dual-distribution-aware-context-prompt-learning-for-large-vision-language-model-abstract/128044/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does dual context prompt learning address in large vision-language models?","Question",{"text":76,"@type":77},"It targets fine-grained classification cases where unified, shared prompts lack subtle class-specific discriminative attributes.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How are class-specific contexts generated in the proposed framework?",{"text":81,"@type":77},"Class-specific contexts are generated by large language models such as GPTs.",{"name":83,"@type":74,"acceptedAnswer":84},"What role does unbalanced optimal transport (UOT) play in the method?",{"text":85,"@type":77},"UOT quantifies relationships between constructed prompts and visual tokens, enabling partial matching and alignment under differing mass distributions while handling irrelevant or noisy elements.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":30,"slug":119},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]