[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-421531-105":59,"doc-detail-421531-en":130},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":123,"head_meta":125,"extra_data":127,"updated_unix":129},105,"en","vocabulary-scaling-law-tuning-open-vocabulary-predictors-for-their-openness","Vocabulary Scaling Law - Tuning Open-vocabulary Predictors for Their Openness","","Open-vocabulary learning on CLIP generalizes across diverse concepts, yet it underperforms in realistic streaming open-world evaluations, especially stability against distractor classes and extensibility to unseen novel classes. The paper introduces a “vocabulary scaling law” that bounds openness measures by performance on the full class-name universe, motivating robust fine-tuning. It proposes Submodular-Vocabulary Fine-tuning (SVFT), a bi-level method that greedily selects an informative subset of class-name embeddings and tunes them with an orthogonality constraint, improving both stability and extensibility through extensive experiments.",{"@graph":69,"@context":122},[70,84,105],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/vocabulary-scaling-law-tuning-open-vocabulary-predictors-for-their-openness/421531/",{"url":83,"name":65,"@type":85,"image":86,"author":91,"headline":65,"publisher":94,"fileFormat":97,"inLanguage":63,"description":67,"dateModified":98,"datePublished":99,"encodingFormat":97,"isAccessibleForFree":100,"interactionStatistic":101},"DigitalDocument",{"url":87,"@type":88,"width":89,"height":90},"https://docshare.wps.com/thumbnails/vocabulary-scaling-law-tuning-open-vocabulary-predictors-for-their-openness/421531.png","ImageObject",300,407,{"name":92,"@type":93},"Adam","Person",{"url":74,"name":95,"@type":96},"DocShare","Organization","application/pdf","2026-09-29","2026-09-28",true,{"@type":102,"interactionType":103,"userInteractionCount":8},"InteractionCounter",{"@type":104},"ViewAction",{"@type":106,"mainEntity":107},"FAQPage",[108,114,118],{"name":109,"@type":110,"acceptedAnswer":111},"What problems does the paper address in open-vocabulary CLIP evaluation?","Question",{"text":112,"@type":113},"It targets failures in streaming open-world settings, focusing on stability against distractor classes and extensibility to novel classes not seen during fine-tuning.","Answer",{"name":115,"@type":110,"acceptedAnswer":116},"What is the “vocabulary scaling law” introduced in the paper?",{"text":117,"@type":113},"It states that openness measures are lower-bounded by performance on the scaled vocabulary containing the full open-set class-name universe, guiding how fine-tuning should be designed.",{"name":119,"@type":110,"acceptedAnswer":120},"How does SVFT improve stability and extensibility?",{"text":121,"@type":113},"SVFT tunes class-name embeddings via a bi-level optimization that selects a small informative subset using constrained submodular maximization, and enforces orthogonality between prompt embeddings to achieve near-optimal open-vocabulary behavior.","https://schema.org",{"og:url":83,"og:type":124,"og:title":65,"og:site_name":95,"og:description":67},"article",{"robots":126,"canonical":83},"index,follow",{"doc_id":128,"site_id":62},421531,1790671964,{"code":4,"msg":5,"data":131},{"doc_id":128,"user_id":132,"nickname":92,"user_avatar":133,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":8,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":52,"language":139,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":140,"faqs":141,"seo_title":142,"seo_description":67,"update_tm":143,"read_time":144},1374404737137,"https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d","This CVPR paper is the Open Access version, provided by the Computer Vision Foundation.  \nExcept for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.  \nVocabulary Scaling Law : Tuning Open-vocabulary Predictors for  \nTheir Openness  \nZiliang Chen 1 , Yulu Li2 , Liangda Fang2 , Jusheng Zhang3 *, Yongsen Zheng4 , Quanlong Guan2 , Xipeng Chen 1 1Research Institute of Multiple Agents and Embodied Intelligence, Peng Cheng Laboratory, 2Jinan University, 3 Sun Yat-sen University, 4NTU  \nAbstract  \nOpen-vocabulary learning on CLIP provides remarkable generalization on diverse concepts, however, falters under the realistic streaming open-world evaluations for Stability against distractor classes and Extensibility to novel classes. Current fine-tuning methods often fail these tests since they are mainly designed for closed-set conditions, leading to the performance gaps while the target vocabulary progressively scales. We formalize a “vocabulary scaling law” showing that these openness measures can be lower-bounded by performance on the full class-name universe, implying that robust fine-tuning should: (i) account for the entire vocabulary, (ii) tune class-name embeddings rather than context, and (iii) enforce orthogonality between prompt embeddings including training and open-set class names. Guided by our analysis, we propose Submodular-Vocabulary Fine-tuning (SVFT), a bi-level optimization framework that approximates the intractable objective of tuning all class name embedding by greedily selecting a small, informative subset of class names via constrained submodular maximization, thus, allows the employment of efficient greedy algorithm for the near-optimal class-name subset selection to finetune CLIP instead of using all open classes. Across extensive experiments, SVFT consistently improves both stability and extensibility, advancing the openness and practical robustness of CLIP-based vision–language models.  \n1. Introduction  \nVision-Language Models (VLMs) such as CLIP [9, 21] have marked a paradigm shift in visual recognition, leveraging natural language supervision from massive imagetext datasets to enable incredible few-shot and even zeroshot inference [34, 35] . Their hallmark capability, openvocabulary prediction, allows for the classification of images using arbitrary, user-defined category names, breaking free from the constraints of predefined, fixed-class datasets. This flexibility has positioned VLMs as a basic technology  \n*indicate corresponding author;  \nFigure 1 . The comparison between fine-tuning for openness and the other task setups for CLIP from the aspects of whether there are new classes join the training and evaluation, as well as whether the training and evaluation are continually executed.  \nfor a wide array of downstream applications that demand generalization to diverse and unforeseen visual concepts.  \nDespite their successes, a critical gap remains between their typical evaluation and the demands of real-world deployment. In particular, most fine-tuning schemes and evaluation protocols for CLIP operate under the unrealistic assumption that the evaluated images are exactly consistent with the classes to construct the open vocabulary. In other words, existing open-vocabulary predictors are mostly evaluated with close-set classes in the stationary setups. In the pursuit of the openness of CLIP (Fig.1), the concepts of stability and extensibility have emerged as more rigorous metrics for re-evaluating CLIP. Stability measures a model’s robustness to maintain accuracy on known classes when the vocabulary is expanded with unseen ”distractor”concepts, while extensibility measures its zero-shot ability to correctly classify close-set and open-set categories along with the same vocabulary expansion (Fig.2) . As shown by [22], the CLIP family and their fine-tuning methods degrade significantly on these metrics while the vocabulary","cbCaigyd3BBzhUra","https://ap.wps.com/l/cbCaigyd3BBzhUra","pdf",4422023,"English","# Introduction\n## Open-vocabulary CLIP and deployment gap\n## Stability and extensibility metrics\n## Vocabulary scaling law and key takeaways\n# Proposed approach: SVFT\n## Submodular subset selection and bi-level optimization","[{\"question\":\"What problems does the paper address in open-vocabulary CLIP evaluation?\",\"answer\":\"It targets failures in streaming open-world settings, focusing on stability against distractor classes and extensibility to novel classes not seen during fine-tuning.\"},{\"question\":\"What is the “vocabulary scaling law” introduced in the paper?\",\"answer\":\"It states that openness measures are lower-bounded by performance on the scaled vocabulary containing the full open-set class-name universe, guiding how fine-tuning should be designed.\"},{\"question\":\"How does SVFT improve stability and extensibility?\",\"answer\":\"SVFT tunes class-name embeddings via a bi-level optimization that selects a small informative subset using constrained submodular maximization, and enforces orthogonality between prompt embeddings to achieve near-optimal open-vocabulary behavior.\"}]","Vocabulary Scaling Law - Tuning Open-vocabulary Predictors for Their Openness | PDF",1790625078,25]