[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82144-en":3,"doc-seo-82144-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82144,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","C-GAP Class-Aware and Online Prompting Improves Vision-Language Models on Imbalanced Classes","Safety-critical perception systems must reliably detect rare object classes within small label spaces, where long-tailed methods built for dense, large-category supervision provide limited help. Open-vocabulary detectors enable natural-language queries at inference time, making prompt quality a direct control knob. C-GAP tests whether iteratively refining prompts for frozen detectors can boost minority-class detection without retraining or extra annotations. It introduces a composite caption baseline and an LLM-driven refinement loop, achieving large AP@0.5 gains over baselines.","C-GAP: Class-Aware and Online Prompting Improves Vision–Language  \nModels on Imbalanced Classes  \nFrancis Fernandez San Diego State University  \n[fafernandez@sdsu.edu](fafernandez@sdsu.edu)  \nArash Jahangiri San Diego State University  \n[ajahangiri@sdsu.edu](ajahangiri@sdsu.edu)  \nSalimeh Sekeh San Diego State University  \n[ssekeh@sdsu.edu](ssekeh@sdsu.edu)  \narXiv :2607 .09008v1 [ cs .CV] 10 Jul 2026  \nAbstract  \nSafety-critical perception systems must reliably detect rare object classes within small label spaces, a setting that long-tailed detection methods, designed for hundreds of classes with dense annotation, fundamentally do not address. Open-vocabulary detectors offer a promising alternative, as they use natural language queries at inference time, making prompt quality a first-class lever for detection performance. We exploit this property to address class imbalance: rather than retraining models or collecting additional annotations, we ask whether iteratively refining the language prompts, fed to frozen detectors, can improve minorityclass detection. We introduce C-GAP (Caption-Guided Augmentation and Prompting), a detector-agnostic, annotation-free framework that operates in two phases. First, we establish a composite caption baseline combining per-image scene descriptions with class-quantity context, which we show outperforms scene-descriptiononly or class-quantity-only prompts across multiple open-vocabulary architectures and benchmarks. Second, an LLM iteratively refines each image’s caption individually, with trials triaged into accept, tentative, or regenerate buckets based on minority-class AP@0.5 against a dynamic threshold derived from the composite baseline. Refinement terminates early once sufficient AP@0.5 gain is achieved. No detector weights are updated at any stage. Our experiments shows that C-GAP improves minority-class average precision up to 53% over the baselines. On COCO, C-GAP improves minority-class AP@0.5 by ∼81% relative over the composite baseline (17 .69 → 32.09). Experiments confirm that composite captions provide the critical foundation for effective refinement: using scene-description-only or class-quantity-only prompts as the refinement starting point yields diminishing returns, supporting both stages of C-GAP as necessary contributions.  \nFigure 1 . Chula Vista intersection: cars dominate (∼80%) while cyclists are severely underrepresented (∼8%), motivating minority-class prompt refinement.  \n1. Introduction  \nSafety-critical perception systems must reliably detect rare object classes within small label space, a setting where long-tailed detection methods, designed for hundreds of categories with dense annotation, offer limited guidance [18, 33] . In urban traffic monitoring, minority classes such as cyclists and pedestrians account for a small fraction of instances yet are the most consequential for collision avoidance and regulatory compliance. Standard resampling and loss-reweighting remedies require labeled minority instances that may simply not exist in sufficient quantity, and collecting additional annotations is expensive.  \nBecause open-vocabulary detectors (OVDs) [51] accept free-form text queries at inference time rather thana fixed classification head, prompt quality becomes a direct lever for detection performance, requiring no annotation, no architectural change, and no retraining. We  \nFigure 2 . Overview of C-GAP. Phase I (top): per-image Scene Description (tSDi) and Class Quantity (tCQi) captions are generated offline and concatenated into a Composite Caption ti,0 = concat(tSDi, tCQi), forming the initial caption set T0 . Phase II (bottom): a VLM refines T0 over K trials. Each candidate set Tk is evaluated by the frozen detector fθ , yielding minority-class AP50 (cm ;Tk) . Trials are triaged into B1 (regenerate, AP50 \u003C τlow ), B2 (tentative, τlow ≤ AP50 ≤ τhigh ), or B3 (keep, AP50 > τhigh ) . The output is Tk ∗ where k ∗ = arg max0≤k≤K AP50 (cm ;Tk) . Detector weights θ","cbCaid2Y2NMlIHsj","https://ap.wps.com/l/cbCaid2Y2NMlIHsj","pdf",6225165,1,18,"English","en",105,"# Abstract\n# Introduction\n# C-GAP Framework Overview\n# Application: Smart Intersection Monitoring\n# Contributions","[{\"question\":\"What is the role of the composite caption baseline in C-GAP?\",\"answer\":\"The composite caption baseline combines per-image scene descriptions with class-quantity context. Experiments show it is the critical foundation for effective refinement, while starting from scene-description-only or class-quantity-only prompts yields diminishing returns.\"}]",1784178425,45,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"c-gap-class-aware-and-online-prompting-improves-vision-language-models-on-imbalanced-classes","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/c-gap-class-aware-and-online-prompting-improves-vision-language-models-on-imbalanced-classes/82144/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What is the role of the composite caption baseline in C-GAP?","Question",{"text":75,"@type":76},"The composite caption baseline combines per-image scene descriptions with class-quantity context. Experiments show it is the critical foundation for effective refinement, while starting from scene-description-only or class-quantity-only prompts yields diminishing returns.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]