[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86101-en":3,"doc-seo-86101-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86101,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Confidence Scores in Open-Vocabulary Detection Are a Biased Mixture of Scale and Semantics","Open-vocabulary object detectors built on image-text foundation models like CLIP generate confidence scores intended to reflect localization certainty. Experiments on COCO using GroundingDINO, OWL-ViT, and YOLO-World, and verification on LVIS with GroundingDINO, show confidence s = cos(v,t) mixes two systematic effects rather than pure detection evidence. Scale bias inflates scores for large objects, while semantic bias suppresses scores for generic prompts. These biases arise from CLIP’s image-level pretraining and cannot be fully removed by thresholding. Temperature scaling improves small-object Recall@10 by 19.6% with measurable precision trade-offs, exposing a fundamental limitation when adapting image-level models to region-level detection.","arXiv :2607 . 10993v1 [ cs .CV] 13 Jul 2026  \nConfidence Scores in Open-Vocabulary Detection Are a Biased Mixture of Scale and Semantics  \nYi Tang Soon 1[0009−0009−8889−8989] and Jun-Wei Hsieh2[0000−0002−5477−4891]  \n1 Institute of Intelligence Systems, National Yang Ming Chiao Tung University,  \nTainan, Taiwan  \n[esthersoon2002.ai14@nycu.edu.tw](esthersoon2002.ai14@nycu.edu.tw)  \n2 College of Artificial Intelligence and Green Energy, National Yang Ming Chiao Tung  \nUniversity, Hsinchu, Taiwan  \n[jwhsieh@nctu.edu.tw](jwhsieh@nctu.edu.tw)  \nAbstract. Foundation models such as CLIP [21] have enabled openvocabulary object detectors that generalise to novel categories via visionlanguage similarity. However, the confidence scores these detectors produce are not reliable localization probability estimates: they conflate visual scale and semantic query specificity with the true detection signal.  \nThrough controlled experiments on COCO across three foundation-modelbased detectors (GroundingDINO, OWL-ViT, YOLO-World), with the scale-bias finding further replicated on LVIS (1,203 categories) using GroundingDINO, we show that s = cos (v, t) is a biased mixture of two effects. Scale bias (αˆ = +0 .064 , r = 0 .579 , p = 1 .29 × 10 −58) systematically inflates scores for large objects. Semantic bias (ˆβ = −0 .705 , p = 5 .23 × 10 −41) suppresses scores for generic queries. Both biases are structurally inevitable from CLIP’s image-level pretraining. Threshold adjustment cannot remove them: oracle per-scale thresholding yields ∆F1 = +0 .001 for small objects versus +0 .102 for large. A parameter-free temperature scaling correction improves small-object Recall@10 by 19.6%(p \u003C 0.01) without retraining. This comes at a modest, measurable cost to pooled-ranking precision, so the bias is partially, not freely, reversible at inference time. These findings reveal a fundamental limitation of adapting image-level foundation models to region-level detection tasks.  \nKeywords: Foundation models · Open-vocabulary detection · CLIP · Confidence calibration · Scale bias · Prompt sensitivity  \n1 Introduction  \nFoundation models such as CLIP [21], ALIGN [8], and SigLIP [29] have transformed computer vision by providing powerful image-text representations that generalise across tasks. Open-vocabulary object detectors build on these models to localise objects described by arbitrary text queries. GroundingDINO [14],  \nThis version of the article has been accepted for publication, after peer review but isnot the Version of Record and does not reflect post-acceptance improvements, or any corrections. The Version of Record will be available online at Springer Link.  \n2 Y.T. Soon and J.W. Hsieh  \nOWL-ViT [17], and YOLO-World [4] all achieve strong zero-shot performance this way. The standard output of these systems is a scalar confidence score s ∈ [0 , 1] per detection, derived from cosine similarity in CLIP space. Practitioners use this score to filter detections by a fixed threshold. The implicit assumption is that s measures localization certainty: a detection with s = 0 .7 should be correct 70% of the time, regardless of object size or query specificity. We show this assumption is fundamentally false.  \nConsider applying such a detector to a scene. At a threshold of 0.3, small objects are systematically discarded. This is not because the detector failed to localise them: it produced correct bounding boxes, but with confidence scores of 0.08, 0.11, and 0.09 . Lowering the threshold to rescue small objects floods the output with large-object false positives. The root cause is not a model failure but a structural property of s = cos (v, t) that conflates two independent signals.  \nWe identify and quantify two systematic biases:  \nScale bias (αˆ = +0 .064 , r = 0 .579 , p = 1 .29 × 10 −58) . Large objects receive systematically higher confidence scores than small objects for identical queries. Averaging CLIP features over fewer pixels yields a noisier, less concen","cbCaivxXs2jCxEct","https://ap.wps.com/l/cbCaivxXs2jCxEct","pdf",689861,4,1,14,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"Why are confidence scores in open-vocabulary object detectors unreliable for localization certainty?\",\"answer\":\"Because the scalar confidence derived from CLIP cosine similarity conflates multiple factors. Scale and semantic prompt specificity bias the score away from true localization evidence.\"},{\"question\":\"What are scale bias and semantic bias in this paper?\",\"answer\":\"Scale bias systematically increases confidence for larger objects, while semantic bias decreases confidence for generic queries compared with specific prompts. Both are tied to CLIP’s pretraining behavior.\"},{\"question\":\"Can adjusting detection thresholds remove the confidence bias?\",\"answer\":\"No. The paper reports that oracle per-scale thresholding yields only negligible improvement for small objects and substantial remaining issues overall, so bias is structurally persistent.\"}]",1784208513,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"confidence-scores-in-open-vocabulary-detection-are-a-biased-mixture-of-scale-and-semantics","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/confidence-scores-in-open-vocabulary-detection-are-a-biased-mixture-of-scale-and-semantics/86101/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are confidence scores in open-vocabulary object detectors unreliable for localization certainty?","Question",{"text":75,"@type":76},"Because the scalar confidence derived from CLIP cosine similarity conflates multiple factors. Scale and semantic prompt specificity bias the score away from true localization evidence.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are scale bias and semantic bias in this paper?",{"text":80,"@type":76},"Scale bias systematically increases confidence for larger objects, while semantic bias decreases confidence for generic queries compared with specific prompts. Both are tied to CLIP’s pretraining behavior.",{"name":82,"@type":73,"acceptedAnswer":83},"Can adjusting detection thresholds remove the confidence bias?",{"text":84,"@type":76},"No. The paper reports that oracle per-scale thresholding yields only negligible improvement for small objects and substantial remaining issues overall, so bias is structurally persistent.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]