[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85093-en":3,"doc-seo-85093-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85093,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","VocaDet Sample-Driven Open Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database Retrieval","Open-vocabulary object detection and segmentation targets arbitrary objects beyond predefined categories, yet prior vision-language and reference-based methods often depend on text prompts, limited examples, or costly feature matching, hindering scalability to expanding repositories. VocaDet introduces a sample-driven framework that learns object concepts from user-provided positive and negative collections without retraining. Visual representations are quantized into multi-granularity tokens via DINOv3 and clustering, stored in a vector database for efficient retrieval-based localization and segmentation. A background filtering mechanism reduces redundant matches for fixed cameras, and UA-DETRAC results show effective training-free open-vocabulary performance with continually expandable recognition.","arXiv :2607 .0854 1v 1 [ cs .CV] 9 Jul 2026  \nVocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database Retrieval  \nZhiXin Sun  \nPowerChina Zhongnan Engineering Corporation Limited  \n[sunzxjdi@gmail.com](sunzxjdi@gmail.com)  \nAbstract  \nOpen-vocabulary object detection and segmentation aim to recognize arbitrary objects beyond predefined categories. Although recent vision-language and referencebased approaches have significantly advanced this field, they often rely on text prompts, limited visual examples, or expensive feature matching procedures, making them difficult to scale to large and continuously expanding object repositories.  \nIn this work, we propose VocaDet, a sample-driven open-vocabulary object detection and segmentation framework that learns object concepts directly from user-provided positive and negative sample collections without model retraining. The key idea is to transform continuous visual representations into discrete visual vocabularies and perform efficient retrieval-based recognition through a scalable vector database. Specifically, we employ DINOv3 as the visual feature extractor and apply agglomerative clustering with adaptive clustering sensitivity to generate multi-granularity visual tokens. These visual tokens, together with position-debiased representations and spatial topology information, are stored as expandable object memories in a vector database. During inference, query images are converted into visual tokens and efficiently matched against the stored object memories for object localization and segmentation. Furthermore, a background filtering mechanism is introduced to remove frequently occurring background patterns and reduce redundant retrieval operations in practical fixed-camera scenarios. Experiments on the UA-DETRAC dataset demonstrate that VocaDet achieves effective open-vocabulary detection performance without conventional detector training, while supporting continuously expandable recognition capability as additional positive and negative samples are accumulated. Code is available at [https://github.com/sunzx97/VocaDet](https://github.com/sunzx97/VocaDet).  \n1 Introduction  \nObject detection and segmentation are among the most fundamental and widely applied tasks in computer vision. Owing to its strong real-time performance and mature engineering deployment, the YOLO series[1, 2] has been extensively adopted in industrial applications. However, such methods typically require large-scale data collection and task-specific training for different scenarios, and inevitably suffer from false positives and false negatives.  \nRecently, methods such as SAM3[3], Grounding DINO[4], T-Rex2[5], DINO-X[6], and INSID3[7] have introduced context-aware detection and segmentation via text prompts or visual prompts, enabling open-vocabulary recognition. Nevertheless, these approaches usually require explicit visual prompt inputs or repeated similarity computation between visual prompts and target images, which limits their scalability to large reference contexts and prevents efficient inference.  \nPreprint. Under review.  \nIn addition, recent approaches such as Rex-Thinker [8] and Rex-Omni [9] exploit multimodal large language models to further improve open-domain detection capabilities. Most relevant to our approach is Training-free [10], which introduces a training-free reference image-driven instance segmentation framework. By leveraging frozen DINOv2[11] semantic features and the categoryagnostic segmentation capability of SAM2[12], it constructs a feature memory bank with feature aggregation and semantic-aware matching. Given only a small number of reference images, this method achieves cross-domain generalization for automatic instance-level segmentation.  \nHowever, existing approaches remain limited in scenarios where users provide large-scale positive and negative sample collections and expect the system to automatically learn from t","cbCaiuAqvEZ8FEwH","https://ap.wps.com/l/cbCaiuAqvEZ8FEwH","pdf",2146561,3,1,6,"English","en",105,"# Introduction\n# Method","[{\"question\":\"What problem does VocaDet address in open-vocabulary detection and segmentation?\",\"answer\":\"VocaDet targets open-vocabulary recognition where arbitrary objects must be detected beyond predefined categories. It addresses scalability issues caused by heavy feature matching, limited visual examples, and repeated similarity computations in existing approaches.\"},{\"question\":\"How does VocaDet learn object concepts without retraining?\",\"answer\":\"It builds object memories directly from user-provided positive and negative sample collections. Continuous visual features are transformed into discrete multi-granularity visual tokens using DINOv3 and agglomerative clustering, then stored for retrieval.\"},{\"question\":\"How does VocaDet perform localization and segmentation during inference?\",\"answer\":\"Query images are converted into visual tokens, then matched against the stored object memories in a vector database. Regions whose token nodes or local topological structures exceed a similarity threshold are identified as target objects, enabling segmentation.\"}]",1784201074,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"vocadet-sample-driven-open-vocabulary-object-detection-and-segmentation-via-visual-tokenization-and-vector-database-retrieval","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/vocadet-sample-driven-open-vocabulary-object-detection-and-segmentation-via-visual-tokenization-and-vector-database-retrieval/85093/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does VocaDet address in open-vocabulary detection and segmentation?","Question",{"text":75,"@type":76},"VocaDet targets open-vocabulary recognition where arbitrary objects must be detected beyond predefined categories. It addresses scalability issues caused by heavy feature matching, limited visual examples, and repeated similarity computations in existing approaches.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does VocaDet learn object concepts without retraining?",{"text":80,"@type":76},"It builds object memories directly from user-provided positive and negative sample collections. Continuous visual features are transformed into discrete multi-granularity visual tokens using DINOv3 and agglomerative clustering, then stored for retrieval.",{"name":82,"@type":73,"acceptedAnswer":83},"How does VocaDet perform localization and segmentation during inference?",{"text":84,"@type":76},"Query images are converted into visual tokens, then matched against the stored object memories in a vector database. Regions whose token nodes or local topological structures exceed a similarity threshold are identified as target objects, enabling segmentation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]