[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84570-en":3,"doc-seo-84570-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84570,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","Personalized Object Identification and Localization via In-Context Inference with Vision-Language Models","Personalized object localization (POL) pinpoints an object instance in a query image using a few annotated reference images and a target label. The pioneering IPLoc approach performs in-context inference with vision-language models (VLMs) but assumes every query contains the target object, restricting real-world use with many irrelevant images. This work introduces personalized object identification and localization (POIL), which localizes the target instance when present and rejects queries that do not contain the reference instance. It also presents IPLoc-ID and datasets.","arXiv :2607 .00357v 1 [ cs .CV] 1 Jul 2026  \nPersonalized Object Identification and Localization via In-Context Inference with Vision-Language Models  \nKensuke Nakamuraa , Byung-Woo Honga,∗  \na Artificial Intelligence Department, Chung-Ang University, Seoul, 06974, Korea  \nAbstract  \nPersonalized object localization (POL) localizes an object instance in a query image based on a few reference images with bounding-box annotations and a target object label. The pioneering method, IPLoc, solves this task through in-context inference with vision-language models (VLMs) . However, it assumes that the query image always contains the target object. This assumption severely limits its applicability to real-world scenarios with many irrelevant images. To address this issue, we formulate a new task, personalized object identification and localization (POIL), by positioning POL within the broader few-shot object detection framework. POIL aims to localize the target object instance while rejecting query images that do not contain the reference object instance. We also present POIL datasets constructed from public sources. We further propose an in-context algorithm named IPLoc-ID for solving POIL with VLMs. IPLocID first predicts a candidate bounding box and then determines whether it corresponds to the reference object instance. We introduce a self-posed query to connect these two steps within a single autoregressive generation framework. Through ablation studies and comprehensive experiments, we show that IPLoc-ID substantially suppresses false-positive detections on negative query images while maintaining localization performance comparable to IPLoc. Overall, IPLoc-ID effectively addresses the practical instance-level POIL task, which cannot be sufficiently solved by conventional object detection, few-shot object detection, orthe localization-only IPLoc method.  \nKeywords: object detection, object identification, bounding-box localization, vision-language models, in-context learning  \n∗ Corresponding author  \nEmail addresses: [kensuke@image.cau.ac.kr](kensuke@image.cau.ac.kr) (Kensuke Nakamura), [hong@cau.ac.kr](hong@cau.ac.kr)[ ](hong@cau.ac.kr)(Byung-Woo Hong)  \n1. Introduction  \nObject detection (OD) is a fundamental visual recognition task that aims to find objects in an image and estimate their locations as bounding boxes. Recent advances in open-vocabulary object detection and few-shot object detection (FSOD) have made it possible to detect objects specified not only by predefined categories but also by text labels or a small number of support examples [1– 5] . However, most of these methods are essentially designed for category-level detection and do not aim to identify a specific object instance indicated by reference data. For example, even when a reference image specifies a particular cat, conventional OD or FSOD methods may regard detecting another cat from the same category as a successful result. In contrast, reference-conditioned instance-level localization aims to detect a specific object instance indicated by reference data in a query image. Such a capability is expected to be useful for future applications such as user-specified image retrieval, video grounding, object re-identification, and personalized object tracking.  \nIn this line of research, IPLoc (in-context personalized object localization) [6] is pioneering work on reference-conditioned instance-level localization. It exploits the contextual understanding ability of transformer-based vision-language models (VLMs) to localize the corresponding object region in a query image based on reference data. IPLoc takes a small number of images with bounding-box (BBOX) annotations and the target label as reference data, and generates the BBOX coordinates for the query image through next-token prediction. This formulation enables reference-conditioned inference with VLMs without fine-tuning to the reference data. However, IPLoc assumes that the target object is present in t","cbCairzBOlHhUaq7","https://ap.wps.com/l/cbCairzBOlHhUaq7","pdf",46437536,1,30,"English","en",105,"# Introduction\n## Object detection limitations for instance-level identification\n## Personalized object identification and localization (POIL)\n## IPLoc-ID approach and evaluation setup","[{\"question\":\"What problem does the paper address in personalized object localization?\",\"answer\":\"IPLoc assumes the target object is always present in the query image, so it still generates bounding boxes even for negative queries where the object is absent, causing many false positives in practical settings.\"},{\"question\":\"What is personalized object identification and localization (POIL)?\",\"answer\":\"POIL extends the POL setting into a few-shot detection framework where the model outputs a bounding box only when the same object instance from the reference data exists in the query image, and rejects the query otherwise.\"},{\"question\":\"How does IPLoc-ID solve POIL with vision-language models?\",\"answer\":\"IPLoc-ID first predicts a candidate bounding box and then checks whether it matches the reference object instance, using a self-posed query to connect these steps within a single autoregressive generation process.\"}]",1784196865,76,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"personalized-object-identification-and-localization-via-in-context-inference-with-vision-language-models","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/personalized-object-identification-and-localization-via-in-context-inference-with-vision-language-models/84570/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in personalized object localization?","Question",{"text":75,"@type":76},"IPLoc assumes the target object is always present in the query image, so it still generates bounding boxes even for negative queries where the object is absent, causing many false positives in practical settings.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is personalized object identification and localization (POIL)?",{"text":80,"@type":76},"POIL extends the POL setting into a few-shot detection framework where the model outputs a bounding box only when the same object instance from the reference data exists in the query image, and rejects the query otherwise.",{"name":82,"@type":73,"acceptedAnswer":83},"How does IPLoc-ID solve POIL with vision-language models?",{"text":84,"@type":76},"IPLoc-ID first predicts a candidate bounding box and then checks whether it matches the reference object instance, using a self-posed query to connect these steps within a single autoregressive generation process.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":21,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]