[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82834-en":3,"doc-seo-82834-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82834,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","Do All Visual Tokens Matter Equally Object-Evidence Preserving Token Merging for Vision-Language Retrieval","Multi-vector vision-language retrieval preserves fine-grained visual evidence via maximum-similarity late interaction, yet dense image-side tokens cause costly storage and scoring. Existing token compression can discard or collapse object- and region-level evidence needed by future query tokens. SaMer proposes object-aware token merging that compresses image-side post-projector tokens into K representative centroids while keeping the late-interaction interface unchanged. Using object annotations only as a training merge prior, SaMer avoids detector needs at inference and adapts only the shared projection layer; with K=64 it removes over 93% of tokens, reduces ColPali storage by 16.09×, and improves R@1 while enhancing phrase-level grounding.","Do All Visual Tokens Matter Equally?  \nObject-Evidence Preserving Token Merging for Vision-Language Retrieval  \nSuhyeong Park 1,3 , Junha Jung2,3 , Jungwoo Park2,3 , Jaewoo Kang2,3∗  \n1The Catholic University of Korea  \n2 Korea University  \n3AIGEN Sciences Inc.  \n[pshpulip22@catholic.ac.kr](pshpulip22@catholic.ac.kr)  \n{goodjungjun,jungwoo-park,[kangj}@korea.ac.kr](kangj}@korea.ac.kr)  \narXiv :2607 .04605v2 [ cs .IR] 13 Jul 2026  \nAbstract  \nMulti-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-side tokens make storage and scoring expensive. Existing token compression methods reduce this cost, yet they can remove or collapse object-and region-level evidence that future query tokens may need to select. We propose SaMer, an object-aware token merging framework that compresses image-side post-projector tokens into K representative centroids while preserving the original late-interaction interface. SaMer uses object annotations only during training as a merge prior to discourage cross-instance mixing, requires no ground-truth bounding boxes or detectors at inference time, and adapts only the shared projection layer with frozen vision and language backbones. With K = 64, SaMer removes more than 93% of image-side tokens and reduces ColPali storage by  \n16.09 ×, while improving R@1 on Flickr30K and MSCOCO. These gains arise because object-aware merging preserves query-selectable object evidence that pruning or feature-only pooling can remove or collapse. SaMer also outperforms compression baselines and shows stronger phrase-level grounding, suggesting that efficient multi-vector retrieval depends not only on reducing token count, but on preserving the evidence future query tokens need to select.  \nCode—[https://github.com/dmis-lab/SaMer](https://github.com/dmis-lab/SaMer)  \n1 Introduction  \nEfficient multi-vector vision-language retrieval requires preserving the object-, attribute-, and relation-level evidence that query tokens can select under late interaction. Multivector retrieval has been widely adopted in text retrieval since ColBERT (Khattab and Zaharia 2020; Santhanam et al. 2022b), which performs maximum-similarity (MaxSim) matching between contextualized query and document token embeddings. Recent vision-language retrievers such as ColPali (Faysse et al. 2024) extend this paradigm to visual inputs by storing image-side patch embeddings and comparing them with query tokens through MaxSim. Similarly, ColQwen2 (Faysse et al. 2024) builds on Qwen2-VL (Wanget al. 2024), whose visual tokenization supports flexible image representation. These models use MaxSim to compare each query token with all image-side tokens and select the  \n∗Corresponding author.  \nmost similar visual token as evidence, making retrieval sensitive to objects, attributes, and relations rather than only global image-text similarity (Li et al. 2019; Diao et al. 2021) .  \nHowever, preserving token-level image representations makes retrieval expensive in both storage and scoring. A multi-vector retriever stores hundreds to over a thousand image-side token embeddings per image, and retrieval requires MaxSim comparisons between query tokens and all stored image tokens. As the index grows, both memory footprint and scoring latency increase rapidly, limiting largescale image search and retrieval-augmented visual question answering. The issue is especially pronounced in patch-based visual representations, where each image produces many local tokens that must be stored and compared during MaxSim scoring, making patch-level embeddings a major bottleneck for multi-vector VLM retrievers (Dosovitskiy et al. 2021; Tang et al. 2022; Ma et al. 2025a) .  \nOne strategy is to compress visual tokens through pruning or merging (Ryoo et al. 2021; Rao et al. 2021; Lianget al. 2022; Marin et al. 2023; Bolya et al. 2023) . Recent multimodal acceleration methods reduce visual tokens for large vision-lan","cbCaikrYarMsLVei","https://ap.wps.com/l/cbCaikrYarMsLVei","pdf",19135976,5,1,14,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"Why can’t generic token compression be safely applied to multi-vector vision-language retrieval?\",\"answer\":\"Late interaction depends on object-, attribute-, and relation-level evidence selected by individual query tokens. If compression prunes or merges phraserelevant tokens, future MaxSim matching may lose the evidence a query requires.\"},{\"question\":\"What is SaMer and how does it preserve object-evidence for retrieval?\",\"answer\":\"SaMer performs object-aware token merging on image-side post-projector tokens, forming K representative centroids with a feature-spatial soft assignment. A merge prior guided by object annotations during training discourages cross-instance mixing so the merged tokens remain selectable for later MaxSim queries.\"},{\"question\":\"What does SaMer require during inference, and what is trained?\",\"answer\":\"At inference, SaMer needs no ground-truth bounding boxes and no object detector. Training freezes the vision and language backbones and adapts only the shared projection layer to keep compatibility with the late-interaction MaxSim scoring interface.\"}]",1784183286,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"do-all-visual-tokens-matter-equally-object-evidence-preserving-token-merging-for-vision-language-retrieval","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/do-all-visual-tokens-matter-equally-object-evidence-preserving-token-merging-for-vision-language-retrieval/82834/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why can’t generic token compression be safely applied to multi-vector vision-language retrieval?","Question",{"text":76,"@type":77},"Late interaction depends on object-, attribute-, and relation-level evidence selected by individual query tokens. If compression prunes or merges phraserelevant tokens, future MaxSim matching may lose the evidence a query requires.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is SaMer and how does it preserve object-evidence for retrieval?",{"text":81,"@type":77},"SaMer performs object-aware token merging on image-side post-projector tokens, forming K representative centroids with a feature-spatial soft assignment. A merge prior guided by object annotations during training discourages cross-instance mixing so the merged tokens remain selectable for later MaxSim queries.",{"name":83,"@type":74,"acceptedAnswer":84},"What does SaMer require during inference, and what is trained?",{"text":85,"@type":77},"At inference, SaMer needs no ground-truth bounding boxes and no object detector. Training freezes the vision and language backbones and adapts only the shared projection layer to keep compatibility with the late-interaction MaxSim scoring interface.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]