[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83030-en":3,"doc-seo-83030-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83030,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Learn to Pool: Lightweight Fine-Tuning for Flexible Multi-Vector Compression","Late interaction models offer strong generalization but require many token vectors per document, creating storage and memory costs for retrieval systems. Token pooling at inference reduces vector count with limited accuracy loss, and pooling-aware large-scale training improves results at high compression. This work proposes lightweight pooling-aware fine-tuning for ColBERT, showing gains even with minimal training using k-means, transfer across pooling methods and datasets, and multi-factor training that yields one model across compression levels. The best model surpasses the unpooled baseline on BEIR SciFact across pool factors 1–6, enabling ~83% vector compression without sacrificing retrieval accuracy.","Learn to Pool: Lightweight Fine-Tuning for Flexible Multi-Vector Compression  \nStefan Josef1, * 1 Independent Researcher  \nAbstract  \nLate interaction models have shown strong generalization capabilities, often outperforming much larger dense embedding models. One challenge to their widespread deployment is the large number of token vectors they produce per document and the associated storage and memory costs. Pooling tokens at inference time has shown great promise to reduce the vector count with limited effects on retrieval accuracy. Large-scale pooling-aware training has demonstrated even more impressive results at high compression rates. We propose lightweight fine-tuning as a practical alternative and find that even minimal pooling-aware training with k-means yields broad gains over inference-only pooling, shows evidence of transfer across pooling methods and datasets, and —with multi-factor training —produces a single model effective across different compression levels. Our strongest model outperforms the unpooled baseline on BEIR SciFact across pool factors 1–6, implying a vector compression rate of 83% at no cost to retrieval accuracy.  \nKeywords  \nLate interaction, Multi-vector, ColBERT, Token Pooling, Vector Compression, Fine-tuning  \n1. Introduction  \nLate interaction models such as ColBERT [1] represent documents as collections of token embeddings, allowing fine-grained token-level interactions that show strong generalization across domains and document lengths. A key challenge for their widespread adoption is the large number of vectors that need to be stored per document compared to single vector dense embedding models.  \nRecent work has focused on reducing the number of vectors that need to be stored while retaining most of the accuracy of multi-vector models. Inference pooling as introduced by Clavié et al. [2] applies hierarchical pooling to ColBERT models without any training and shows that the number of vectors can be reduced by 2–3× without significantly lowering retrieval accuracy. Veneroso et al. [3] has focused on improving the efficiency-accuracy curve by training pooling-aware late interaction models with k-means clustering. Models trained this way achieve impressive compression rates without performance degradation, yet they rely on large-scale contrastive training which may not be accessible or practical for many practitioners that care mostly about their in-domain datasets.  \nThis paper investigates whether lightweight pooling-aware fine-tuning of existing ColBERT models can achieve gains over pure inference pooling techniques. We first evaluate a small and modern ColBERT model on selected datasets using training-free inference pooling techniques — sequential (span), hierarchical or k-means clustering — across various pool factors. We then fine-tune the model using different pool factors and methods. We evaluate all fine-tuned models against the untrained baseline model, assess cross-method spill-over and cross-dataset effects to capture potential degradation on other datasets.  \nWe find that:  \n1. Hierarchical pooling is the strongest inference pooling technique overall, consistent with Clavié et al. [2] .  \nThe 1st Late Interaction Workshop (LIR) @ ECIR 2026, April 02, 2026, Delft, NL  \n* Corresponding author.  \n$ [stefan.m.josef@gmail.com](stefan.m.josef@gmail.com) (S. Josef)  \n􀀚 0009-0005-4387-6912 (S. Josef)  \n © 2026 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0) .  \n2. Lightweight fine-tuning is an effective technique for making models pooling-aware. K-means is the strongest training method overall, producing the most consistent gains over untrained baselines. However, the strength of the gains depends on pool factor and dataset.  \n3. Fine-tuning without pooling can severely degrade pooled performance, demonstrating that the observed gains are specifically attributable to including pooling in the training loop.  ","cbCaicGkxMFkE3wF","https://ap.wps.com/l/cbCaicGkxMFkE3wF","pdf",870612,2,1,14,"English","en",105,"# Introduction\n# Related Work\n# Pooling-Aware Fine-Tuning Experiments\n# Results and Findings","[{\"question\":\"What problem do pooling methods address in late interaction models like ColBERT?\",\"answer\":\"They address the large number of token vectors produced per document, which increases storage and memory costs in retrieval systems.\"},{\"question\":\"How does the proposed lightweight fine-tuning improve over inference-only pooling?\",\"answer\":\"It enables pooling-aware behavior during training; even minimal pooling-aware training with k-means yields broad gains over inference-only pooling.\"},{\"question\":\"What is the performance impact and compression level of the strongest model?\",\"answer\":\"On BEIR SciFact, the strongest model outperforms the unpooled baseline across pool factors 1–6, implying about 83% vector compression at no cost to retrieval accuracy.\"}]",1784184769,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"learn-to-pool-lightweight-fine-tuning-for-flexible-multi-vector-compression","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/learn-to-pool-lightweight-fine-tuning-for-flexible-multi-vector-compression/83030/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem do pooling methods address in late interaction models like ColBERT?","Question",{"text":75,"@type":76},"They address the large number of token vectors produced per document, which increases storage and memory costs in retrieval systems.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed lightweight fine-tuning improve over inference-only pooling?",{"text":80,"@type":76},"It enables pooling-aware behavior during training; even minimal pooling-aware training with k-means yields broad gains over inference-only pooling.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the performance impact and compression level of the strongest model?",{"text":84,"@type":76},"On BEIR SciFact, the strongest model outperforms the unpooled baseline across pool factors 1–6, implying about 83% vector compression at no cost to retrieval accuracy.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]