[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84484-en":3,"doc-seo-84484-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84484,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","Retrieval with Multiple Query Vectors through Anomalous Pattern Detection","Classical vector retrieval typically uses a single query embedding to find nearest neighbors in a vector database, but complex reasoning tasks often require multiple query vectors. This work introduces a multi-query retrieval method based on anomalous pattern detection. Given a query set Q (|Q|≥1), the approach identifies an anomalous subset of vector dimensions, scans the database for vectors that match the same joint anomaly, and returns them as results. Experiments on two image datasets, one text dataset, and one tabular dataset show that larger query sets improve retrieval, with the biggest gains moving from 1 to 8 queries.","Retrieval with Multiple Query Vectors through Anomalous Pattern Detection  \nAllassan Tchangmena A Nken 1 * , Baimam Boukar Jean Jacques2, * , Miriam Rateike3, 4 , Celia Cintas3 , Skyler Speakman3  \n1IBM,  \n2University of Galway, 3 Carnegie Mellon University Africa, 4University of T¨ubingen,  \n[a.tchangmenaanken1@universityofgalway.ie](a.tchangmenaanken1@universityofgalway.ie), [bbaimamb@andrew.cmu.edu](bbaimamb@andrew.cmu.edu),  \n[miriam.rateike@ibm.com](miriam.rateike@ibm.com), [celia.cintas@ibm.com](celia.cintas@ibm.com), [skyler@ke.ibm.com](skyler@ke.ibm.com)  \narXiv :2605 .0 1965v2 [ cs .LG] 13 Jul 2026  \nAbstract  \nA classical vector retrieval problem typically considers a single query embedding vector as input and retrieves the most similar embedding vectors from a vector database. However, complex reasoning and retrieval tasks frequently require multiple query vectors, rather than a single one. In this work, we propose a retrieval method that considers multiple query vectors simultaneously and retrieves the most relevant vectors from the database using concepts from anomalous pattern detection. Specifically, our approach leverages a set of query vectors Q (with |Q| ≥ 1), and identifies the subset of vector dimensions within Q that standout (anomalous) from the rest of dimensions. Next, we scan the vector database to retrieve the set of vectors that are also anomalous across the previously identified vector dimensions and return them as our retrieved set of vectors. We validate our approach on two image datasets, a text dataset, and a tabular dataset. Overall, we observe that, across most datasets, larger query sets lead to improved retrieval performance. The improvement is most pronounced when increasing the query sets from 1 to 8, while the gains become smaller beyond that.  \n1 Introduction  \nInformation retrieval (IR) underpins a wide range of applications such as question answering (Karpukhin et al. 2020), image search (Oquab et al. 2024), and more recently has become central in retrieval-augmented generation (RAG) for large language models (LLMs), including retrieval asa reasoning action (Yao et al. 2022), retrieval for selfcritique/self-verification (Asai et al. 2024), and retrieval-inthe-loop during LLM training (Tang et al. 2024). In addition, an increasing number of companies are adopting IR systems as they migrate from traditional relational databases to vector databases 1 , owing to their ability to efficiently handle low-latency queries.2  \nMost retrieval systems operate in a single-query setting. For example, assume a news retrieval scenario, where a user provides a query sentence about an event, which is encoded  \n*This work was done during an internship at IBM Research. Copyright © 2026, Association for the Advancement of Artificial Intelligence ([www.aaai.org](www.aaai.org)). All rights reserved.  \n1Vector databases store data points as high-dimensional numerical vectors.  \n2[https://www.marketresearch.com/Global-Industry-Analysts](https://www.marketresearch.com/Global-Industry-Analysts)v1039/Vector-Databases-41409421/, last accessed on 05 .11.2025.  \ninto a single vector representation, such as the final token embedding from a large language model (LLM) . The retrieval system then returns articles from a database, which are relevant to this query vector. In practice, however, retrieval tasks often involve multiple query vectors. For instance, a user may instead provide a paragraph summarizing the event, where each sentence captures a different aspect (such as the location, participants, or outcomes) and is encoded as a separate query vector. Existing retrieval methods typically address this challenge by either aggregating multiple query vectors into a single embedding via simple heuristics (Li and Lu 2016; Wang et al. 2022), or by treating each query independently (Xu et al. 2016) . Both lines of work then apply standard single-query retrieval methods such as cosine similarity (Venkatesh Sharma et al. 20","cbCaiqD8GzG8B2F1","https://ap.wps.com/l/cbCaiqD8GzG8B2F1","pdf",776968,1,7,"English","en",105,"# Abstract\n# Introduction\n# Retrieval through Anomalous Pattern","[{\"question\":\"How does the proposed method handle multiple query vectors compared with classical retrieval?\",\"answer\":\"Instead of combining or independently querying multiple vectors using distance metrics, the method processes a set of query vectors jointly to find a shared anomalous pattern and then retrieves database vectors exhibiting the same pattern.\"},{\"question\":\"What is the core idea behind defining relevance in this work?\",\"answer\":\"Relevance is determined by an anomaly scoring function derived from anomalous dimensions shared across the query set, rather than by proximity under cosine similarity, inner product, or Euclidean distance.\"},{\"question\":\"How does increasing the number of query vectors affect retrieval performance?\",\"answer\":\"Across most datasets, larger query sets lead to improved retrieval performance. The improvement is most significant when increasing the query set size from 1 to 8, with diminishing gains beyond that.\"}]",1784195966,18,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"retrieval-with-multiple-query-vectors-through-anomalous-pattern-detection","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/retrieval-with-multiple-query-vectors-through-anomalous-pattern-detection/84484/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the proposed method handle multiple query vectors compared with classical retrieval?","Question",{"text":75,"@type":76},"Instead of combining or independently querying multiple vectors using distance metrics, the method processes a set of query vectors jointly to find a shared anomalous pattern and then retrieves database vectors exhibiting the same pattern.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the core idea behind defining relevance in this work?",{"text":80,"@type":76},"Relevance is determined by an anomaly scoring function derived from anomalous dimensions shared across the query set, rather than by proximity under cosine similarity, inner product, or Euclidean distance.",{"name":82,"@type":73,"acceptedAnswer":83},"How does increasing the number of query vectors affect retrieval performance?",{"text":84,"@type":76},"Across most datasets, larger query sets lead to improved retrieval performance. The improvement is most significant when increasing the query set size from 1 to 8, with diminishing gains beyond that.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]