[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82986-en":3,"doc-seo-82986-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82986,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Retrieving a Set Not Independent Passages Set-Level Compatibility Learning for Efficient Set Exploration","Multi-hop question answering requires selecting evidence passages that are jointly useful for a query, yet many retrievers score passages independently or rely on locally supervised sequential decisions, failing when usefulness depends on compatibility across passages. The work frames multi-hop retrieval as query–set compatibility scoring and introduces a set-level retrieval framework. Training ranks complete, compatible evidence sets above incomplete noisy alternatives, improving robustness to variable-length and partially noisy contexts. The approach instantiates ParaSet and SetCE, and experiments show consistent gains on multi-hop QA benchmarks and complementary benefits over document-level retrieval.","Retrieving a Set, Not Independent Passages: Set-Level Compatibility Learning for Efficient Set Exploration  \nMooho Song, Jay-Yoon Lee *  \nSeoul National University {anmh9161, [lee.jayyoon}@snu.ac.kr](lee.jayyoon}@snu.ac.kr)  \narXiv :2607 .057 12v 1 [ cs .IR] 7 Jul 2026  \nAbstract  \nMulti-hop question answering and retrievalaugmented reasoning require selecting evidence passages that are jointly useful for answering a query. However, most retrievers still score passages independently or make locally supervised sequential decisions, which can fail when evidence usefulness depends on compatibility among passages. LLM-based set selection can model such interactions, but its computational cost limits practical use. We address this gap by formulating multi-hop retrieval as query–set compatibility scoring and propose a set-level retrieval framework. Our training objective teaches retrievers to rank complete and compatible evidence sets above incomplete, noisy alternatives, making set scoring more robust to variable-length and partially noisy contexts. We instantiate the framework with two complementary set scorers: ParaSet, a lightweight late-interaction scorer that applies self-attention over precomputed bi-encoder embeddings for fast candidate-set exploration, and SetCE, a cross-encoder-based reranker trained with the same set-level objective. Experiments on various multi-hop QA benchmarks show that set-level compatibility learning improves retrieval performance and downstream QA task performance. We further show that the proposed set-level retrievers not only outperform document-level retrievers, but also exhibit complementary retrieval characteristics: combining their outputs yields stronger performance than simply retrieving more passages from a single document-level retriever.  \n1 Introduction  \nDocument retrieval has traditionally focused on selecting the single most relevant passage for a given query. However, many reasoning-intensive tasks such as multi-hop QA require retrieving a set of evidence passages, where the usefulness of each passage is inherently dependent on the others. In  \n* Correspondence author  \nsuch settings, retrieving individual passages in independent manner is insufficient: a passage that is useful in isolation may become uninformative or even misleading when combined with an incompatible context. This interdependence makes multi-hop retrieval fundamentally a set-level problem, rendering standard single-passage retrieval approaches inadequate.  \nRecent work such as Lee et al. (2025) explicitly highlights this limitation and proposes retrieving sets of passages rather than individual passages. However, their approach relies on LLM-based set selection, resulting in substantial computational and memory overhead that limits its practicality in large-scale or latency-sensitive settings. Similarly, although Chen et al. (2025) does not explicitly discuss the limitations of single-passage retrieval, it also aims to retrieve multiple objects jointly by leveraging structured relationships across heterogeneous sources. However, its retrieval pipeline depends on both an LLM and an MIP solver, incurring high computational cost. These approaches suggest the importance of set-level reasoning, but also highlight a central challenge: making set-level retrieval computationally practical. A useful setlevel retrieval model should be able to account for global compatibility among passages while remaining efficient enough to search over many candidate passage sets.  \nNon-LLM sequential retrieval methods provide a more computationally efficient alternative. For example, although Zhang et al. (2024) does not explicitly frame its motivation around set retrieval, its cross-encoder scorer can be directly used for multi-hop QA by iteratively constructing a retrieval chain using beam search. This procedure enables conditioning on previously retrieved evidence without relying on LLM-based passage selection. However, the training sign","cbCaiq92LgefiLmH","https://ap.wps.com/l/cbCaiq92LgefiLmH","pdf",643160,1,19,"English","en",105,"# Introduction\n## Evidence retrieval as a set-level problem\n## Limits of LLM-based and sequential retrievers\n## Proposed query–set compatibility framework\n## Set scorers: ParaSet and SetCE\n## Main empirical findings","[{\"question\":\"Why is multi-hop retrieval treated as a set-level problem rather than independent passage selection?\",\"answer\":\"Each passage’s usefulness can depend on the other passages, so independently retrieved evidence may become uninformative or misleading when combined. Multi-hop QA therefore needs evidence sets whose members are jointly compatible with the query context.\"},{\"question\":\"What problem does the proposed method address compared with prior retrievers?\",\"answer\":\"Prior retrievers often score passages independently or use locally supervised sequential decisions with limited training signals. This can make them brittle under variable-length or noisy retrieval chains where global query–set interactions matter.\"},{\"question\":\"How does the framework score evidence sets and what models implement it?\",\"answer\":\"It formulates multi-hop retrieval as query–set compatibility scoring and trains retrievers to rank complete, compatible evidence sets over incomplete noisy alternatives. It is instantiated with ParaSet (lightweight late-interaction scorer) and SetCE (cross-encoder reranker) trained with the same set-level objective.\"}]",1784184473,48,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"retrieving-a-set-not-independent-passages-set-level-compatibility-learning-for-efficient-set-exploration","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/retrieving-a-set-not-independent-passages-set-level-compatibility-learning-for-efficient-set-exploration/82986/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is multi-hop retrieval treated as a set-level problem rather than independent passage selection?","Question",{"text":75,"@type":76},"Each passage’s usefulness can depend on the other passages, so independently retrieved evidence may become uninformative or misleading when combined. Multi-hop QA therefore needs evidence sets whose members are jointly compatible with the query context.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What problem does the proposed method address compared with prior retrievers?",{"text":80,"@type":76},"Prior retrievers often score passages independently or use locally supervised sequential decisions with limited training signals. This can make them brittle under variable-length or noisy retrieval chains where global query–set interactions matter.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the framework score evidence sets and what models implement it?",{"text":84,"@type":76},"It formulates multi-hop retrieval as query–set compatibility scoring and trains retrievers to rank complete, compatible evidence sets over incomplete noisy alternatives. It is instantiated with ParaSet (lightweight late-interaction scorer) and SetCE (cross-encoder reranker) trained with the same set-level objective.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]