[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84677-en":3,"doc-seo-84677-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84677,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","HETERQA：多种异构来源下的基准记录检索","HETERQA is a comprehensive benchmark for record retrieval in question-answering settings where target records are grounded by multiple heterogeneous sources. The benchmark provides 857 QA pairs over five source types, instantiated with Yelp business records linked to relational tables, text documents, image repositories, spatial databases, and knowledge graphs. It uses an answer-driven construction: candidate records are constrained by record fields, then enriched and cross-verified across required sources, followed by contradiction detection and human validation. Experiments evaluate sparse, dense, hybrid, late-interaction, and agentic retrievers, showing strong challenge and limited saturation.","arXiv :2607 .03028v 1 [ cs .IR] 3 Jul 2026  \nHETERQA: Benchmarking Record Retrieval over  \nMultiple Heterogeneous Sources  \nYaodong Su 1 Hanchang Li 1 Quanqing Xu2 Chuanhui Yang2 Yixiang Fang 1  \n1 CUHK-Shenzhen 2 OceanBase, Ant Group  \n[yaodongsu@link.cuhk.edu.cn](yaodongsu@link.cuhk.edu.cn) [hanchangli@link.cuhk.edu.cn](hanchangli@link.cuhk.edu.cn)[ ](hanchangli@link.cuhk.edu.cn)[xuquanqing.xqq@oceanbase.com](xuquanqing.xqq@oceanbase.com) [rizhao.ych@oceanbase.com](rizhao.ych@oceanbase.com)[ ](rizhao.ych@oceanbase.com)[fangyixiang@cuhk.edu.cn](fangyixiang@cuhk.edu.cn)  \nAbstract  \nIn emerging systems (e.g., social media and e-commerce platforms), data records are often drawn from heterogeneous sources, such as relational tables, text documents, image repositories, spatial databases, and knowledge graphs. Accordingly, retrieving target records for question-answering (QA) tasks requires us to jointly exploit these heterogeneous sources. However, most existing benchmarks are constructed from individual sources, and only a very few recent benchmarkshave considered two or three sources. To alleviate this issue, we introduce HETER QA, a comprehensive benchmark with 857 QA pairs for record retrieval over five heterogeneous sources. HETERQA instantiates this setting with Yelp business records, each of which is grounded by multiple sources. We build HETER QA in an answer-driven manner: candidate records are first initialized with record-field constraints, then enriched through heterogeneous sources, and finally cross-verified across required sources before the natural-language question is retained. We validate the benchmark through contradiction detection and human validation, and further evaluate sparse, dense, hybrid, late-interaction, and agentic retrievers under the same metrics. The results show that HETERQA is challenging: hybrid retrieval achieves the strongest Recall@10, Self-RAG achieves the best MRR@10, and all evaluated methods remain far from saturating the benchmark. These findings indicate that HETERQA provides an effective testbed for record retrieval over heterogeneous sources and leaves substantial room for future retrieval methods. The benchmark dataset and source code are publicly available at [https://huggingface.co/datasets/hanchang02/HeterQA](https://huggingface.co/datasets/hanchang02/HeterQA and)[ and](https://huggingface.co/datasets/hanchang02/HeterQA and)  \n[https://github.com/hanchang02/HeterQA](https://github.com/hanchang02/HeterQA), respectively.  \n1 Introduction  \nIn emerging systems (e.g., social media and e-commerce platforms), natural-language interfaces have been increasingly employed to retrieve data records from heterogeneous sources, such as relational tables, textual documents, image repositories, spatial databases, and knowledge graphs (KGs) . For example, a user may ask for chicken wing restaurants that are highly rated, commented with positive words, close to the current location, and visually similar to a referenced venue. To identify target records for such question-answering (QA) tasks, we have to jointly consider multiple heterogeneous sources, rather than individual sources which suffer from incomplete coverage and reliability.  \nFormally, in the record retrieval of heterogeneous sources, it often assumes there is a collection of records R, and each record r ∈ R is drawn from multiple heterogeneous sources. Given a natural language question q, it aims to return a ranked list of records matching with q. A returned record is correct only when its source bundle satisfies all constraints expressed by q; retrieving an isolated text, image, spatial, or relational clue is insufficient if the record violates another required constraint.  \nPreprint.  \nTable 1: Benchmark comparison where P denotes partial coverage.  \n\n| Benchmark | Relational |  | Data Sources |  |  | Task Feature |\n| --- | --- | --- | --- | --- | --- | --- |\n|  |  | Text | Image | Spatial | Knowledge | Missing-Value |\n|  | Tables |","cbCaitaRNckqt8Xb","https://ap.wps.com/l/cbCaitaRNckqt8Xb","pdf",2646781,3,1,20,"English","en",105,"# Abstract\n# Introduction\n## Problem Setting: Record Retrieval over Heterogeneous Sources\n## Related Work and Benchmark Comparison\n## HETERQA Overview and Evaluation Goal","[{\"question\":\"What problem does HETERQA target?\",\"answer\":\"HETERQA benchmarks record retrieval for QA tasks when a correct record must satisfy constraints expressed in natural language and is grounded by multiple heterogeneous sources.\"},{\"question\":\"How is the HETERQA dataset constructed?\",\"answer\":\"Candidates are initialized using record-field constraints, then enriched through heterogeneous sources and cross-verified across required sources before the question is retained.\"},{\"question\":\"What retrievers and metrics are used to evaluate HETERQA?\",\"answer\":\"The evaluation covers sparse, dense, hybrid, late-interaction, and agentic retrievers, using metrics including Recall@10 and MRR@10, showing the benchmark remains challenging with results far from saturation.\"}]",1784197617,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"heterqa-benchmarking-record-retrieval-over-multiple-heterogeneous-sources","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/heterqa-benchmarking-record-retrieval-over-multiple-heterogeneous-sources/84677/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-20","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does HETERQA target?","Question",{"text":75,"@type":76},"HETERQA benchmarks record retrieval for QA tasks when a correct record must satisfy constraints expressed in natural language and is grounded by multiple heterogeneous sources.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the HETERQA dataset constructed?",{"text":80,"@type":76},"Candidates are initialized using record-field constraints, then enriched through heterogeneous sources and cross-verified across required sources before the question is retained.",{"name":82,"@type":73,"acceptedAnswer":83},"What retrievers and metrics are used to evaluate HETERQA?",{"text":84,"@type":76},"The evaluation covers sparse, dense, hybrid, late-interaction, and agentic retrievers, using metrics including Recall@10 and MRR@10, showing the benchmark remains challenging with results far from saturation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]