[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86126-en":3,"doc-seo-86126-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86126,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","ResearchQA 基于引用的科学论文问答基准评测","Large language models support scientific reading, yet standard evaluations often miss whether answers are backed by verifiable citations. ResearchQA introduces a benchmark of 6,211 single-paper question-answer pairs from 494 open-access papers across eight domains and four question types. The benchmark targets citation-grounded evaluation, allowing multiple valid supporting passages and rewarding grounded refusal when a paper does not support an answer. Experiments on leading closed- and open-weight models use a deterministic citation matcher and an LLM rubric evaluator, and release the dataset and evaluation assets.","arXiv :2607 . 1 1074v 1 [ cs .CL] 13 Jul 2026  \nResearchQA: Benchmarking Citation-Grounded Question-Answering on Scientific Papers  \nSaba Imran Debanjum Singh Solanky  \nKhoj Inc.  \nMay 4, 2026  \nAbstract  \nLarge language models are increasingly used to assist scientific reading, but existing evaluation methods often fail to detect whether answers are supported by verifiable citations. We introduce ResearchQA, a benchmark of 6,211 single-paper question-answer pairs from 494 open-access papers spanning eight domains and four question types: lookup, comprehension, multi-hop, and adversarial. ResearchQA is designed for citation-grounded evaluation: it permits multiple valid supporting passages for a claim and rewards grounded refusal when the source paper does not support an answer. We evaluate eight leading closed-and open-weight models in a citation-grounded chat-with-paper setting using a deterministic citation matcher and an LLM-based rubric evaluator. Citation-based metrics separate systems more clearly than LLM-evaluator scores: section coverage and citation accuracy vary substantially across models, while evaluator scores remain tightly compressed. We further find that open-weight models approach the best closed-model citation accuracy while achieving 3 to 6 times lower per-example latency. We release the benchmark, evaluation harness, and evaluator prompt.  \nIntroduction  \nResearchers increasingly rely on large language models to read and reason over scientific literature, where two properties matter more than fluency: (1) every claim should be traceable to verifiable evidence in the source, and (2) the answer should surface the complete, relevant context rather than a confident subset of it. Fabricated or misattributed citations propagate false assumptions and erode trust in the findings that build on them. Standard evaluation methods fail to measure these attributes: LLM evaluator scores reward plausible prose, and embedding similarity passes fabricated-but-on-topic text. We introduce ResearchQA, a benchmark that measures citation-grounded answering directly, pairing LLM-generated questions with a deterministic check that a cited passage actually appears in the paper it claims to.  \nInformation retrieval benchmarks have a long history, but the existing experiments did not satisfy our research constraints. They either target a narrower corpus, the different granularity of evidence, or omit the failure modes that matter most for a citation-grounded product. HotPotQA is grounded in Wikipedia and constructed for multi-hop reasoning; it is well-curated and widely used, but Wikipedia articles lack the dense numerical and methodological content—statistics, tables, methods sections, findings—that a “chat with this paper” feature is constantly asked to ground. We eliminated it for that reason.  \nThe closest prior art to our experiment is QASPER: a curated corpus of NLP papers with questions, answers, and supporting evidence extracted by human annotators. QASPER’s design choices—paper-level scope, evidence-anchored answers, multiple question types—directly inform ours. The two reasons it was not enough on its own are that it covers only NLP papers (we need a domain-diverse corpus to measure a general-purpose research dataset), and that its annotation pipeline used paid graduate-student labelers, which conflicted with the scale of the rows-per-paper density we wanted. The open question that motivatesa substantial part of this work is whether a strong frontier LLM, paired with a deterministic verification step against the source text, can stand in for human annotators well enough to extend QASPER’s design to  \na multi-domain dataset at low cost.  \nTwo methodology pitfalls shape the design that follows. LLM evaluator scoring suffers from rubric collapse (the evaluators concentrate at the endpoints of a 1–5 scale unless every level is anchored); citation evaluation suffers from a verbatim-vs-paraphrase tradeoff between substring match","cbCaijY1ljs73OyX","https://ap.wps.com/l/cbCaijY1ljs73OyX","pdf",2757559,4,1,20,"English","en",105,"# Abstract\n# Introduction\n# Methods\n## Dataset construction\n## Evaluation setup and scoring\n## Metrics and robustness considerations","[{\"question\":\"ResearchQA评测的核心目标是什么？\",\"answer\":\"评测重点是答案是否有可验证的引文支撑，而不仅仅是语言是否流畅或内容是否看似合理。对无法由论文支持的答案，基准会奖励“带依据的拒答”。\"},{\"question\":\"ResearchQA包含哪些规模与覆盖范围？\",\"answer\":\"基准包含6,211条单论文问答对，来源于494篇开放获取论文，覆盖八个领域与四类问题：lookup、comprehension、多跳以及对抗式问题。\"},{\"question\":\"评测如何判断模型回答是否被论文支持？\",\"answer\":\"采用确定性的引用匹配器，核查声称引用的支持段落是否确实出现在原论文中；同时还使用LLM驱动的rubric评分器对输出进行评分。\"}]",1784208694,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"researchqa-citation-grounded-question-answering-benchmarking-on-scientific-papers","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/researchqa-citation-grounded-question-answering-benchmarking-on-scientific-papers/86126/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"ResearchQA评测的核心目标是什么？","Question",{"text":75,"@type":76},"评测重点是答案是否有可验证的引文支撑，而不仅仅是语言是否流畅或内容是否看似合理。对无法由论文支持的答案，基准会奖励“带依据的拒答”。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"ResearchQA包含哪些规模与覆盖范围？",{"text":80,"@type":76},"基准包含6,211条单论文问答对，来源于494篇开放获取论文，覆盖八个领域与四类问题：lookup、comprehension、多跳以及对抗式问题。",{"name":82,"@type":73,"acceptedAnswer":83},"评测如何判断模型回答是否被论文支持？",{"text":84,"@type":76},"采用确定性的引用匹配器，核查声称引用的支持段落是否确实出现在原论文中；同时还使用LLM驱动的rubric评分器对输出进行评分。","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]