[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83125-en":3,"doc-seo-83125-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83125,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Healthier LLMs Retrieval-Augmented Generation for Public Health Question Answering","Large language models (LLMs) perform well on medical question answering benchmarks, yet applying them to public health faces major constraints from hallucinations and rapidly changing official guidance. Retrieval-Augmented Generation (RAG) addresses these issues by grounding answers in an explicitly maintained corpus, but end-to-end quality depends heavily on retrieval configuration and on evaluation methods beyond multiple-choice formats. The work extends PubHealthBench to a retrieval-augmented setting and compares dense, sparse, and hybrid retrieval across embedding models and corpus variants. Hybrid retrieval improves recall and ranking quality, and retrieved context raises multiple-choice accuracy across diverse LLMs, enabling smaller open-weight models to match or exceed larger models without retrieval. A rubric-based LLM-as-a-judge evaluates free-form faithfulness, completeness, clarity, and factual consistency, validated against dual human annotations.","arXiv :2607 .0664 1v 1 [ cs .CL] 7 Jul 2026  \nHealthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering  \nFelix Feldman  \nFan Grayson  \nJoshua Harris Timothy Laurence Leo Loman Ollie Higgins  \nPoonam Soma Bethany Pace-Bonello Michael Borowitz  \nToby Nonnenmacher  \nAbstract  \nLarge language models (LLMs) achieve promising results on medical question answering benchmarks, yet their use in public health is constrained by hallucinations and the rapid evolution of official guidance. Retrieval-Augmented Generation (RAG) mitigates these risks by grounding responses in an explicitly maintained corpus, but end-to-end performance depends critically on retrieval configuration and on evaluation beyond multiple-choice formats. We extend PubHealthBench, a question answering (QA) benchmark of 7,929 questions derived from UK Government public health guidance, into a retrieval-augmented setting and systematically evaluate retrieval and generation choices. We compare dense, sparse, and hybrid retrieval across multiple embedding models and corpus variants, and show that hybrid retrieval consistently improves recall and ranking quality, with chunk length and topic interacting with ranking performance. Providing retrieved context substantially increases multiple-choice accuracy across a diverse set ofLLMs, enabling smaller open-weight models to match or outperform larger models used without retrieval, with gains primarily driven by retrieval quality and careful context selection.  \nTo assess realistic free-form answering, we introduce a rubric-based LLM-as-ajudge covering faithfulness, completeness, clarity, and factual consistency, and validate it against dual human annotations. Judge–human agreement is strongest for faithfulness and completeness, while factual consistency and clarity are less reliably reproduced, motivating caution when interpreting those dimensions at scale. Overall, our results highlight retrieval as a primary lever for reliable public health QA and provide practical guidance for building and evaluating RAG systems grounded in official guidance.  \n1 Introduction  \nArtificial intelligence (AI) is playing an expanding role in public health, from chatbots that answer health queries [1] to decision-support tools for public health professionals [2] . Large language models (LLMs) can already generate coherent answers based on extensive training data, in some cases approaching expert performance on medical question answering tasks [3] . This potential has spurred interest in deploying LLMs in public health, both as public-facing tools and as decision aids for public health professionals. Unlike many clinical decision-support settings—where tools are designed to support decisions for an individual patient at the point of care, public health guidance is population-level, often precautionary, and closely tied to official recommendations that are updated as evidence and policy evolve [4, 5, 6] . However, LLMs can hallucinate information or provide outdated advice, which in public health can have severe implications: even small inaccuracies may harm individual health decision-making [7] and, at scale, drive inappropriate or dangerous behaviours  \n39th Conference on Neural Information Processing Systems (NeurIPS 2025) .  \nacross populations [8] . Therefore, ensuring that LLM responses are reliable and up to date is a prerequisite for safe AI adoption in public health.  \nOne established approach to improving LLM performance and reliability is Retrieval-Augmented Generation (RAG) [9] . In RAG systems, an LLM is coupled with an external retrieval component that selects relevant documents from a knowledge base; the model then conditions its output on this retrieved context rather than relying solely on its parametric memory [10, 11] . By incorporating retrieval, an LLM can expand and update its effective knowledge beyond what is stored in its frozen parameters, mitigating hallucinations and reducing the impact of outd","cbCaimJ28Gex50cb","https://ap.wps.com/l/cbCaimJ28Gex50cb","pdf",938420,4,1,19,"English","en",105,"# Abstract\n# Introduction\n## Public health challenges for LLM reliability\n## Retrieval-Augmented Generation (RAG) and grounding\n## Motivation for realistic QA benchmarks\n# Benchmark extension and research questions","[{\"question\":\"Why is reliable public health question answering difficult for LLMs?\",\"answer\":\"Public health guidance is population-level and precautionary, updates frequently as evidence and policy evolve, and LLMs may hallucinate or give outdated advice. Small inaccuracies can have serious consequences for both individuals and behavior across populations.\"},{\"question\":\"How does Retrieval-Augmented Generation (RAG) improve public health QA?\",\"answer\":\"RAG adds an external retrieval component that selects relevant documents and conditions the LLM output on the retrieved context. This reduces reliance on frozen parametric knowledge and helps answers stay grounded in trusted, up-to-date guidance.\"},{\"question\":\"What retrieval approach and evaluation strategy does the work emphasize?\",\"answer\":\"The paper extends PubHealthBench to a retrieval-augmented setting and systematically compares dense, sparse, and hybrid retrieval, showing hybrid retrieval improves recall and ranking quality. It also introduces a rubric-based LLM-as-a-judge for free-form responses, validated against dual human annotations.\"}]",1784185449,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"healthier-llms-retrieval-augmented-generation-for-public-health-question-answering","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/healthier-llms-retrieval-augmented-generation-for-public-health-question-answering/83125/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is reliable public health question answering difficult for LLMs?","Question",{"text":75,"@type":76},"Public health guidance is population-level and precautionary, updates frequently as evidence and policy evolve, and LLMs may hallucinate or give outdated advice. Small inaccuracies can have serious consequences for both individuals and behavior across populations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Retrieval-Augmented Generation (RAG) improve public health QA?",{"text":80,"@type":76},"RAG adds an external retrieval component that selects relevant documents and conditions the LLM output on the retrieved context. This reduces reliance on frozen parametric knowledge and helps answers stay grounded in trusted, up-to-date guidance.",{"name":82,"@type":73,"acceptedAnswer":83},"What retrieval approach and evaluation strategy does the work emphasize?",{"text":84,"@type":76},"The paper extends PubHealthBench to a retrieval-augmented setting and systematically compares dense, sparse, and hybrid retrieval, showing hybrid retrieval improves recall and ranking quality. It also introduces a rubric-based LLM-as-a-judge for free-form responses, validated against dual human annotations.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]