[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83159-en":3,"doc-seo-83159-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83159,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Evaluating LLM Robustness Under Domain-Specific Prompt Perturbations in Public Health Applications","Large language models (LLMs) are increasingly used for public health applications, but robustness to non-clinical user inputs is not well studied. A domain-specific benchmark is proposed to test two realistic perturbations: misinformation framing (false health claims injected into prompts) and layperson rewriting (symptom descriptions in everyday language). Results show misinformation framing harms accuracy by −7.2 pp on average with 9–38% prediction flips, while layperson rewriting causes only −1.4 pp degradation, revealing distinct deployment risks.","Evaluating LLM Robustness Under Domain-Specific Prompt Perturbations in Public Health Applications  \nChuqing Zhao  \nSchool of Engineering and Applied Sciences Harvard University Boston, MA  \nHaochen Yang  \nSchool of Engineering and Applied Sciences Harvard University Boston, MA  \narXiv :2607 .069 13v 1 [ cs .CY] 8 Jul 2026  \nAbstract—Large language models (LLMs) are increasingly applied in public health applications, yet their robustness to non-clinical user inputs remains underexplored. We propose a domain-specific robustness benchmark that evaluates LLMs under two perturbation types that commonly arise when nonclinical users interact with health AI systems: misinformation framing (MF), where prompt might be injected by false health claims, and layperson rewriting (LR), where patients describe symptoms in everyday language rather than medical terminology. Our goal is to evaluate the stability of LLMs under these perturbation. Experiments show that MF degrades accuracy by −7.2 pp on average with prediction flip rates of 9–38%, even when claims are explicitly labelled as unsupported; LR causes only −1.4 pp degradation. These findings highlight two distinct deployment risks in public health settings: models may produce incorrect outputs when users unintentionally carry misinformation into their queries, and may misinterpret clinically relevant details when patients use informal language. Both risks call for perturbation-aware robustness evaluation beyond clean baseline benchmark.  \nIndex Terms—Natural Language Processing , Large Language Model, Robustness  \nI. INTRODUCTION  \nLarge language models (LLMs) are increasingly deployed in public health and clinical decision support, from answering biomedical literature questions to assisting with exam-style clinical reasoning and monitoring health-related discourse on social media [1], [2] . Benchmarks such as PubMedQA and MedQA report strong performance on clean, professionally phrased prompts, and recent systems are often evaluated as if users always present well-formed, evidence-aligned queries. However, in real-world settings patients and the general public frequently use colloquial or imprecise health vocabulary in prompts. These conditions are not captured by standard benchmarks, leaving a critical gap in our understanding of LLM robustness under real-world public health application scenarios.  \nPrompt robustness, the ability of LLMs to maintain accurate and consistent output under input variation, remains underexplored in the public health domain. Existing robustness benchmarks such as PromptRobust [3] evaluate general-purpose NLP tasks under character-level noise and synonym substitution, while biomedical evaluations such as RAmBLA [4] focus primarily on paraphrase variation. A few studies have examined LLM robustness in prompt injection, such as embedded  \nmisinformation, and lexical-level perturbations, such as layperson terminology, within a unified public health benchmark.  \nIn this paper, we address this gap by evaluating four lightweighted LLMs (Llama-3.1-8B, Mistral-7B, Qwen2.5-7B, and GPT-4.1-Nano) across three public health datasets under two perturbation types: (1) misinformation framing (MF), in which a contradictory or false claim is injected into the prompt; (2) layperson rewriting (LR), in which medical terminology is replaced with colloquial equivalents drawn from the Consumer Health Vocabulary [5] . We measure accuracy drop and output consistency to characterize robustness profiles across model families. Our contributions are as follows:  \n• Domain-Specific Benchmark. We construct a domainspecific robustness benchmark covering three public health tasks and two perturbation types grounded in realistic deployment scenarios.  \n• Perturbation Comparison. We provide a new, novel empirical comparison of misinformation framing and layperson rewriting as structurally distinct perturbation types, revealing divergent failure patterns across multiple models.  \n• Practical ","cbCaiesnUAAJKGgv","https://ap.wps.com/l/cbCaiesnUAAJKGgv","pdf",229359,3,1,5,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n## LLMs in Public Health and Clinical NLP\n## Prompt Robustness in Explainable NLP","[{\"question\":\"What problem does the paper address in public health LLM deployments?\",\"answer\":\"It addresses the lack of understanding about how well LLMs handle non-clinical user inputs that differ from clean, professionally phrased prompts.\"},{\"question\":\"What are the two perturbation types evaluated in the benchmark?\",\"answer\":\"The benchmark evaluates misinformation framing (injecting contradictory or false claims) and layperson rewriting (replacing medical terminology with colloquial equivalents).\"},{\"question\":\"How do the two perturbations affect model accuracy and prediction stability?\",\"answer\":\"Misinformation framing reduces accuracy by −7.2 pp on average and produces prediction flip rates of 9–38%, even when unsupported claims are explicitly labeled. Layperson rewriting causes a smaller −1.4 pp degradation.\"}]",1784185665,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"evaluating-llm-robustness-under-domain-specific-prompt-perturbations-in-public-health-applications","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/evaluating-llm-robustness-under-domain-specific-prompt-perturbations-in-public-health-applications/83159/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in public health LLM deployments?","Question",{"text":75,"@type":76},"It addresses the lack of understanding about how well LLMs handle non-clinical user inputs that differ from clean, professionally phrased prompts.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the two perturbation types evaluated in the benchmark?",{"text":80,"@type":76},"The benchmark evaluates misinformation framing (injecting contradictory or false claims) and layperson rewriting (replacing medical terminology with colloquial equivalents).",{"name":82,"@type":73,"acceptedAnswer":83},"How do the two perturbations affect model accuracy and prediction stability?",{"text":84,"@type":76},"Misinformation framing reduces accuracy by −7.2 pp on average and produces prediction flip rates of 9–38%, even when unsupported claims are explicitly labeled. Layperson rewriting causes a smaller −1.4 pp degradation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":22,"slug":137},19,"General","general"]