[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82960-en":3,"doc-seo-82960-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82960,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Prompt Robustness Is Task Dependent: Comparing Objective and Belief Style Questions in LLM Evaluation","Survey-style evaluations often treat an LLM’s prompted responses as evidence of the model’s values or beliefs, but this link is fragile when answers are interpreted as political values or social attitudes. This work tests whether prompt robustness differs between objective questions with fixed correct answers and subjective opinion or value queries. Four instruction-tuned model families are evaluated across objective (MMLU, ARC, CulturalBench) and subjective (Political Compass Test, ValueBench, World Values Survey) datasets using controlled prompt perturbations.","Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation  \nSadia Kamal†, Arefa Patwary‡, Anthony Marchiafava†, Atriya Sen†, Sagnik Ray Choudhury‡†Oklahoma State University, ‡University of North Texas  \n{sadia.kamal,anmarch,[atriya.sen}@okstate.edu](atriya.sen}@okstate.edu) , [arefapatwary@my.unt.edu](arefapatwary@my.unt.edu) ,[sagnik.raychoudhury@unt.edu](sagnik.raychoudhury@unt.edu)  \narXiv :2607 .05554v 1 [ cs .CL] 6 Jul 2026  \nAbstract  \nSurvey-style evaluations of large language models often treat a prompted response as a measure of a model’s values or beliefs. This assumption is particularly fragile when responses are read as evidence of political values, social attitudes, or beliefs. We ask whether prompt robustness differs between objective questions with fixed answers and subjective questions that ask for opinions or values. We evaluate four instruction-tuned model families on three objective datasets (MMLU, ARC, and CulturalBench) and three subjective datasets (Political Compass Test, ValueBench, and World Values Survey) . For each question/statement, we apply multiple types of prompt changes, such as variations in wording, framing, and format, and measure whether the model gives the same answer across variants. Using a binomial generalized estimating equation, we find significant effects of model, dataset, prompt category, and their interactions. The dataset type effect is also significant, and the interaction between dataset type and prompt category is large. These results show that prompt robustness depends on the question type, the prompt change, and the model.  \n1 Introduction  \nMost large language model evaluations rest on a fragile assumption that one prompt gives a stable and meaningful measure of model behavior. Prior work shows this is often not the case, small changes in prompt format, wording, and answer presentation can change model behavior (Sclar et al., 2024 ; Chatterjee et al., 2024 ; Ismithdeen et al., 2025) . Ifa model changes its answer when the task meaning stays the same, then the evaluation is measuring both task ability and prompt sensitivity.  \nThis is a critical concern for survey-style tests. Recent work uses political and value surveys to infer what LLMs “believe” or which human groups they resemble. But survey responses from  \nLLMs are unstable under ordering, labeling, forcedchoice wording, and framing changes (DominguezOlmedo et al., 2024 ; Röttger et al., 2024 ; Rupprechtet al., 2025) . A model’s answer can be shaped by the way the prompt asks the question.  \nThis raises a question about whether prompt sensitivity behaves the same way for different types of questions. We study this problem by comparing two kinds of questions. Type-I questions consist of objective, multiple-choice items with a single correct answer. Type-II questions are subjective survey items that ask for opinions, values, or degrees of agreement. When a prompt is reworded without altering its underlying meaning, a robust model is expected to produce the same answer across both versions. The same applies for objective questions as well. For subjective questions, however, models may interpret minor changes in wording or response options as cues about how to respond. As a result, the model’s answer can shift even when the survey item itself has not changed in substance.  \nWe ask three research questions: RQ1: Does response consistency differ between objective and subjective question types? RQ2: Do objective and subjective questions show different sensitivity patterns across prompt categories? RQ3: Is prompt robustness a model-level property, or does it also depend on dataset and prompt category?  \nOur study makes three contributions. First, we curate a broad perturbation set from prior work on prompt sensitivity, multiple-choice formatting, survey response bias, and value measurement. Second, we give a unified robustness test across objective and subjective datasets. Third, we","cbCaisZ9Ms5CzWLu","https://ap.wps.com/l/cbCaisZ9Ms5CzWLu","pdf",329647,4,1,7,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"How does the document define objective versus subjective questions in LLM evaluation?\",\"answer\":\"Objective questions are multiple-choice items with a single correct answer. Subjective questions are survey items that ask for opinions, values, or degrees of agreement.\"},{\"question\":\"What kinds of prompt changes are used to test robustness?\",\"answer\":\"For each question or statement, multiple prompt perturbations are applied, including variations in wording, framing, and format, and then consistency across prompt variants is measured.\"},{\"question\":\"What does the analysis conclude about prompt robustness differences?\",\"answer\":\"Prompt robustness depends on the question type, the specific prompt change category, and the model. Subjective questions are less stable under prompt variation than objective questions, with option-order changes causing the largest consistency drop.\"}]",1784184345,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"prompt-robustness-is-task-dependent-comparing-objective-and-belief-style-questions-in-llm-evaluation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/prompt-robustness-is-task-dependent-comparing-objective-and-belief-style-questions-in-llm-evaluation/82960/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the document define objective versus subjective questions in LLM evaluation?","Question",{"text":75,"@type":76},"Objective questions are multiple-choice items with a single correct answer. Subjective questions are survey items that ask for opinions, values, or degrees of agreement.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What kinds of prompt changes are used to test robustness?",{"text":80,"@type":76},"For each question or statement, multiple prompt perturbations are applied, including variations in wording, framing, and format, and then consistency across prompt variants is measured.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the analysis conclude about prompt robustness differences?",{"text":84,"@type":76},"Prompt robustness depends on the question type, the specific prompt change category, and the model. Subjective questions are less stable under prompt variation than objective questions, with option-order changes causing the largest consistency drop.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]