[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83559-en":3,"doc-seo-83559-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83559,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","Persona Non Grata: LLM Persona-Driven Generations in MCQA Are Unstable in Distinct Dimensions","Persona-driven generations (PDGs) use an LLM “persona” to complete tasks, yet stability concerns remain underexplored when persona is expressed in non-text-heavy outputs such as multiple-choice question answering (MCQA). This work studies PDG instability in MCQA by introducing three metrics that measure performance, outcome, and question correctness stability across dimensions. Results show instability varies by model family and size, and by question domain, with math/commonsense yielding higher instability. Prompt format changes instability more than temperature.","Persona Non Grata: LLM Persona-Driven Generations in MCQA are Unstable in Distinct Dimensions  \nCésar Guerra-Solano, Xiang Lorraine Li  \nDepartment of Computer Science, University of Pittsburgh  \nCorrespondence: {cguerrasol, [xianglli}@pitt.edu](xianglli}@pitt.edu)  \narXiv :2607 .00937v 1 [ cs .CL] 1 Jul 2026  \nAbstract  \nPersona-driven generations (PDGs) have seen prolific use in research and industry applications, where a large language model (LLM) takes on a “persona” while completing some task. While persona expressed through freeform text (like dialogue) has substantial work investigating stability or consistency, relatively, persona expressed in non-text-heavy outputs (like in multiple-choice question answering, or MCQA) is often overlooked. We work to address this gap, seeking to understand the instability of LLM PDGs in MCQA tasks. We develop three metrics investigating the per  \nformance, outcome, and question correctness stability, evaluating three distinct dimensions. Using these metrics, we find that instability varies consistently between model families and model size, and across question domains, with math/commonsense questions leading to greater instability. We also find task prompt format introduces more prediction instability than other hyperparameters, like temperature. Finally, we find that instability is related to task accuracy, and using our instability metrics, find different experimental settings that result in different best and worst personas for tasks, despite their similarity. This reveals the importance of checking hyperparameter instability in PDGs.  \n1 Introduction  \nLarge language models (LLMs) have seen prolific use in a variety of domains, with their adaptability and broad capabilities lending themselves to high performance across a range of tasks and opportunities for personalization for user-and task-specific needs (Kojima et al., 2022 ; Brown et al., 2020 ; Grattafiori et al., 2024 ; Yang et al., 2025 ; Zhang et al., 2024) . With this, persona-driven generations (PDGs) have become prevalent – here, by leveraging the power of prompting, LLMs can roleplay as a \"persona\" while carrying out some task, such as general user assistance (e.g. \"You are a  \nTPF = justanswer  \na lawyer  \nan old person  \nFigure 1: Depicting instability in non-text-heavy persona-driven generations (PDGs) . Given different, although still reasonable, experiment configurations, represented by different settings for the task prompt format (TPF) hyperparameter, an LLM’s persona-driven performance on a multiple-choice evaluation differs greatly.  \nteacher. Explain [...]\") or more specific task completion (e.g. \"You are a hiring manager. Rate these resumes [...]\") . As PDGs see use in high-stakes domains, from medicine to education, evaluating their stability is critical (Sun et al., 2025 ; Yuan et al., 2025 ; Kyung et al., 2025 ; Li et al., 2024) .  \nWith PDGs, two main uses can be observed: in free-form text generation settings, such as in dialogue systems like Character.AI (Character.AI), and non-text-heavy outputs, such as in persona-based multiple-choice question answering (MCQA) . In the latter case, LLM persona is expressed through task completion: rather than producing text that reflects the style of a persona, an LLM answers questions in accordance with the capabilities associated with that persona, when relevant (e.g., a biologist answering biology questions), or to remain unaffected by persona choice when the persona is irrelevant (e.g., race should not influence performance on mathematical questions) .  \nHowever, relative to text generation, little work has characterized or improved PDGs in MCQA. Prior studies indicate that PDGs in this setting are  \nimplicitly biased and unpredictable (Gupta et al., 2024 ; Zheng et al., 2024), motivating further work to characterize them and their potential flaws. Additionally, these studies bear little standardization, with large differences in model choice and hyperparameters a","cbCaijTP87PBoN3z","https://ap.wps.com/l/cbCaijTP87PBoN3z","pdf",12758140,1,23,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"What problem does the paper study about persona-driven generations in MCQA?\",\"answer\":\"The paper investigates how unstable LLM persona-driven generations can be when the persona affects performance in multiple-choice question answering tasks, rather than free-form dialogue outputs.\"},{\"question\":\"How does the paper measure instability in MCQA?\",\"answer\":\"It develops three metrics that evaluate stability in performance, overall outcomes, and question correctness across distinct experimental dimensions.\"},{\"question\":\"What factors most affect instability according to the results?\",\"answer\":\"Instability differs consistently by model family and size, varies across question domains (with math/commonsense more unstable), and is more influenced by task prompt format than by hyperparameters like temperature.\"}]",1784188814,58,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"persona-non-grata-llm-persona-driven-generations-in-mcqa-are-unstable-in-distinct-dimensions","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/persona-non-grata-llm-persona-driven-generations-in-mcqa-are-unstable-in-distinct-dimensions/83559/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper study about persona-driven generations in MCQA?","Question",{"text":75,"@type":76},"The paper investigates how unstable LLM persona-driven generations can be when the persona affects performance in multiple-choice question answering tasks, rather than free-form dialogue outputs.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper measure instability in MCQA?",{"text":80,"@type":76},"It develops three metrics that evaluate stability in performance, overall outcomes, and question correctness across distinct experimental dimensions.",{"name":82,"@type":73,"acceptedAnswer":83},"What factors most affect instability according to the results?",{"text":84,"@type":76},"Instability differs consistently by model family and size, varies across question domains (with math/commonsense more unstable), and is more influenced by task prompt format than by hyperparameters like temperature.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]