[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82705-en":3,"doc-seo-82705-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82705,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","A Scalable Approach to Evaluating Moral Sensitivity in LLMs","Moral sensitivity is the capacity to detect morally relevant features in a decision situation and use them to guide action, underpinning wider moral competence. Existing evaluations of LLM moral reasoning often show shortcomings due to limited scenario coverage, reliance on highly curated vignettes, and poor scalability. This paper proposes a new, more scalable evaluation that avoids expensive human baselines and avoids using an LLM judge requiring the very capability under test.","arXiv :2607 .02972v 1 [ cs .CY] 3 Jul 2026  \nA Scalable Approach to Evaluating Moral Sensitivity  \nin LLMs  \nDaniel Kilov 1,∗,†, Secil Yanik Guyot 1,∗, Caroline Hendy 1 , Sichao Li2 , and Seth Lazar 1, 3  \n1Australian National University  \n2The University of Sydney  \n3Johns Hopkins University  \n∗These authors contributed equally to this work.  \n†[Correspondence to](Correspondence to daniel.kilov@anu.edu.au)[ daniel.kilov@anu.edu.au](Correspondence to daniel.kilov@anu.edu.au)  \nAbstract  \nMoral sensitivity is the ability to identify the morally relevant features of a decision situation and use them as the basis for action. It is the foundation of broader moral competence: any other moral reasoning capabilities will be irrelevant if an agent lacks sensitivity to the relevant facts. Recent research on LLMs’ moral reasoning has largely concluded that the models fall short. In this paper, we offer a new evaluation of LLM moral sensitivity which presents a more optimistic picture. In doing so, we address and resolve a central problem in AI alignment research: how to scale behavioural evaluations beyond expensive and sometimes metaethically dubious comparisons with a human baseline, without adopting an LLM judge that must be assumed to have the very capability that you are attempting to evaluate. Our central question is this: can LLMs successfully identify the morally relevant features of noisy cases, in which various kinds of morally irrelevant information have been introduced to distract the respondent? To explore this, we introduce MORPH-1K (MOral Robustness under Perturbed Hypotheticals), a procedurally-generated 1,000-case benchmark spanning 50 moral foundationpole combinations across four social domains. MORPH-1K is paired with a suite of textual moral distractors, irrelevant detail additions, and embedded chat histories, along with a method for validating that the distractors do not change the morally salient content of the case. We claim that morally competent respondents should identify the same morally relevant features in the perturbed cases as they do in the unperturbed cases, measured by focusing on per-pair semantic stability between responses to clean and perturbed cases, using a fixed embedding. We apply MORPH-1K to eight contemporary LLMs, and show that while morally irrelevant perturbations often changed the number of features listed, the semantic content of those features remained stable across all noise conditions, with similarity scores above our calibrated floor threshold. More broadly, our invariance framework extends to evaluative domains where ground truth is difficult to specify but relevant and irrelevant features can be separated by design.  \n1 Introduction  \nLLMs are already taking morally significant actions. They offer personal moral advice to millions of users (McCain et al., 2025; Shen et al., 2026) . They are used in content moderation systems that adjudicate permissible speech (Franco et al., 2025) . They guide medical doctors’ diagnoses (Brodeur et al., 2026; Blease et al., 2025), and lawyers’ research and litigation decision-making (Terzidou, 2025; Legal Services Research Centre, 2026) . And they are increasingly not just advising human  \nPreprint.  \nagents, but acting as agents themselves (Janjeva et al., 2026) . For AI systems to engage acceptably in these morally-freighted behaviors, they must be morally competent—that is, able to identify the morally relevant features of the situations before them, and reason from them to a sensible conclusion about what to do (Snoswell et al., 2026; Haas et al., 2026) . The foundation for moral competence is moral sensitivity, the ability to pick out the morally relevant features of a choice situation (Railton, 2020; Kilov et al., 2025; Kwon et al., 2023; Chiu et al., 2025) . Evaluating LLM moral sensitivity is therefore an urgent and important challenge.  \nExisting evaluations of LLM moral sensitivity have at least three limitations (Snoswell et al., 2026) .","cbCaitt0dSU1Bx3Q","https://ap.wps.com/l/cbCaitt0dSU1Bx3Q","pdf",682351,1,24,"English","en",105,"# Abstract\n# 1 Introduction\n## Motivation and limitations of existing evaluations\n## MORPH-1K benchmark and perturbation design\n## Semantic stability method for evaluation","[{\"question\":\"What is “moral sensitivity” in the context of LLMs?\",\"answer\":\"Moral sensitivity is the ability to identify the morally relevant features of a decision situation and base action on them, serving as a foundation for broader moral competence.\"},{\"question\":\"Why do the authors argue existing evaluations are insufficient?\",\"answer\":\"They claim current evaluations are limited by narrow scenario coverage, reliance on highly curated moral vignettes that miss noisy realistic conditions, and weak scalability due to costly human baselines or LLM-judge bootstrapping issues.\"},{\"question\":\"What is MORPH-1K and what does it measure?\",\"answer\":\"MORPH-1K is a procedurally generated 1,000-case benchmark spanning moral foundation–pole combinations across multiple social domains, measuring whether models preserve morally relevant feature identification under perturbed hypotheticals.\"}]",1784182406,60,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"a-scalable-approach-to-evaluating-moral-sensitivity-in-llms","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/a-scalable-approach-to-evaluating-moral-sensitivity-in-llms/82705/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is “moral sensitivity” in the context of LLMs?","Question",{"text":75,"@type":76},"Moral sensitivity is the ability to identify the morally relevant features of a decision situation and base action on them, serving as a foundation for broader moral competence.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why do the authors argue existing evaluations are insufficient?",{"text":80,"@type":76},"They claim current evaluations are limited by narrow scenario coverage, reliance on highly curated moral vignettes that miss noisy realistic conditions, and weak scalability due to costly human baselines or LLM-judge bootstrapping issues.",{"name":82,"@type":73,"acceptedAnswer":83},"What is MORPH-1K and what does it measure?",{"text":84,"@type":76},"MORPH-1K is a procedurally generated 1,000-case benchmark spanning moral foundation–pole combinations across multiple social domains, measuring whether models preserve morally relevant feature identification under perturbed hypotheticals.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":28,"slug":108},5,"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]