[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82961-en":3,"doc-seo-82961-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82961,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment","The study investigates why large language models display a yes–no bias, showing that apparent moral verdict shifts originate from how answers are presented rather than from changes in underlying moral judgment. It argues that LLM evaluations using single or limited wordings misattribute framing effects to ethics, since repetition fails to average surface-form sensitivity. A psychometric battery uses crossed symmetrization over logically equivalent dilemma forms to disentangle stance from artifacts, decomposing the bias into order and lexical components.","The yes–no bias of large language models reflects answer order and wording,  \nnot shifts in moral judgment  \nHaonan Huang  \nPrinceton University, Princeton, NJ 08540, USA  \n[hnhuang@princeton. edu](hnhuang@princeton. edu)  \nattached bias vanishes; the pull follows the printed label. Two interpretable numbers, a framing susceptibility and amoral decisiveness, summarize the artifact, and deliberation typically shrinks it. Measuring what an AI values requires crossing the frames of the question, not asking once.  \nKeywords: large language models | moral judgment | psychometrics | framing effects | AI evaluation  \nHow an agent ranks a moral dilemma depends on how it is asked. In people this framing sensitivity is real but mild [1– 8]; in large language models (LLMs) it is severe—moral and social judgments move under changes of wording, option order, and answer format that carry no logical content [9–23] . The stakes are practical: a model’s judgment is typically read out as a binary verdict—a safety gate, a survey’s forced choice, a judge model’s approve/reject [18, 24–27]—so what a shifted verdict means determines what the evaluation measured: a changed moral judgment, a preference for an answer word, or a pull toward a position on the page. Two properties of LLMs make the question urgent and the naive remedy fail. Repetition does not help: at temperature 1, replicates of a single form are nearly unanimous, while the variance across logically equivalent forms runs up to an order of magnitude larger (Methods) [28, 29] . Resampling one wording thus underrepresents the variation that matters, and that variation is large: LLMs are hypersensitive to surface form [15–17, 21], so a questionnaire built for humans—who need only a handful of framings—does not transfer.  \nWe build the instrument the measurement requires: apsychometric battery whose core operation is crossed symmetrization—every logically irrelevant perturbation applied in balanced flip-pairs, the perturbations crossed so that their contributions separate—repeated across a corpus of logically equivalent question forms. A single form, even symmetrized, is a noisy instrument; a family of forms is a measurement. The sharpest instance, and the anchor for everything below, is the finding of Cheung, Maier, and Lieder [9] that on matched moral dilemmas LLMs show amplified cognitive biases—among them a yes–no bias absent in their human comparison: the models’ verdicts shift with logically irrelevant features of the yes/no question far more than people’s do—and hundreds of resamples per item did not average the shift away. We take that finding as established and ask what the shift is made of. When a model’s verdict moves toward “no,” is the model drawn to the last-printed option  \n(an order bias)? To the word itself (a lexical bias)? Or to the negative verdict (a logical bias)? Repetition cannot tell these apart, and neither can any single wording, however carefully chosen: in the standard question—“. . . answer yes or no”—the three candidates coincide on the same token. Most value-probing evaluations elicit each judgment in exactly this way, through a single fixed wording [24, 30–33] or a handful of format or order variants that leave the scenario’s substance unchanged [10, 34–36]; however precisely such a design measures the shift, only crossing can attribute it—the question’s verb against the printed order of the answers against the label that carries the verdict.  \nThe battery poses the same twenty dilemmas—nineteen of them verbatim from Cheung et al.’s materials, so the stance axis and the human anchors remain comparable (Methods)—through three instruments. A graded rating (I-1) elicits moral acceptability under crossed, logically equivalent forms: two scales (0–10 and 0–100), both anchor directions, three wordings, and both poles of the action (rating the action and rating its complement) . A two-alternative free choice (I-2) elicits a committed decision with no scale at all. A","cbCais6fxiMgvBm1","https://ap.wps.com/l/cbCais6fxiMgvBm1","pdf",2475672,5,1,15,"English","en",105,"# Measuring framing sensitivity in LLM moral judgments\n## Why naive repetition and single wording fail\n## Crossed symmetrization psychometric battery\n## Instruments: graded rating, free choice, forced binary\n## Results and interpretable decomposition of the bias","[{\"question\":\"Why does the yes–no bias appear in large language models during moral evaluation?\",\"answer\":\"Because the model’s binary verdict is strongly influenced by answer order and wording/labels rather than by genuine shifts in moral judgment. The study shows these surface-form artifacts can dominate what evaluators interpret as ethics.\"},{\"question\":\"How does the proposed evaluation method separate stance from measurement artifacts?\",\"answer\":\"It uses a psychometric battery built on crossed symmetrization, applying logically irrelevant perturbations in balanced flip-pairs across many logically equivalent question forms. This crossing assigns observed shifts to specific sources like verb order or labels.\"},{\"question\":\"What components make up the decomposed yes–no bias in the results?\",\"answer\":\"The apparent yes–no bias splits into an order bias toward the last-printed option and a lexical pull toward the word “no,” concentrated in a particular model family. Under arbitrary answer labels, the verdict-attached bias nearly vanishes.\"}]",1784184345,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"the-yes-no-bias-of-large-language-models-reflects-answer-order-and-wording-not-shifts-in-moral-judgment","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/the-yes-no-bias-of-large-language-models-reflects-answer-order-and-wording-not-shifts-in-moral-judgment/82961/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why does the yes–no bias appear in large language models during moral evaluation?","Question",{"text":76,"@type":77},"Because the model’s binary verdict is strongly influenced by answer order and wording/labels rather than by genuine shifts in moral judgment. The study shows these surface-form artifacts can dominate what evaluators interpret as ethics.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the proposed evaluation method separate stance from measurement artifacts?",{"text":81,"@type":77},"It uses a psychometric battery built on crossed symmetrization, applying logically irrelevant perturbations in balanced flip-pairs across many logically equivalent question forms. This crossing assigns observed shifts to specific sources like verb order or labels.",{"name":83,"@type":74,"acceptedAnswer":84},"What components make up the decomposed yes–no bias in the results?",{"text":85,"@type":77},"The apparent yes–no bias splits into an order bias toward the last-printed option and a lexical pull toward the word “no,” concentrated in a particular model family. Under arbitrary answer labels, the verdict-attached bias nearly vanishes.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]