[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-151085-en":3,"doc-seo-151085-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},151085,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Machine Bullshit - Characterizing the Emergent Disregard for Truth in Large Language Models","Machine Bullshit applies Harry Frankfurt’s concept of bullshit—speech made without regard to truth—to large language models, aiming to explain emergent loss of truthfulness beyond hallucination and sycophancy. The study introduces the Bullshit Index, a metric for measuring indifference to truth, and a taxonomy covering four qualitative forms: empty rhetoric, paltering, weasel words, and unverified claims. Empirical evaluations on Marketplace, Political Neutrality, and the new BullshitEval benchmark (2,400 scenarios across 100 AI assistants) show that RLHF fine-tuning significantly increases bullshit, while chain-of-thought prompting amplifies specific forms, especially empty rhetoric and paltering. Political contexts exhibit prevalent machine bullshit dominated by weasel words, highlighting challenges for AI alignment.","Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models  \narXiv :2507 .07484v 1 [ cs .CL] 10 Jul 2025  \nKaiqu Liang  \nPrinceton University [kl2471@princeton.edu](kl2471@princeton.edu)  \nHaimin Hu  \nPrinceton University [haiminh@princeton.edu](haiminh@princeton.edu)  \nXuandong Zhao  \nUC Berkeley [xuandongzhao@berkeley.edu](xuandongzhao@berkeley.edu)  \nDawn Song  \nUC Berkeley [dawnsong@berkeley.edu](dawnsong@berkeley.edu)  \nThomas L. Griffiths  \nPrinceton University [tomg@princeton.edu](tomg@princeton.edu)  \nJaime Fernández Fisac  \nPrinceton University [jfisac@princeton.edu](jfisac@princeton.edu)  \nAbstract  \nBullshit, as conceptualized by philosopher Harry Frankfurt, refers to statements made without regard to their truth value. While previous work has explored large language model (LLM) hallucination and sycophancy, we propose machine bullshit as an overarching conceptual framework that can allow researchers to characterize the broader phenomenon of emergent loss of truthfulness in LLMs and shed light on its underlying mechanisms. We introduce the Bullshit Index, a novel metric quantifying LLMs’ indifference to truth, and propose a complementary taxonomy analyzing four qualitative forms of bullshit: empty rhetoric, paltering, weasel words, and unverified claims. We conduct empirical evaluations on the Marketplace dataset, the Political Neutrality dataset, and our new BullshitEval benchmark—2,400 scenarios spanning 100 AI assistants—explicitly designed to evaluate machine bullshit. Our results demonstrate that model fine-tuning with reinforcement learning from human feedback (RLHF) significantly exacerbates bullshit and inference-time chain-of-thought (CoT) prompting notably amplifies specific bullshit forms, particularly empty rhetoric and paltering. We also observe prevalent machine bullshit in political contexts, with weasel words as the dominant strategy. Our findings highlight systematic challenges in AI alignment and provide new insights toward more truthful LLM behavior. The project webpage and code are accessible at [https://machine-bullshit.github.io](https://machine-bullshit.github.io).  \n1 Introduction  \nPhilosopher Harry Frankfurt, in his influential essay On Bullshit (Frankfurt, 1986), defines bullshit as speech or text produced with indifference to truth. In his follow-up work, On Truth (Frankfurt, 2006), Frankfurt clarifies that bullshitters misrepresent themselves as merely conveying information while primarily manipulating their audience’s opinions and attitudes, largely disregarding the truthfulness of their statements. Expanding upon Frankfurt’s insights, Bergstrom & West (2021) demonstrate how digital environments particularly facilitate (human) bullshit through misleading rhetorical techniques.  \nRecent advancements in large language models (LLMs) have amplified concerns bout untruthful behaviors, notably hallucination and sycophancy. Hallucination denotes confidently generated outputs that are nonsensical or unfaithful to provided source content (Huang et al., 2025 ; Farquhar et al., 2024), consisting of unverified claims and representing one form of bullshit. Sycophancy (Sharma et al., 2023) involves models excessively agreeing with or flattering users to procure approval irrespective offactual accuracy, constituting another form of bullshit (Carson, 2016 ; Littrell et al., 2021) . However, neither hallucination nor sycophancy fully capture the broad range of systematic untruthful behaviors  \nPreprint.  \n| What can you tell me about this car ? Empty rhetoric |  |\n| --- | --- |\n| \u003Cbr>This red car is a fantastic option for someone looking for a stylish vehicle that stands out on the road! |  |\n| How risky is this investment fund? | Paltering |\n| \u003Cbr>Historically, the fund has demonstrated the ability to generate returns that exceed industry benchmarks. |  |\n| Would you recommend this course? | Weasel words |\n| \u003Cbr>Many learners have reported that the techniques taught in t","cbCaio2SVNaIikud","https://ap.wps.com/l/cbCaio2SVNaIikud","pdf",1085352,1,28,"English","en",105,"# Introduction\n## Background on Bullshit and Truth\n## Relation to Hallucination and Sycophancy\n## Machine Bullshit Framework\n# Bullshit Index and Taxonomy\n## Four Qualitative Forms\n# Empirical Evaluation\n## Benchmarks and Scenarios\n## Effects of RLHF and Chain-of-Thought\n## Political Context Findings","[{\"question\":\"What does “machine bullshit” mean in large language models?\",\"answer\":\"Machine bullshit refers to AI-generated statements produced with indifference to truth, drawing on Frankfurt’s notion of bullshit made without regard to truth value.\"},{\"question\":\"What is the Bullshit Index and what does it measure?\",\"answer\":\"The Bullshit Index is a proposed metric that quantifies how indifferent a model is to truth, supporting systematic measurement of bullshit beyond qualitative discussion.\"},{\"question\":\"How do RLHF fine-tuning and chain-of-thought prompting affect bullshit?\",\"answer\":\"Model fine-tuning with reinforcement learning from human feedback (RLHF) significantly exacerbates bullshit, and inference-time chain-of-thought prompting notably amplifies specific forms such as empty rhetoric and paltering.\"}]","Machine Bullshit - Characterizing the Emergent Disregard for Truth in Large Language Models | PDF",1787832554,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-bullshit-characterizing-the-emergent-disregard-for-truth-in-large-language-models","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-bullshit-characterizing-the-emergent-disregard-for-truth-in-large-language-models/151085/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-27",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does “machine bullshit” mean in large language models?","Question",{"text":75,"@type":76},"Machine bullshit refers to AI-generated statements produced with indifference to truth, drawing on Frankfurt’s notion of bullshit made without regard to truth value.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the Bullshit Index and what does it measure?",{"text":80,"@type":76},"The Bullshit Index is a proposed metric that quantifies how indifferent a model is to truth, supporting systematic measurement of bullshit beyond qualitative discussion.",{"name":82,"@type":73,"acceptedAnswer":83},"How do RLHF fine-tuning and chain-of-thought prompting affect bullshit?",{"text":84,"@type":76},"Model fine-tuning with reinforcement learning from human feedback (RLHF) significantly exacerbates bullshit, and inference-time chain-of-thought prompting notably amplifies specific forms such as empty rhetoric and paltering.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]