[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85945-en":3,"doc-seo-85945-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85945,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Articulate Intuition or Genuine Analysis? Benchmarking Epistemic Reliability in LLM-as-a-Judge Peer Review","When an LLM judge labels a peer review as “analytical” while human committees call a different review “high quality,” the two assessments may not reflect the same underlying construct. The work operationalizes Kahneman’s dual-process theory into a structured peer-review rubric and releases Kahneman4Review, scoring 3,563 reviews across nine textual dimensions, eight bias diagnostics, and a continuous reasoning-quality metric. Results show weak alignment between decision tiers and the text-grounded proxy, explainability limits for length/venue effects, and temporal shifts in review diagnostics around 2022–2023.","arXiv :2607 . 105 1 1v 1 [ cs .CL] 12 Jul 2026  \nArticulate Intuition or Genuine Analysis? Benchmarking Epistemic Reliability in LLM-as-a-Judge Peer Review  \nNuo Chen Qian Wang Qingyun Zou Bingsheng He  \nNational University of Singapore  \nAbstract  \nWhen an LLM judge calls a peer review “analytical” and a human committee calls another review “high quality,” are they tracking the same thing? We argue they are not, and that the difference matters philosophically. We operationalise Kahneman’s dual-process theory into a structured rubric for peer review and release Kahneman4Review, a benchmark of 3 ,563 rated reviews scored along nine theoretically motivated textual dimensions, eight bias diagnostics, and a continuous reasoning-quality score. Three findings bear on trustworthiness: decision tier is not detectably aligned with the rubric’s text-grounded epistemic-quality proxy; publicshowcase agentic reviews receive higher raw scores than pooled human reviews, but length and venue explain most of the gap and the samples are not paper-paired; and ICLR review-text diagnostics shift at the 2022–2023 transition, temporally coincident with widespread LLM availability but without identifying its cause. A matched function-probe pilot further shows that the rubric distinguishes textual probes designed to contrast genuine fault-finding with surface fluency. We argue that a trustworthy reliability benchmark for LLM judges must separate analytical form from epistemic function, and propose concrete design choices toward that goal. An interactive demo is available at [https://huggingface.co/spaces/nuojohnchen/Kahneman4Review](https://huggingface.co/spaces/nuojohnchen/Kahneman4Review).  \n1 Introduction  \nPeer review is the most widely used epistemic gatekeeping mechanism in science, and large language models are increasingly proposed both to produce reviews 1,2 and to judge review quality 3,4. Deploying LLMs in either role invites a philosophical question that the machine learning community rarely poses directly: what counts as a good review? An answer is usually smuggled in through training objectives and benchmark design (acceptance prediction, score agreement, humanpreference alignment), yet philosophers have long argued that epistemic quality is not a scalar outcome but a property of the reasoning itself5,6,7. A review may be technically correct and still textually under-justified; a negative review may be high-quality if it is grounded, falsifiable, and responsive to rebuttal.  \nPeer review offers an unusually clean test bed for this problem. First, the review is a written artefact amenable to fine-grained analysis. Second, there exist outcomes (acceptance, meta-review, AC flags) that are typically treated as ground truth but which our data will show are not detectably aligned with the rubric’s text-grounded epistemic-quality proxy. Third, the recent appearance of agentic AI reviewers 2, 1 creates a natural textual contrast between reviews written by humans under deadline and outputs from retrieval-augmented LLM pipelines with full paper access. The contrast motivates a dual-process-inspired question for ML evaluation: can an LLM judge identify textual evidence that separates articulate intuition from genuine analysis without mistaking analytical form for verified epistemic function?  \nContributions. (1) We propose nine theoretically motivated textual dimensions, eight bias diagnostics, a four-label taxonomy, and a continuous reasoning-quality score, grounded in analytic epistemology and philosophy of explanation. (2) We rate 3 ,563 reviews using claude-sonnet-4-6 as the judge: 1 , 155 stratified-sample human reviews from ICLR, ICML, and NeurIPS 2025; 2 ,307 ICLR 2021/2022/2023 human reviews for the longitudinal study; and 101 reviews from an open agentic AI-reviewer pipeline. The released code, rubric prompt, ratings, and provenance metadata are available through the Kahneman4Review demo. (3) We report three empirical findings bearing dire","cbCaif18vRHK5V48","https://ap.wps.com/l/cbCaif18vRHK5V48","pdf",432611,3,1,15,"English","en",105,"# Introduction\n## Contributions\n# Related Work\n## Philosophy and ML evaluation\n# A Dual-Process Rubric for Review Epistemics","[{\"question\":\"What question does the paper address about LLM judge reliability in peer review?\",\"answer\":\"It asks whether LLMs judging peer reviews as “analytical” track the same concept that human committees use when calling reviews “high quality,” and why the discrepancy matters for epistemic reliability.\"},{\"question\":\"How is Kahneman4Review built and what does it measure?\",\"answer\":\"The paper turns Kahneman’s dual-process theory into a structured rubric, rating 3,563 reviews along nine textual dimensions, running eight bias diagnostics, and adding a continuous reasoning-quality score.\"},{\"question\":\"What key evidence is reported about trustworthiness of the rubric-based LLM judge?\",\"answer\":\"Decision tier shows no detectable alignment with the rubric’s text-grounded epistemic-quality proxy; agentic review samples score higher, but most gaps are explained by length and venue and the data are not paper-paired.\"}]",1784207305,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"articulate-intuition-or-genuine-analysis-benchmarking-epistemic-reliability-in-llm-as-a-judge-peer-review","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/articulate-intuition-or-genuine-analysis-benchmarking-epistemic-reliability-in-llm-as-a-judge-peer-review/85945/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What question does the paper address about LLM judge reliability in peer review?","Question",{"text":75,"@type":76},"It asks whether LLMs judging peer reviews as “analytical” track the same concept that human committees use when calling reviews “high quality,” and why the discrepancy matters for epistemic reliability.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is Kahneman4Review built and what does it measure?",{"text":80,"@type":76},"The paper turns Kahneman’s dual-process theory into a structured rubric, rating 3,563 reviews along nine textual dimensions, running eight bias diagnostics, and adding a continuous reasoning-quality score.",{"name":82,"@type":73,"acceptedAnswer":83},"What key evidence is reported about trustworthiness of the rubric-based LLM judge?",{"text":84,"@type":76},"Decision tier shows no detectable alignment with the rubric’s text-grounded epistemic-quality proxy; agentic review samples score higher, but most gaps are explained by length and venue and the data are not paper-paired.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]