[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-160356-en":3,"doc-seo-160356-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},160356,962085564381,"Bintang","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",8,"Research & Report","AI Achieves a Perfect LSAT Score - Research Report","The paper documents the first officially disclosed instance of a language model achieving a perfect Law School Admission Test (LSAT) score. Controlled tests across eight reasoning models show prompt changes, answer-choice shuffling, and multi-sample selection do not meaningfully change performance. Removing the generated thinking phase reduces frontier accuracy by up to 8 percentage points, mainly in logical reasoning. Distilled models reproduce full thinking-trace formats yet plateau below frontier results. A pilot QLoRA fine-tuned reward model using Best-of-5 selection narrows the gap predominantly for logical reasoning.","arXiv :2604 . 10034v1 [ cs .AI] 11 Apr 2026  \nAI Achieves a Perfect LSAT Score  \nBonmu Ku  \n[bku@alumni.harvard.edu](bku@alumni.harvard.edu)  \nAbstract  \nThis paper reports the first documented instance of a language model achieving a perfect score on an officially disclosed Law School Admission Test (LSAT) .  \nControlled experiments on eight reasoning models show that varying the prompt, shuffling answer choices, and sampling multiple responses have no meaningful effect as drivers of performance. Ablating the thinking phase that models generate before answering, however, lowers frontier accuracy by up to 8 percentage points, predominantly in logical reasoning. Distilled models produce full thinking traces in the same format yet plateau far below frontier performance. A pilot process reward model fine-tuned via QLoRA on official LSAT explanations narrows this gap through Best-of-5 selection, with gains again predominantly in logical reasoning.  \nThe gatekeeper of elite legal education since 1948, the LSAT has not merely been passed but answered without a single error by models that reason. The upper bound of the cognitive capacities it has tested is no longer exclusive to human cognition.  \n1 Introduction  \n1.1 The LSAT as a Benchmark for Logical Reasoning  \nThe Law School Admission Test (LSAT) stands as the most rigorous standardized examination of logical reasoning. The test requires no prior legal training. It instead requires exceptional reasoning capability to identify logical structures, evaluate arguments, and draw valid inferences from complex linguistic input. This exclusive measure of inferential skill makes the LSAT a particularly clean benchmark for evaluating the reasoning capabilities of language models as well.  \nEven the strongest human performers fall below the perfect score of 180, with a median of 174 among Harvard Law School admits. Language models have been measured against the same standard. GPT-3.5 scored 149 at the 40th percentile, and GPT-4 scored 163 at the 88th percentile [1] . Each generation of language models narrows the gap to the human ceiling. A perfect score, however, demands the capacity for extended multi-step reasoning and deliberation, a level of performance that has thus far remained unmet for language models.  \n1.2 From Language Models to Reasoning Models  \nThe leap from GPT-4 to the models that followed was not a continuation of prior scaling trends [2, 3] . It marked the emergence of a fundamentally new inference paradigm: the reasoning model [4, 5, 6] . The central capability was generating extended thinking traces at inference time before producing a final answer [7, 8, 9, 10] . This advance developed incrementally, through three phases that moved explicit reasoning progressively deeper into the model.  \nReasoning via Prompting  \nChain-of-thought prompting [11] dramatically improved performance on reasoning tasks by eliciting intermediate reasoning steps from the model. Zero-shot chain-of-thought [12] extended this result to the setting without few-shot exemplars, and self-consistency [13] further improved accuracy by  \nPreprint.  \nmajority-voting over multiple sampled reasoning paths. Still, the model did not yet internalize the reasoning process. Intermediate steps were elicited externally and remained sensitive to prompt phrasing.  \nReasoning via Training  \nSTaR [14] bootstrapped reasoning ability by generating chain-of-thought rationales, fine-tuning on those that yielded correct answers, and iterating. Quiet-STaR [15] generalized this by training models to produce internal rationales at every token position during general text prediction. On the evaluation side, outcome reward models (ORMs) [16] scored candidate solutions by final-answer correctness, while process reward models (PRMs) [17] scored each reasoning step independently and provided a stronger selection signal.  \nReasoning via Inference  \nModern reasoning models are trained end-to-end via reinforcement learning to gene","cbCaine4j6QzBHDV","https://ap.wps.com/l/cbCaine4j6QzBHDV","pdf",525862,1,35,"English","en",105,"# Introduction\n## The LSAT as a Benchmark for Logical Reasoning\n## From Language Models to Reasoning Models\n# Models\n## Evaluated Reasoning Models","[{\"question\":\"What does the paper claim about AI performance on the LSAT?\",\"answer\":\"It reports the first officially disclosed case where a language model achieves a perfect LSAT score.\"},{\"question\":\"Which experimental changes were found to not meaningfully affect performance?\",\"answer\":\"Changing prompts, shuffling answer choices, and sampling multiple responses showed no meaningful effect across the evaluated reasoning models.\"},{\"question\":\"How does removing the model’s thinking phase affect results?\",\"answer\":\"Ablating the thinking phase lowers frontier accuracy by up to 8 percentage points, especially for logical reasoning tasks.\"}]","AI Achieves a Perfect LSAT Score - Research Report | PDF",1788062230,88,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"ai-achieves-a-perfect-lsat-score-research-report","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/ai-achieves-a-perfect-lsat-score-research-report/160356/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-30",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does the paper claim about AI performance on the LSAT?","Question",{"text":75,"@type":76},"It reports the first officially disclosed case where a language model achieves a perfect LSAT score.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which experimental changes were found to not meaningfully affect performance?",{"text":80,"@type":76},"Changing prompts, shuffling answer choices, and sampling multiple responses showed no meaningful effect across the evaluated reasoning models.",{"name":82,"@type":73,"acceptedAnswer":83},"How does removing the model’s thinking phase affect results?",{"text":84,"@type":76},"Ablating the thinking phase lowers frontier accuracy by up to 8 percentage points, especially for logical reasoning tasks.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]