[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85566-en":3,"doc-seo-85566-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85566,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Filtered Reasoning Score Evaluating Reasoning Quality on a Model’s Most Confident Traces","Large Language Models can reach high accuracy on reasoning benchmarks, yet outcome-based evaluation fails to reveal the reasoning quality behind correct answers. This paper introduces a reasoning score that evaluates reasoning traces across faithfulness, coherence, utility, and factuality, aiming to distinguish models with similar accuracy and remain robust to prompt and generation changes. It further proposes the Filtered Reasoning Score (FRS), aggregating only the top-K% most confident traces to better reflect deployed selection behavior. Results show FRS correlates with both accuracy and reasoning transfer across benchmarks.","arXiv :2604 . 1 1996v2 [ cs .CL] 13 Jul 2026  \nFiltered Reasoning Score: Evaluating Reasoning Quality on a Model’s Most Confident Traces  \nManas Pathak, Xingyao Chen, Shuozhe Li, Amy Zhang, Leqi Liu  \nUniversity of Texas at Austin  \n{manaspathak, caspar.cxy06, [shuozhe.li](shuozhe.li) , [leqiliu](leqiliu}@utexas.edu)[}](leqiliu}@utexas.edu)[@utexas.edu](leqiliu}@utexas.edu) , [amy.zhang@austin.utexas.edu](amy.zhang@austin.utexas.edu)  \nAbstract  \nShould we trust Large Language Models (LLMs) with high accuracy? LLMs  \nachieve high accuracy on reasoning benchmarks, but correctness alone does  \nnot reveal the quality of the reasoning used to produce it. This highlights a  \nfundamental limitation of outcome-based evaluation: models may arrive at  \ncorrect answers through flawed reasoning, and models with substantially  \ndifferent reasoning capabilities can nevertheless exhibit similar benchmark  \naccuracy, for example due to memorization or over-optimization. In this  \npaper, we ask: given existing benchmarks, can we move beyond outcome  \nbased evaluation to assess the quality of reasoning itself? We seek metrics  \nthat (1) differentiate models with similar accuracy and (2) are robust to  \nvariations in input prompts and generation configurations. To this end, we  \npropose a reasoning score that evaluates reasoning traces along dimensions  \nsuch as faithfulness, coherence, utility, and factuality. A remaining question  \nis how to aggregate this score across multiple sampled traces. Naively  \naveraging them is undesirable, particularly in long-horizon settings, where  \nthe number of possible trajectories grows rapidly, and low-confidence  \ncorrect traces are more likely to be coincidental. To address this, we in  \ntroduce the Filtered Reasoning Score (FRS), which computes reasoning  \nquality using only the top-K% most confident traces. Evaluating with  \nFRS, models that are indistinguishable under standard accuracy exhibit  \nsignificant differences in reasoning quality. Moreover, models with higher  \nFRS on one benchmark tend to perform better on other reasoning bench  \nmarks, in both accuracy and reasoning quality. Together, these findings  \nsuggest that FRS complements accuracy by capturing a model’s trans  \nferable reasoning capabilities. We open source our evaluation codebase:  \n[https://github.com/Manas2006/benchmark](https://github.com/Manas2006/benchmark)   reproducibility.  \nTwo correct answers, very different reasoning quality  \nTrace A  Score: 100/100   \nProblem: Octagon with same perimeter as hexagon of side 16 cm. Find octagon side length. Perimeter = 6 × 16 = 96 cm. Octagon has 8 sides: 8s = 96 ⇒ s = 12. 12  \nTrace B  Score: 25/100   \nProblem: GCF of 6432 and 132, increased by 11 .  \nLists factors, concludes GCF = 4, gets 4+11=15 [wrong] → ... 23  \nFigure 1: Two traces from different models produce correct final answers and receive the same Pass@1 score, yet their reasoning quality is vastly different.  \n1 Introduction  \nLarge language models have advanced rapidly in recent years, with newer models achieving higher scores on an expanding set of reasoning benchmarks (Hendrycks et al., 2021; Cobbe  \net al., 2021; Rein et al., 2024) . Yet the way we evaluate these models has not kept pace. The dominant paradigm remains final-answer accuracy: a model is scored by how often it produces the correct output, with no regard for the reasoning process that produced it.  \nThis paradigm is increasingly inadequate. Models can produce flawed reasoning that still leads to correct answers (Lightman et al., 2024; Uesato et al., 2022; Turpin et al., 2023) . As a result, accuracy gains do not reliably reflect improvements in reasoning quality (Xia et al., 2025), and benchmark saturation further reduces their ability to distinguish between models (Deveci et al., 2025) . Moreover, outcome-based evaluation can be sensitive to prompt choice and generation configuration, further obscuring differences in underlying reasoning ability (Hochlehner","cbCaiaPsj2OXJcTv","https://ap.wps.com/l/cbCaiaPsj2OXJcTv","pdf",1844130,2,1,33,"English","en",105,"# Abstract\n# Introduction\n## Motivation: limitations of outcome-based evaluation\n## Shift to evaluating reasoning traces\n## Aggregation challenge and proposed FRS\n# Figure 1\n# Figure 2","[{\"question\":\"Why is final-answer accuracy insufficient for evaluating reasoning models?\",\"answer\":\"Correct answers can be produced using flawed reasoning, so outcome-based scores do not reflect reasoning quality. Benchmark saturation and sensitivity to prompt/generation choices can also obscure differences between models.\"},{\"question\":\"What does the proposed reasoning score evaluate?\",\"answer\":\"The rubric-based reasoning score rates reasoning traces along dimensions including faithfulness, coherence, utility, and factuality.\"},{\"question\":\"How does the Filtered Reasoning Score (FRS) work and why is it used?\",\"answer\":\"FRS samples multiple reasoning traces, estimates per-trace confidence from token-level probabilities, keeps only the top-K% most confident traces, and computes the final reasoning quality score. This focuses evaluation on the high-confidence region that better matches what deployed systems select.\"}]",1784204641,83,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"filtered-reasoning-score-evaluating-reasoning-quality-on-a-models-most-confident-traces","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/filtered-reasoning-score-evaluating-reasoning-quality-on-a-models-most-confident-traces/85566/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is final-answer accuracy insufficient for evaluating reasoning models?","Question",{"text":75,"@type":76},"Correct answers can be produced using flawed reasoning, so outcome-based scores do not reflect reasoning quality. Benchmark saturation and sensitivity to prompt/generation choices can also obscure differences between models.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the proposed reasoning score evaluate?",{"text":80,"@type":76},"The rubric-based reasoning score rates reasoning traces along dimensions including faithfulness, coherence, utility, and factuality.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the Filtered Reasoning Score (FRS) work and why is it used?",{"text":84,"@type":76},"FRS samples multiple reasoning traces, estimates per-trace confidence from token-level probabilities, keeps only the top-K% most confident traces, and computes the final reasoning quality score. This focuses evaluation on the high-confidence region that better matches what deployed systems select.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]