[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85195-en":3,"doc-seo-85195-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85195,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language Models","Reasoning failures in large language models are difficult to diagnose from final answers alone because the same wrong output may reflect missing capability, an unstable reasoning trajectory, or failure to activate latent reasoning states. SPARK introduces a hidden-state susceptibility diagnostic to determine whether a model enters an effective reasoning state, then applies lightweight test-time steering. It addresses strong confounding from prompt length by using length-controlled residual susceptibility and cross-layer coordination, improving Qwen3 series accuracy on GSM8K and MATH-500.","arXiv :2607 . 10296v 1 [ cs .AI] 11 Jul 2026  \nSPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language Models  \nDongxu Zhang1 Yiding Sun1 Zihao Guo1 Xiangyang Yang1 Kai Tang2 Lin Chen1 Cheng Tan3 Jihua Zhu1  \n1Xi’an Jiaotong University 2Peking University 3Tencent  \nReasoning failures in large language models (LLMs) are usually evaluated from final answers, but a wrong answer does not reveal why the model failed. The same incorrect output may reflect missing capability, an unstable reasoning trajectory, or a failure to activate a reasoning state that is already available in the frozen model. Existing prompting and benchmark-based evaluation methods mostly operate at the output level, while generic activation-steering methods typically apply global directions without diagnosing which examples require intervention. In this paper, we introduce SPARK, which uses hidden-state response to diagnose whether a model internally enters an effective reasoning state and to guide lightweight test-time steering. The key observation is that raw hidden-state susceptibility is strongly confounded by prompt length, especially in programmatic and algorithmic reasoning where harder serialized instances naturally become longer. SPARK therefore uses length-controlled susceptibility to separate input-scale effects from residual reasoning activation, and combines this signal with cross-layer coordination to select reasoning-active anchors and under-activated hard examples. We use FRONTIER-4.5K as a controlled programmatic reasoning suite for latent profiling and difficulty-aware analysis, and evaluate SPARK-Steering on GSM8K and MATH-500 with forward-only benchmark profiling. Our method improves Qwen3 series models consistently; on MATH-500, accuracy rises from 82.0% to 84.6% for Qwen3-4B and from 82.4% to 85.6% for Qwen3-8B. These results suggest that susceptibility can serve not only as a diagnostic signal for reasoning failures, but also as a practical guide for targeted test-time intervention.  \nKeywords: Large Language Models, Chain-of-Thought, Latent Reasoning States  \nContact: Dongxu Zhang, [zhangdongxu@stu.xjtu.edu.cn](zhangdongxu@stu.xjtu.edu.cn)  \n1 Introduction  \nLarge language models (LLMs) have made rapid progress on mathematical, symbolic, and algorithmic reasoning tasks [33, 34, 12], yet their reasoning failures remain difficult to interpret from final answers alone [40] . A model may solve easy instances reliably, become unstable on moderately difficult inputs, and fail sharply as problem structure grows [26, 27, 9] . Benchmark accuracy captures this behavioral pattern only at the output level [24, 16] . It tells us when performance drops, but not what internal condition has changed. This matters because two incorrect answers may correspond to different internal states. In one case, the model may not possess the required capability [10] . In another, it may have partial capability but fail to activate the latent computation needed for the current input [31] . Moreover, hard reasoning problems often become longer and more structurally dense [43], so weak  \nWhy Benchmark Accuracy is Not Enough?  \nFigure 1: Motivation and overview of SPARK. Benchmark evaluation observes correct and wrong answers and summarizes performance with accuracy curves, but leaves the latent reasoning state hidden. SPARK complements this output-level view with latent profiling and test-time steering. It diagnoses hidden-state response, selects reasoning-active anchors and under-activated hard targets, and evaluates whether steering improves harder examples while preserving easier cases.  \nhidden-state response may reflect input scale rather than a true collapse of reasoning.  \nCurrent evaluation practice creates a gap between observing failure and deciding how to intervene. Accuracy curves can locate a behavioral capability boundary, but they do not specify what should be changed when a hard example fails. Prompting treats th","cbCaiaVWuLDMgxJJ","https://ap.wps.com/l/cbCaiaVWuLDMgxJJ","pdf",1607640,3,1,14,"English","en",105,"# Introduction\n## Why Benchmark Accuracy is Not Enough?\n## Hidden-state Susceptibility and Length Control\n## SPARK for Profiling and Steering","[{\"question\":\"Why can’t benchmark accuracy alone explain LLM reasoning failures?\",\"answer\":\"Accuracy curves summarize correct vs. wrong outputs, but they do not reveal which internal condition changed. Different wrong answers may correspond to different internal states such as missing capability or failure to activate latent computation.\"},{\"question\":\"What does SPARK measure to diagnose latent reasoning states?\",\"answer\":\"SPARK measures hidden-state response by assessing how strongly transformer hidden states react when small noise is applied to input embeddings. This yields a susceptibility signal tied to internal reasoning activation.\"},{\"question\":\"How does SPARK avoid confusing susceptibility with prompt length effects?\",\"answer\":\"Raw susceptibility is strongly confounded by prompt length, especially in longer serialized algorithmic instances. SPARK uses length-controlled residual susceptibility to separate input-scale effects from residual reasoning activation and then guides steering accordingly.\"}]",1784201677,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"spark-susceptibility-guided-profiling-and-steering-of-latent-reasoning-states-in-large-language-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/spark-susceptibility-guided-profiling-and-steering-of-latent-reasoning-states-in-large-language-models/85195/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why can’t benchmark accuracy alone explain LLM reasoning failures?","Question",{"text":75,"@type":76},"Accuracy curves summarize correct vs. wrong outputs, but they do not reveal which internal condition changed. Different wrong answers may correspond to different internal states such as missing capability or failure to activate latent computation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does SPARK measure to diagnose latent reasoning states?",{"text":80,"@type":76},"SPARK measures hidden-state response by assessing how strongly transformer hidden states react when small noise is applied to input embeddings. This yields a susceptibility signal tied to internal reasoning activation.",{"name":82,"@type":73,"acceptedAnswer":83},"How does SPARK avoid confusing susceptibility with prompt length effects?",{"text":84,"@type":76},"Raw susceptibility is strongly confounded by prompt length, especially in longer serialized algorithmic instances. SPARK uses length-controlled residual susceptibility to separate input-scale effects from residual reasoning activation and then guides steering accordingly.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]