[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81672-en":3,"doc-seo-81672-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81672,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","Code Correctness Signals in LLM Hidden States: Pre-Generation Probing and Repair Geometry","Large language models encode information in hidden states, and this study tests whether code correctness is legible there before generation and during self-repair. Experiments use Qwen3-4B-Instruct-2507 on 444 LiveCodeBench tasks. The prompt-final hidden state linearly decodes first-attempt code correctness with leakage-free AUC 0.955±0.006 and remains high after removing prompt-length effects. On cleaned repair cases, hidden-state shifts toward repairs form a stable geometric direction of success, significant under magnitude and split-half tests and robust to conditional residualization against repair-context covariates.","arXiv :2606 . 14530v2 [ cs .LG] 10 Jul 2026  \nCode Correctness Signals in LLM Hidden States: Pre-Generation Probing and Repair Geometry  \nCarlo Di Cicco∗  \nAbstract  \nLarge language models encode rich information in their hidden states. This work asks whether code correctness is legible in the hidden states of Qwen3-4B-Instruct-2507, before it generates and as it repairs a failed attempt, studied on 444 LiveCodeBench tasks. It reports two findings connected by a single confound-control tool: residualization. First, the correctness of the model’s first-attempt code is linearly decodable from the prompt-final hidden state, with a leakage-free held-out AUC of 0.955 ± 0.006 across 50 outer splits. After the linear effect of prompt length is removed from each hidden state dimension, the probe still reaches 0.940±0 .009 , well above a prompt-length baseline of 0.720 ± 0.015. Second, on 246 cleaned cases where the model attempts to repair a failed first attempt, the hidden state shift from the failing attempt to its repair carries a robust geometric direction of repair success, significant on both a magnitude and a split-half test against label-shuffled nulls. This direction remains significant on both tests after a conditional residualization against three repair-context covariates that differ between successful and failed repairs, indicating a genuine repair-success signal that the observable repair context does not account for. The contribution is as much methodological as empirical, a confound-control diagnostic that reports each signal only to the extent it survives the control.  \n1 Introduction  \nReasoning has become a central axis of large language model development. With each generation, models train on more and better data with improved training methods, and become capable of tackling problems that would have been out of reach a few years ago. Their internals, in turn, grow more informative: a richer internal computation is doing the work behind richer external behavior, which makes the question of what those internals encode increasingly worth asking. Code generation is a particularly clean field for that question: outputs are easily verifiable against unit tests, with no human-judgment loop required to decide whether an answer is correct. As model performance on code-generation benchmarks has improved rapidly, the question is no longer only whether a model can produce correct code, but also what happens internally as it does so. This shifts the focus toward mechanistic interpretability in models that now exhibit genuine capability in code generation.  \nA standard tool for asking what is encoded inside a frozen language model is the linear probe: a simple linear classifier trained on a hidden state vector to predict some property of interest, where probe performance is read as a measure of how linearly decodable that property is from the representation [1] . Two recent works have studied code correctness through hidden states [2, 3], reading hidden states of already-generated code to judge whether it is correct without running the unit tests. One trains a classifier on those states; the other contrasts correct and incorrect samples  \n∗ Independent researcher. Contact: [cdicicco2001@gmail.com](cdicicco2001@gmail.com. Code)[. Code](cdicicco2001@gmail.com. Code), data, and analysis scripts: [https://github](https://github). com/CarloDiCicco/ReasoningLab  \nto extract a correctness direction and rank candidate generations. In both, the hidden state is read after the model has produced an answer.  \nThis work asks a related but earlier question: whether correctness is already decodable from the hidden state at the position of the last prompt token, captured on a single forward pass over the prompt and before any output token is sampled. Whereas those works operate on hidden states of generated code, the probe in this paper operates on the hidden state of the prompt alone. This is amore upstream interpretability claim: the informat","cbCairdp0X4KRpJh","https://ap.wps.com/l/cbCairdp0X4KRpJh","pdf",407952,10,1,12,"English","en",105,"# Introduction\n## Linear probes and hidden-state correctness\n## Pre-generation interpretability via prompt states\n## Repair-success geometry and representation engineering\n## Confound control and residualization","[{\"question\":\"How is code correctness evaluated in the hidden states before any output is generated?\",\"answer\":\"The study trains a linear probe on the prompt-final hidden state to predict correctness of the model’s first-attempt code without using generated code states. It reports a leakage-free held-out AUC across multiple splits.\"},{\"question\":\"What geometric signal is used to detect whether a repair will succeed?\",\"answer\":\"For cases where the model repairs a failed attempt, the method computes hidden-state deltas between the failed attempt and the repair attempt. It then tests whether the contrastive direction between success and failure is stable using magnitude and split-half tests against label-shuffled nulls.\"},{\"question\":\"Why does the paper emphasize residualization and confound-control?\",\"answer\":\"The paper explains that some residualizations can collapse signals if covariate means match across groups, while other covariates can drive the direction when their means differ. A distribution test selects residualizations, and the main repair-success direction remains significant after conditional residualization against repair-context covariates.\"}]",1784175329,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"code-correctness-signals-in-llm-hidden-states-pre-generation-probing-and-repair-geometry","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/code-correctness-signals-in-llm-hidden-states-pre-generation-probing-and-repair-geometry/81672/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"How is code correctness evaluated in the hidden states before any output is generated?","Question",{"text":76,"@type":77},"The study trains a linear probe on the prompt-final hidden state to predict correctness of the model’s first-attempt code without using generated code states. It reports a leakage-free held-out AUC across multiple splits.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What geometric signal is used to detect whether a repair will succeed?",{"text":81,"@type":77},"For cases where the model repairs a failed attempt, the method computes hidden-state deltas between the failed attempt and the repair attempt. It then tests whether the contrastive direction between success and failure is stable using magnitude and split-half tests against label-shuffled nulls.",{"name":83,"@type":74,"acceptedAnswer":84},"Why does the paper emphasize residualization and confound-control?",{"text":85,"@type":77},"The paper explains that some residualizations can collapse signals if covariate means match across groups, while other covariates can drive the direction when their means differ. A distribution test selects residualizations, and the main repair-success direction remains significant after conditional residualization against repair-context covariates.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":122},"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":20,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]