[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84936-en":3,"doc-seo-84936-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84936,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories","Latent reasoning methods perform multi-step inference within a model’s continuous hidden states, aiming for more compact and efficient reasoning than token-level chain-of-thought. Hidden-state opacity raises faithfulness concerns: whether latent steps causally drive the final answer. Prior work studies faithfulness only at converged checkpoints, leaving training-time formation unexplored. This work tracks faithfulness across saved checkpoints using counterfactual input edits and noise-ablation activation patches to reveal stage- and answer-format dependent causal decay.","Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories  \nHengyu Jin1,2,3 , Shu Yang2,3 , Di Wang2,3 *  \n1Tongji University  \n2Provable Responsible AI and Data Analytics (PRADA) Lab  \n3 King Abdullah University of Science and Technology  \narXiv :2607 .06648v 1 [ cs .LG] 7 Jul 2026  \nAbstract  \nLatent reasoning methods perform multi-step inference entirely in the model’s continuous hidden states, promising more compact and efficient reasoning. However, these opaque hidden states raise a question of faithfulness: whether these latent reasoning steps causally drive the final answer. Prior work investigates this question at converged checkpoints and reports several unfaithful behaviors, such as latent reasoning steps that can be replaced without changing the answer, but leaves how these behaviors form during training unexamined. We instead track how faithfulness evolves across saved checkpoints for different latent reasoning paradigms, applying a verifiable counterfactual edit on the input and a noise-ablation activation patch on the latent reasoning steps. We find that (i) at the output level, latent reasoning methods can look similarly unfaithful at convergence under counterfactual edits while following qualitatively divergent trajectories;  \n(ii) at the activation level, the causal contribution of latent reasoning steps to the final answer decays across training for both paradigms, with the examples that flip on the output side in (i) also being the examples on which this contribution decays; and (iii) the activation-level trajectory diverges by answer format, decaying on binary choice and rising on open-ended decoding. These findings highlight that latent reasoning faithfulness depends on training stage and answer format.  \n1 Introduction  \nChain-of-thought reasoning represents intermediate computation as natural-language tokens, improving multi-step reasoning by exposing a stepby-step trace (Wei et al., 2022 ; Nye et al., 2021) . This representation comes at a cost: every reasoning step must be committed to discrete textual tokens, and the resulting traces lengthen inference.  \n* Corresponding author.  \nLatent reasoning methods instead perform multistep inference entirely in the model’s continuous hidden states, promising more compact and efficient reasoning by removing the explicit trace that chain-of-thought requires (Hao et al., 2024 ; Shenet al., 2025) .  \nThis opacity raises a question of faithfulness analogous to that studied in explicit CoT, where generated rationales need not reflect the underlying computation (Turpin et al., 2023 ; Lanham et al., 2023) . Latent reasoning produces no such rationale to inspect, so the question is reframed as whether the latent reasoning steps themselves causally drive the final answer: do interventions that change or erase these steps actually change the prediction? Prior work investigates this question at converged checkpoints and reports several unfaithful behaviors of trained latent reasoners, including latent reasoning steps whose hidden states can be replaced or perturbed without changing the final prediction, and evidence that the answer can be produced by alternative paths in the network that do not pass through these steps (Zhang et al., 2025 ; Lin et al., 2025 ; Cui et al., 2026 ; Li et al., 2026) .  \nThese analyses share a methodological choice: each trained latent reasoner is treated as a fixed object and evaluated only at its final checkpoint. This leaves how the reported unfaithful behaviors form during training unexamined, a question that is complex because latent reasoning training paradigms differ substantially in mechanism. A final-checkpoint observation also cannot tell whether different paradigms reach a similar endpoint through similar trajectories or through qualitatively different ones. Treating a trained model as the endpoint of a training trajectory aligns with a broader line of work on model properties across traini","cbCaiiM7XubuJTt6","https://ap.wps.com/l/cbCaiiM7XubuJTt6","pdf",684573,1,15,"English","en",105,"# Introduction\n## Method and Interventions\n## Research Questions (RQ1–RQ3)\n## Figure 1: Trajectory Analysis","[{\"question\":\"What faithfulness question does the paper focus on for latent reasoning methods?\",\"answer\":\"Whether latent reasoning steps in hidden states causally drive the final answer, i.e., whether interventions that change or erase these steps change the prediction.\"},{\"question\":\"How does the paper measure faithfulness across training checkpoints?\",\"answer\":\"It applies two interventions: a verifiable counterfactual edit to the input (on ProsQA) and a noise-ablation activation patch to latent reasoning steps, then tracks resulting changes across saved checkpoints.\"},{\"question\":\"What overall trends does the paper find about faithfulness during training?\",\"answer\":\"Output-level unfaithfulness can look similar under counterfactual edits at convergence while following qualitatively different trajectories; at the activation level, causal contribution decays across training for both paradigms, with trajectory divergence tied to answer format.\"}]",1784199523,38,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"final-checkpoints-are-not-enough-analyzing-latent-reasoning-faithfulness-along-training-trajectories","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/final-checkpoints-are-not-enough-analyzing-latent-reasoning-faithfulness-along-training-trajectories/84936/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What faithfulness question does the paper focus on for latent reasoning methods?","Question",{"text":74,"@type":75},"Whether latent reasoning steps in hidden states causally drive the final answer, i.e., whether interventions that change or erase these steps change the prediction.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does the paper measure faithfulness across training checkpoints?",{"text":79,"@type":75},"It applies two interventions: a verifiable counterfactual edit to the input (on ProsQA) and a noise-ablation activation patch to latent reasoning steps, then tracks resulting changes across saved checkpoints.",{"name":81,"@type":72,"acceptedAnswer":82},"What overall trends does the paper find about faithfulness during training?",{"text":83,"@type":75},"Output-level unfaithfulness can look similar under counterfactual edits at convergence while following qualitatively different trajectories; at the activation level, causal contribution decays across training for both paradigms, with trajectory divergence tied to answer format.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]