[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85602-en":3,"doc-seo-85602-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85602,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","NITP: Next Implicit Token Prediction for LLM Pre-training","Standard next-token prediction (NTP) trains language models with sparse one-hot supervision in the output logit space, leaving the latent representation geometry insufficiently constrained. This under-constraining can drive hidden states toward degenerate, anisotropic configurations that harm generalization. Next Implicit Token Prediction (NITP) augments discrete prediction with dense continuous supervision in representation space, using implicit semantic targets from shallow-layer outputs. The method consistently improves downstream results across dense and MoE scales with minimal extra training compute and no added inference cost.","NITP: Next Implicit Token Prediction for LLM Pre-training  \nXiangdong Zhang 1 2 Debing Zhang 2 Shaofeng Zhang 3 Xiaohan Qin 1 2 Yu Cheng 4 Junchi Yan 1  \narXiv :2605 .24956v 3 [ cs .CL] 12 Jul 2026  \nAbstract  \nStandard Next-Token Prediction (NTP) supervises language models solely through discrete labels in the output logit space. We argue that this sparse, one-hot supervision leaves the latent representation space under-constrained, allowing hidden states to drift into degenerate and anisotropic configurations that limit generalization. To address this issue, we propose Next Implicit Token Prediction (NITP), which augments discrete prediction with dense, continuous supervision directly in the representation space. NITP requires the model to predict the implicit semantic content of the next token, using shallow-layer representations from the same model as stable self-supervised targets. Theoretically, we show that NITP regularizes the optimization landscape by mitigating under-constrained degrees of freedom and enforcing a compact, structured representation geometry. Empirically, across dense and MoE models ranging from 0.5B to 9B parameters, NITP consistently improves downstream performance with negligible computational overhead. Notably, on the 9B MoE model, NITP achieves a 5.7% absolute improvement on MMLU-Pro, along with gains of 6.4% on C3 and 4.3% on CommonsenseQA, with ∼2% additional training FLOPs and no additional inference cost. Our implementation is available at [https://github.com/aHapBean/NITP](https://github.com/aHapBean/NITP).  \n1. Introduction  \nLarge language models (LLMs) have achieved remarkable success through large-scale pre-training with the NextToken Prediction (NTP) (Liu et al., 2024b ; Yang et al., 2025 ;  \n1 School of AI, Shanghai Jiao Tong University 2Dots Studio, Xiaohongshu Inc. 3University of Science and Technology of China 4The Chinese University of Hong Kong. Correspondence to: Junchi Yan \u003C[yanjunchi@sjtu.edu.cn](yanjunchi@sjtu.edu.cn) >.  \nProceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \n(a) Effective rank (b) Cosine similarity  \n(c) Avg performance (MoE) (d) Avg performance (Dense)  \nFigure 1. Top: Representation geometry of the last hidden states under NTP and NITP. Bottom: Average downstream performance of 9B MoE and 2B dense models (details in Appendix B) .  \nTeam et al., 2025 ; He & Su, 2025 ; Chen et al., 2024) . By maximizing the likelihood of the next token over massive corpora, this paradigm learns general-purpose representations that support a wide range of downstream tasks (Guo et al., 2025) . Although subsequent fine-tuning stages further adapt these models, their effectiveness critically depends on the representations learned during base pre-training (Yue et al., 2025) . As a result, the pre-training objective plays a central role in determining the capability of LLMs.  \nDespite its empirical success, standard NTP provides supervision only through discrete one-hot targets in the output token space. While gradients propagate to hidden states via the output projection, the objective primarily constrains representations along the target logit direction, leaving many degrees of freedom in the latent space that remain weakly constrained. This raises a fundamental question: Does nexttoken prediction alone sufficiently supervise the geometry of hidden representations?  \nPrior work has identified a phenomenon known as representation degeneration, where likelihood-based training drives learned embeddings to collapse into a narrow, anisotropic cone (Ethayarajh, 2019 ; Wang et al., 2020 ; Barbero et al., 2024 ; Gao et al., 2019) . Such geometric collapse limits the expressive capacity of representations and has been linked to  \ndegraded generalization on downstream tasks (Ethayarajh, 2019 ; Zhao et al., 2024) . To examine this issue, we track the geometric evolution of last hidden states throughout ","cbCaisOdXT0hDAQ3","https://ap.wps.com/l/cbCaisOdXT0hDAQ3","pdf",1785820,3,1,21,"English","en",105,"# Abstract\n# 1. Introduction\n## Motivation: under-constrained geometry in NTP\n## Related work: representation degeneration\n## Proposed method: Next Implicit Token Prediction (NITP)","[{\"question\":\"What problem does the paper identify with standard next-token prediction (NTP)?\",\"answer\":\"The paper argues that NTP’s one-hot supervision under-constrains the latent representation space, allowing hidden states to drift toward degenerate, anisotropic geometries that limit generalization.\"},{\"question\":\"How does NITP differ from NTP in training signals?\",\"answer\":\"NITP adds dense continuous supervision directly in the representation space by predicting an implicit semantic representation of the next token, derived from the model’s shallow-layer outputs.\"},{\"question\":\"What evidence supports NITP’s effectiveness?\",\"answer\":\"The paper reports that across dense and MoE models (0.5B–9B), NITP consistently improves downstream performance, including a 5.7% absolute improvement on MMLU-Pro for a 9B MoE model, with negligible inference overhead.\"}]",1784204862,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"nitp-next-implicit-token-prediction-for-llm-pre-training","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/nitp-next-implicit-token-prediction-for-llm-pre-training/85602/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper identify with standard next-token prediction (NTP)?","Question",{"text":75,"@type":76},"The paper argues that NTP’s one-hot supervision under-constrains the latent representation space, allowing hidden states to drift toward degenerate, anisotropic geometries that limit generalization.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does NITP differ from NTP in training signals?",{"text":80,"@type":76},"NITP adds dense continuous supervision directly in the representation space by predicting an implicit semantic representation of the next token, derived from the model’s shallow-layer outputs.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence supports NITP’s effectiveness?",{"text":84,"@type":76},"The paper reports that across dense and MoE models (0.5B–9B), NITP consistently improves downstream performance, including a 5.7% absolute improvement on MMLU-Pro for a 9B MoE model, with negligible inference overhead.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]