[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86018-en":3,"doc-seo-86018-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86018,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","The First ChineseBabyLM Challenge","The paper introduces the first ChineseBabyLM challenge to be held at the 2026 NLPCC conference. It requires participants to train Chinese language models from scratch using a 100 million-token data budget, with no constraints on tokenizer, model architecture, or training epochs. Evaluation spans three tracks: Natural Language Understanding, Hanzi knowledge, and cognitive alignment. The setup uses a standardized corpus or an equivalent self-compiled alternative, plus an open evaluation pipeline and reproducible baselines.","arXiv :2607 . 10745v 1 [ cs .CL] 12 Jul 2026  \nThe First ChineseBabyLM Challenge:  \ntraining data-efficient and cognitively plausible language models for Chinese  \nSiyuan Song∗1 Zhiheng Qian∗2 Yunhao Zhang3 Linyang He4 Xiaozhe Ji5 Yingxin Lin6 Hongao Zhu7 Chongtian Shao2 Chuhan Lang8 Luan Li2 Rui Wang2 Renfen Hu5 Shaonan Wang8 Hai Hu8  \n1 Princeton University; 2 Shanghai Jiao Tong University; 3 Chinese Academy of Sciences; 4 Columbia University  \n5 Beijing Normal University; 6 Tsinghua University; 7 University of California San Diego; 8 The Hong Kong Polytechnic University  \n∗ : Equal contributions  \nCorrespondence: [chinese.babylm@gmail.com](chinese.babylm@gmail.com) ; [ss1280@princeton.edu](ss1280@princeton.edu) ; [hai.hu@polyu.edu.hk](hai.hu@polyu.edu.hk) ;  \nAbstract  \nThis paper describes the first ChineseBabyLM challenge, which will be held in the 2026 NLPCC conference. The challenge calls for researchers to train language models from scratch with 100 million Chinese tokens and evaluates the models on 3 tracks of tasks: NLU, cognitive alignment and Hanzi knowledge. There is no restriction on tokenizer, model architecture and the number of training epochs. Details of the challenge can be found in [https://chinese-babylm](https://chinese-babylm) . [github.io/](github.io/) .  \n1 Introduction  \nLarge language models have achieved strong performance across many language understanding tasks, but their success has typically depended on pretraining with very large corpora and substantial compute. This data-hungry paradigm contrasts sharply with human language acquisition: children acquire robust linguistic competence from comparatively limited, naturalistic input. The BabyLM Challenge was created to make this contrast experimentally useful by asking participants to train language models under developmentally motivated data budgets, encouraging research on sampleefficient pretraining, cognitively plausible data, and evaluation protocols that go beyond raw scale (Warstadt et al., 2023b) .  \nThe original BabyLM shared tasks focused primarily on English. Recent multilingual efforts such as BabyBabelLM extend the same motivation to a wider set of languages, showing that data-efficient language modeling should be studied under typologically and orthographically diverse conditions (Jumelet et al., 2026) . Chinese is an especially important test case. It lacks explicit word boundaries, makes extensive use of compounding, permits flexible syntactic configurations, and uses a logographic writing system in which characters carry visual, structural, and phonological regularities. These properties make Chinese language learning a poor fit for evaluation protocols designed only around alphabetic, whitespacedelimited languages, and they raise distinct questions about tokenization, character-level representation, and the amount of data needed to acquire linguistic generalizations.  \nThe ChineseBabyLM Challenge is a shared task for studying data-efficient and cognitively plausible language modeling for Chinese. Participating systems must be pretrained from randomly initialized weights, either on the official 102-millionword corpus derived from the Chinese portion of BabyBabelLM (Jumelet et al., 2026) or on a self-compiled corpus within the same word budget. This setup preserves the central BabyLM constraint while allowing participants to explore different model architectures, tokenizers, data selection strategies, and training methods for Chinese.  \nThe challenge evaluates models along three complementary dimensions. The Natural Language Understanding track combines zero-shot minimal-pair evaluation, including ZhoBLiMP (Liu et al., 2026), with supervised Chinese understanding tasks from CLUE (Xu et al., 2020) . The Hanzi track directly tests whether models acquire structural and phonological knowledge about Chinese characters from limited input. The Cognitive Modeling track uses Chinese fMRI benchmarks from MulCogBench (Zhang et al., 2025) to me","cbCaiuJX9vdeCRCm","https://ap.wps.com/l/cbCaiuJX9vdeCRCm","pdf",399026,4,1,"English","en",105,"# Introduction\n# Evaluation Tasks","[{\"question\":\"What is the core goal of the ChineseBabyLM challenge?\",\"answer\":\"To study data-efficient and cognitively plausible language modeling for Chinese under child-development-inspired data budgets and evaluation criteria beyond raw scale.\"},{\"question\":\"What training setup do participating systems have to follow?\",\"answer\":\"Systems must be pretrained from randomly initialized weights using 100 million Chinese tokens, either on the official 102-million-word corpus derived from BabyBabelLM or on a self-compiled corpus within the same token budget.\"},{\"question\":\"How are models evaluated in the challenge?\",\"answer\":\"Models are assessed across three tracks: Natural Language Understanding (including minimal-pair and CLUE-style tasks), Hanzi knowledge (character structure and phonology from limited input), and Cognitive Modeling using Chinese fMRI benchmarks.\"}]",1784207843,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"the-first-chinesebabylm-challenge","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":20},"https://docshare.wps.com/document/the-first-chinesebabylm-challenge/86018/",{"url":51,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What is the core goal of the ChineseBabyLM challenge?","Question",{"text":74,"@type":75},"To study data-efficient and cognitively plausible language modeling for Chinese under child-development-inspired data budgets and evaluation criteria beyond raw scale.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What training setup do participating systems have to follow?",{"text":79,"@type":75},"Systems must be pretrained from randomly initialized weights using 100 million Chinese tokens, either on the official 102-million-word corpus derived from BabyBabelLM or on a self-compiled corpus within the same token budget.",{"name":81,"@type":72,"acceptedAnswer":82},"How are models evaluated in the challenge?",{"text":83,"@type":75},"Models are assessed across three tracks: Natural Language Understanding (including minimal-pair and CLUE-style tasks), Hanzi knowledge (character structure and phonology from limited input), and Cognitive Modeling using Chinese fMRI benchmarks.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]