[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-160215-en":3,"doc-seo-160215-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},160215,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Train-before-Test Harmonizes Language Model Rankings - Research Paper","Language model benchmarks often produce contradictory rankings even when they target similar skills, complicating model selection and comparison. This paper proposes train-before-test, a fair evaluation method that gives every model identical benchmark-specific fine-tuning before testing, enabling comparisons based on model potential rather than raw out-of-the-box scores. Experiments cover 24 benchmarks and 61 models, showing consistent rankings across benchmarks, restoring the link between perplexity and downstream performance, and indicating model potential is dominated by one latent factor.","arXiv :2507 .05195v2 [ cs .LG] 13 Oct 2025  \nTrain-before-Test Harmonizes Language Model Rankings  \nGuanhua Zhang*, Ricardo Dominguez-Olmedo, Moritz Hardt Max Planck Institute for Intelligent Systems, Tübingen and Tübingen AI Center  \nAbstract  \nExisting language model benchmarks provide contradictory model rankings, even for benchmarks that aim to capture similar skills. This dilemma of conflicting rankings hampers model selection, clouds model comparisons, and adds confusion to a growing ecosystem of competing models. In this paper, we take a different perspective on model comparison: instead of relying on out-of-the-box performance via direct evaluation, we compare model potential by providing each model with identical benchmark-specific fine-tuning before evaluation. We call this approach train-before-test. Our primary contribution is a comprehensive empirical evaluation of model potential across 24 benchmarks and 61 models. First, we demonstrate that model potential rankings obtained through train-before-test exhibit remarkable consistency across all benchmarks. Whereas traditional rankings demonstrate little external validity under direct evaluation, they enjoy a significant degree of external validity when applying train-beforetest: model potential rankings transfer gracefully from one benchmark to another. Second, train-before-test restores the connection between perplexity and downstream task performance, lost under direct evaluation. Remarkably, even pre-finetuning perplexity of a base model predicts post-finetuning downstream performance, suggesting that ranking consistency reflects inherent model potential rather than fine-tuning artifacts. Finally, train-before-test reduces the model-score matrix to essentially rank one, indicating that model potential is dominated by one latent factor, uncovered by train-before-test. Our work supports the recommendation to make train-before-test a default component of LLM benchmarking†.  \n1 Introduction  \nExisting language model benchmarks provide contradictory model rankings, even for benchmarks that aim to capture similar skills [47, 6, 22] . This inconsistency poses a serious challenge: how can we reliably compare, rank, and select models when different benchmarks yield conflicting information? While this ranking disagreement is often attributed to the diverse capabilities of large language models [68], it creates a conundrum in practice that muddles model development decisions [93] .  \nCurrent evaluation methodology works from direct evaluation, probing models via black-box function calls. However, large language models are trained on diverse, often proprietary data mixes that vary significantly across models [31, 28, 32] . Recent work showed that this leads to the problem of training on the test task [20]: the extent to which a model has encountered data similar  \n* Corresponding author: [guanhua.zhang@tuebingen.mpg.de](guanhua.zhang@tuebingen.mpg.de)  \n†Code is available at [https://github. com/socialfoundations/lm-harmony](https://github. com/socialfoundations/lm-harmony).  \nto the test task during training confounds model comparisons, rankings, and scaling laws [40] . Put simply, an otherwise inferior model may have simply prepared better for a specific task.  \nIn this paper, we take a fresh perspective on evaluation methodology: in contrast with direct evaluation, we compare model potential by giving each model the same task-specific fine-tuning. We call this approach train-before-test. Its goal is to achieve valid model comparisons by ensuring that all models receive equal preparation for the test.  \nWe envision train-before-test as a tool for regret-free model selection for downstream applications. Increasingly, practitioners select one from many available models with the goal of adapting fora specific task. Under direct evaluation the best model to begin with may no longer be the best model after task-specific preparation. In contrast, we show that train-before-task y","cbCaicC0uRMZYLN1","https://ap.wps.com/l/cbCaicC0uRMZYLN1","pdf",927797,1,33,"English","en",105,"# Abstract\n# 1 Introduction\n## Evaluation methodology and the ranking problem\n## Goal of train-before-test\n# 1.1 Our Contributions\n## Ranking consistency across benchmarks\n## Alignment of perplexity with downstream tasks","[{\"question\":\"Why do existing language model benchmarks produce conflicting rankings?\",\"answer\":\"Benchmarks can yield contradictory rankings because direct evaluation reflects differences in training data exposure across models, which may include data similar to the test task.\"},{\"question\":\"What is the train-before-test approach?\",\"answer\":\"Train-before-test compares model potential by applying identical benchmark-specific fine-tuning to each model before evaluation, ensuring equal preparation for the test.\"},{\"question\":\"How does train-before-test affect the relationship between perplexity and downstream task performance?\",\"answer\":\"It restores the connection: perplexity rankings after fine-tuning align with downstream task rankings, and even pre-finetuning perplexity of a base model predicts post-finetuning downstream performance.\"}]","Train-before-Test Harmonizes Language Model Rankings - Research Paper | PDF",1788052022,83,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"train-before-test-harmonizes-language-model-rankings-research-paper","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/train-before-test-harmonizes-language-model-rankings-research-paper/160215/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-30",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do existing language model benchmarks produce conflicting rankings?","Question",{"text":75,"@type":76},"Benchmarks can yield contradictory rankings because direct evaluation reflects differences in training data exposure across models, which may include data similar to the test task.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the train-before-test approach?",{"text":80,"@type":76},"Train-before-test compares model potential by applying identical benchmark-specific fine-tuning to each model before evaluation, ensuring equal preparation for the test.",{"name":82,"@type":73,"acceptedAnswer":83},"How does train-before-test affect the relationship between perplexity and downstream task performance?",{"text":84,"@type":76},"It restores the connection: perplexity rankings after fine-tuning align with downstream task rankings, and even pre-finetuning perplexity of a base model predicts post-finetuning downstream performance.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]