[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83390-en":3,"doc-seo-83390-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},83390,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Stop Guessing When to Stop Testing Efficient Model Evaluation with Just Enough Data","Fixed-size benchmarks limit model evaluation efficiency because they cannot match different goals such as ranking, selection, or in-development testing, leading to either excessive computation or reduced statistical reliability. The work advocates sequential testing and introduces an adaptive evaluation framework that sets principled stopping criteria aligned with needs like diminishing-returns detection and minimum detectable effect size. Experiments on the Open VLM Leaderboard show large cost savings while preserving statistical significance, with publicly available code.","Stop Guessing When to Stop Testing: Efficient Model Evaluation with Just Enough Data  \nOfir Arviv* Kristjan Greenewald Yotam Perlitz Hadar Mulian Michal Shmueli-Scheuer Leshem Choshen*  \nIBM Research  \n*Corresponding authors: [ofir.arviv@ibm.com](ofir.arviv@ibm.com) , [leshem.choshen@ibm.com](leshem.choshen@ibm.com)  \narXiv :2607 .08522v 1 [ cs .LG] 9 Jul 2026  \nAbstract  \nThe inherent rigidity of fixed-size benchmarks makes them an inefficient tool for model evaluation. Diverse evaluation objectives, including model ranking, model selection and testing throughout development, demand varying levels of statistical power. The mismatch between fixed sample sizes and these diverse needs results in either excessive computational cost or compromised reliability – a critical concern for model evaluation. To overcome these limitations, we call for adoption of sequential testing in our field. We provide an adaptive evaluation framework, that provides a principled way to navigate the trade-off between efficiency and reliability in model evaluation. Our framework combines the established statistical paradigm of sequential testing with stopping criteria tailored to common evaluation needs such as diminishing returns detection, and minimum detectable effect size. We demonstrate its ability to adaptively manage the efficiency-reliability trade-off on the Open VLM Leaderboard, including, for example, a 80% reduction in computational cost compared to fixed-size evaluation (with a 2.5-point CI width allowance) while maintaining statistical significance. Code is publicly available at [github.com/OfirArviv/adaptive-eval](github.com/OfirArviv/adaptive-eval).  \n1 Introduction  \nThe rapid advancement of large language and vision models (VLMs and LLMs) has spurred the creation of numerous benchmarks to assess their capabilities (Li and Lu, 2024 ; Chang et al., 2023) . However, processing high-resolution images, handling large contexts, comparing performance across multiple datasets, and utilizing expensive metrics like LLM-as-Judge have drastically increased evaluation costs (Zhao et al., 2024 ; Perlitz et al., 2024) . Current evaluation practices, typically employing fixed-size benchmarks, are inherently wasteful, continuing to the predetermined sample size even  \nFigure 1: Half-width of the confidence interval as a function of sample size for 206 models in the Open VLM Leaderboard benchmark, averaged over 10 random seeds. As sample size increases, the CI narrows, reducing uncertainty in model performance estimates. Two stopping strategies are illustrated: (1) stopping when the CI reaches ±2 .5, saving 80% of the evaluation cost, and (2) stopping when diminishing returns plateau, reducing cost by 44% while sacrificing only 0.132 points in precision. Our framework can detect and stop evaluation based on such rules, ensuring both statistical rigor and evaluation efficiency.  \nwhen the outcome is statistically clear. While practitioners often reduce costs by using fewer samples (Perlitz et al., 2024 ; Polo et al., 2024a ; Fogliato et al., 2024 ; Zhao et al., 2024), these heuristic approaches lack statistical guarantees.  \nCritically, fixed-size approaches, whether highcost and precise or low-cost and imprecise, fail to align the evaluation effort with the evaluation’s objective. Debugging a model may only require a inexpensive coarse approximation, while definitively determining a superior model among close contenders demands sufficient data to achieve statistical significance.  \nIn this work, we propose a statistically grounded solution: an adaptive evaluation framework based  \non sequential testing. Rather than enforcing a fixed sample size, adaptive evaluation stops when the practical and statistical needs are met. Thus, users explicitly define their needs, and the method ensures that evaluations are neither underpowered nor excessive, effectively balancing reliability and efficiency. In addition, when efficiency is prioritized, users are fully awa","cbCaiqjzUBsleYIW","https://ap.wps.com/l/cbCaiqjzUBsleYIW","pdf",363037,1,11,"English","en",105,"# Abstract\n# Introduction\n# Use Case Examples\n## Compute-Constrained Evaluation with Statistical Guarantees\n## Meaningful Change Achieved\n## Model Development: Efficient Candidate Model Selection","[{\"question\":\"为什么固定样本评测会让模型评测变得低效且不可靠？\",\"answer\":\"固定样本基准无法适配不同评测目标所需的统计能力，导致计算成本被浪费或统计结论不够可靠，从而出现效率与可靠性的错配。\"},{\"question\":\"这项工作提出了什么方法来决定何时停止评测？\",\"answer\":\"作者提出采用顺序检验（sequential testing）的自适应评测框架，用与评测需求相匹配的停止准则来在可靠性与效率之间做权衡。\"},{\"question\":\"框架在哪些具体需求下设置停止准则，并取得了什么效果？\",\"answer\":\"停止准则可用于识别收益递减、检测最小可检出效应大小等常见需求。在 Open VLM Leaderboard 上，框架能在保证统计显著性的同时显著降低计算成本，例如相对固定样本评测减少约 80%。\"}]",1784187173,28,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"stop-guessing-when-to-stop-testing-efficient-model-evaluation-with-just-enough-data","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/stop-guessing-when-to-stop-testing-efficient-model-evaluation-with-just-enough-data/83390/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么固定样本评测会让模型评测变得低效且不可靠？","Question",{"text":75,"@type":76},"固定样本基准无法适配不同评测目标所需的统计能力，导致计算成本被浪费或统计结论不够可靠，从而出现效率与可靠性的错配。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"这项工作提出了什么方法来决定何时停止评测？",{"text":80,"@type":76},"作者提出采用顺序检验（sequential testing）的自适应评测框架，用与评测需求相匹配的停止准则来在可靠性与效率之间做权衡。",{"name":82,"@type":73,"acceptedAnswer":83},"框架在哪些具体需求下设置停止准则，并取得了什么效果？",{"text":84,"@type":76},"停止准则可用于识别收益递减、检测最小可检出效应大小等常见需求。在 Open VLM Leaderboard 上，框架能在保证统计显著性的同时显著降低计算成本，例如相对固定样本评测减少约 80%。","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]