[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82892-en":3,"doc-seo-82892-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82892,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","CollabEval Statistically Efficient Collaborative Model Evaluation via Matrix Completion","CollabEval introduces a Collaborative Evaluation method that improves the statistical efficiency of generative AI model evaluation by exploiting dependencies across historical runs on the same tasks. Evaluation is framed as matrix completion over a models-by-prompts score matrix, where only a small fraction of entries for target models are labeled. Low-rank reconstruction provides control variates for cross-prediction-powered inference, yielding unbiased estimates with asymptotically valid confidence intervals. Experiments across diverse datasets and sparsity show reduced confidence-interval size and lower mean squared error under equal annotation budgets.","arXiv :2607 .05046v 1 [ cs .LG] 6 Jul 2026  \nCollabEval: Statistically Efficient Collaborative Model Evaluation via Matrix Completion  \nAdam Fisch1 , Daniel Deutsch1 , Joshua Maynez1 , Alekh Agarwal2 , Jonathan Berant2 , William Cohen1 , Amir Globerson2 and Jacob Eisenstein1  \n1 Google DeepMind, 2 Google Research  \nEvaluating generative AI models is a routine, but resource-intensive, process that is conducted over andover again during the course of model development. In this work, we propose Collaborative Evaluation (CollabEval), a simple, effective, and principled method for exploiting dependencies between historical runs of different models on the same tasks to improve statistical efficiency. Specifically, our approach treats model evaluation as a matrix completion problem over an 􀀢 × 􀀣 matrix of evaluation scores, where 􀀢 is the total number of models and 􀀣 is the total number of evaluation prompts. We assume that a subset of these 􀀢 models are targeted for evaluation. For these target models only a small fraction, 􀀾, of prompts has been annotated with evaluation scores. Leveraging recent results in prediction-powered inference, we build a low-rank approximation of the score matrix, and use the reconstructed values as control variates in a manner that guarantees unbiased estimates of the true evaluation metric mean, in addition to statistically valid confidence intervals. Empirically, across a wide range of datasets, models, and sparsity levels 􀀾, we find that CollabEval substantially reduces the mean confidence interval size, and the mean squared error of the point estimate, compared to baseline methods at the same annotation budget.  \n1. Introduction  \nEvaluating large-scale generative AI models typically involves gathering prompts from an input distribution, generating model outputs, and scoring all outputs with a reliable rater (e.g., an LLM-as-a-judge or human annotator) . Though straightforward in principle, obtaining reliable results requires gathering a non-trivial number of data points, which can make evaluation costly. While recent work has explored the viability of using more efficient automatic metrics and autoraters as proxies for expensive human annotations, even the inference step for most large AI models carries substantial overhead (e.g., especially when long tool-call chains are involved), motivating methods that bypass generation entirely.  \nEvaluation costs compound when many evaluations are conducted simultaneously or continuously overlong periods of time—for instance, when comparing fine-tuned model variants, or when monitoring for regressions against new checkpoints. The systems under comparison often share architectures, training data, or system prompts, so their performances are correlated. Standard methods, however, usually treat each of these evaluations as a distinct estimation problem, and fail to leverage this correlational structure. This leads to a suboptimal use of the available evaluation budget.  \nIn this work, we propose Collaborative Evaluation (CollabEval): a lightweight approach that exploits dependencies between model output scores to increase the effective sample size of an evaluation, thereby reducing the set of evaluated prompts while preserving accuracy and confidence. Unlike methods that reduce labeling costs via cheaper autoraters but still incur generation costs [6; 9; 3], skipped prompts in our setting incur nearly zero cost—requiring neither inference nor scoring. Our method has two parts. First, we reframe evaluation with partially observed scores as a collaborative filtering task: we represent the full set of models × prompts → scores as an 􀀢 × 􀀣 matrix 􀀨, where only a fraction 􀀾 of entries are observed for the target models we want to evaluate (a subset of the 􀀢 total models; other non-target models may be historical) . We use matrix completion to try to impute the unobserved scores, transferring knowledge across related evaluations. Then we use these  \nFigure 1 | An illus","cbCaib08jjo6Ppqo","https://ap.wps.com/l/cbCaib08jjo6Ppqo","pdf",1588999,2,1,33,"English","en",105,"# Introduction\n## Problem motivation\n## Proposed method: CollabEval\n## Matrix completion and control variates\n## Theoretical validity and extensions","[{\"question\":\"What problem does CollabEval address in generative AI model evaluation?\",\"answer\":\"It targets the high cost and inefficiency of repeatedly estimating evaluation metrics for many model runs, especially when evaluations share correlated systems and prompts.\"},{\"question\":\"How does CollabEval model evaluation data for its efficiency gains?\",\"answer\":\"It represents model scores as a matrix of models × prompts, with dense historical scores from anchor models and sparse labeled entries for target models, then uses matrix completion to impute missing scores.\"},{\"question\":\"How does CollabEval guarantee reliable estimates and confidence intervals?\",\"answer\":\"Imputed values are used as control variates within a prediction-powered inference framework, ensuring unbiased estimation and asymptotically valid confidence intervals under the control variate theory.\"}]",1784183736,83,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"collabeval-statistically-efficient-collaborative-model-evaluation-via-matrix-completion","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/collabeval-statistically-efficient-collaborative-model-evaluation-via-matrix-completion/82892/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does CollabEval address in generative AI model evaluation?","Question",{"text":75,"@type":76},"It targets the high cost and inefficiency of repeatedly estimating evaluation metrics for many model runs, especially when evaluations share correlated systems and prompts.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does CollabEval model evaluation data for its efficiency gains?",{"text":80,"@type":76},"It represents model scores as a matrix of models × prompts, with dense historical scores from anchor models and sparse labeled entries for target models, then uses matrix completion to impute missing scores.",{"name":82,"@type":73,"acceptedAnswer":83},"How does CollabEval guarantee reliable estimates and confidence intervals?",{"text":84,"@type":76},"Imputed values are used as control variates within a prediction-powered inference framework, ensuring unbiased estimation and asymptotically valid confidence intervals under the control variate theory.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]