[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125961-en":3,"doc-seo-125961-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},125961,137451207643,"Noah","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Verifiable evaluations of machine learning models using zkSNARKs - Preprint","Increasingly closed-source machine learning models make it difficult for end users to trust benchmark claims on accuracy, bias, or safety without re-running evaluations on black-box outputs. This work introduces a method for verifiable model evaluation using model inference through zkSNARKs. Zero-knowledge computational proofs can be packaged as evaluation attestations over public inputs, verifying that fixed private-weight models achieve stated performance or fairness metrics. The framework supports arbitrary neural networks with varying compute requirements and demonstrates results across real-world models.","Verifiable evaluations of machine learning models  \nusing zkSNARKs  \narXiv :2402 .02675v2 [ cs .LG] 22 May 2024  \nTobin South∗ MIT  \nRobert Mahari  \nMIT  \nAlexander Camuto  \nEZKL  \nChristian Paquin  \nMicrosoft Research  \nShrey Jain  \nMicrosoft Research  \nShayla Nguyen  \nUniversity of Adelaide  \nJason Morton Alex ‘Sandy’ Pentland  \nEZKL MIT  \nAbstract  \nIn a world of increasing closed-source commercial machine learning models, model evaluations from developers must be taken at face value. These benchmark results—whether over task accuracy, bias evaluations, or safety checks—are traditionally impossible to verify by a model end-user without the costly or impossible process of re-performing the benchmark on black-box model outputs. This work presents a method of verifiable model evaluation using model inference through zkSNARKs.  \nThe resulting zero-knowledge computational proofs of model outputs over datasets can be packaged into verifiable evaluation attestations showing that models with fixed private weights achieve stated performance or fairness metrics over public inputs. We present a flexible proving system that enables verifiable attestations to be performed on any standard neural network model with varying compute requirements. For the first time, we demonstrate this across a sample of real-world models and highlight key challenges and design solutions. This presents a new transparency paradigm in the verifiable evaluation of private models.  \n1 Introduction  \nModel transparency, bias checking, and result reproducibility are at odds with the creation and use of closed-sourced machine learning (ML) models. It is common practice for researchers to release model weights and architectures to provide experimental reproducibility, foster innovation, and iteration, and facilitate model auditing of biases. However, the drive towards commercialization of models by industry (and, in the case of extremely large language models, the concern over the safety of open-source models) has led to the increasing practice of keeping model weights private [1, 2] .  \nKeeping model weights private (regardless of whether the model architecture is public) limits the ability of external observers to examine the model and its performance properties. This presents two key concerns. Firstly, it can be hard to verify any claims that are made about a model’s performance, either in the form of scientific evaluation results or commercial marketing. Secondly, performance characteristics that were not intended or tested for are hard to examine. Instances of racial and gendered bias have been found post-hoc in production ML models [3] . These unwanted model performance characteristics, such as bias, come in many forms and are often not initially tested for, leaving it to the public and interested parties to determine the characteristics of a model.  \nTools such as algorithmic audits [3, 4, 5] can partially allow auditing via API, but come at a significant expense and are not possible when models are not publicly facing. Many high-risk and customerimpacting models exist without public APIs, such as resume screeners [6] or models used internally  \n∗[tsouth@mit.edu](tsouth@mit.edu)  \nPreprint. Under review.  \nFigure 1: A high level overview of the motivations and system design, which is augmented by the flexible ezkl proving system that can handle any ML model.  \nby law enforcement or governments. Even cases where models are public can benefit from verifiable evaluations given the technically challenging and expensive process of running these models.  \nWhile a model creator or provider can always state claims about a specific model’s performance on an evaluation task or set of characteristics (often in the forms of documentation or model cards such as Mitchell et al. [7]), this is insufficient to verify a model being used has a performance characteristic. To verify the claimed capability, an end user must be able to confirm that the model in use was run on the ","cbCaipbmuecaRyNy","https://ap.wps.com/l/cbCaipbmuecaRyNy","pdf",1301658,5,1,21,"English","en",105,"# Abstract\n# 1 Introduction\n## Problem: private model weights and unverifiable claims\n## Contributions and approach","[{\"question\":\"Why is it hard to verify benchmark results for closed-source ML models?\",\"answer\":\"When model weights are private and outputs are black-box, end users cannot re-run or audit the benchmark on the deployed system, making claims about accuracy, bias, or safety difficult to confirm.\"},{\"question\":\"How do zkSNARKs enable verifiable model evaluations?\",\"answer\":\"zkSNARKs are used to generate zero-knowledge computational proofs that model outputs over datasets can be packaged into verifiable evaluation attestations, without revealing private weights.\"},{\"question\":\"What does the proposed system demonstrate and support?\",\"answer\":\"The work presents a flexible proving system that can produce verifiable attestations for standard neural network models with different compute requirements, demonstrating the approach across sample real-world models while discussing design challenges and solutions.\"}]","Verifiable evaluations of machine learning models using zkSNARKs - Preprint | PDF",1785902255,53,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"verifiable-evaluations-of-machine-learning-models-using-zksnarks-preprint","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/verifiable-evaluations-of-machine-learning-models-using-zksnarks-preprint/125961/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-24","2026-08-05",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"Why is it hard to verify benchmark results for closed-source ML models?","Question",{"text":77,"@type":78},"When model weights are private and outputs are black-box, end users cannot re-run or audit the benchmark on the deployed system, making claims about accuracy, bias, or safety difficult to confirm.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"How do zkSNARKs enable verifiable model evaluations?",{"text":82,"@type":78},"zkSNARKs are used to generate zero-knowledge computational proofs that model outputs over datasets can be packaged into verifiable evaluation attestations, without revealing private weights.",{"name":84,"@type":75,"acceptedAnswer":85},"What does the proposed system demonstrate and support?",{"text":86,"@type":78},"The work presents a flexible proving system that can produce verifiable attestations for standard neural network models with different compute requirements, demonstrating the approach across sample real-world models while discussing design challenges and solutions.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":20,"slug":139},19,"General","general"]