[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81599-en":3,"doc-seo-81599-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81599,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","Global Sequential Testing for Multi-Stream Auditing","Continuous auditing of machine learning systems is crucial in high-stakes domains as new data arrives, enabling timely decisions on whether performance matches intended behavior. The paper formulates multi-stream auditing as sequential hypothesis testing with k data streams and a global null asserting correctness across all streams. It compares global sequential tests and derives improved expected stopping times by merging martingales through averaging and product rules. Results show a balanced test matches Bonferroni in sparse regimes and improves under dense alternatives, validated on synthetic and real data.","Global Sequential Testing for Multi-Stream Auditing  \nBeepul Bharti 1 ∗ [bbharti1@jhu.edu](bbharti1@jhu.edu)  \nAmbar Pal2† [ambarpal@amazon.com](ambarpal@amazon.com)  \nJeremias Sulam 1 [jsulam1@jhu.edu](jsulam1@jhu.edu)  \narXiv :2602 .21479v3 [ stat .ML] 10 Jul 2026  \n1 Johns Hopkins University  \n2 Amazon Responsible AI  \nJuly 13, 2026  \nAbstract  \nAcross many risk-sensitive areas, it is critical to continuously audit machine learning systems as we receive more data to quickly determine if they are performing as designed. This auditing task can be modeled as a sequential hypothesis testing problem with k data streams and a global null hypothesis that asserts the system operates as intended across all k streams. Under the alternative, the standard global sequential test, which uses a Bonferroni correction, hasan expected stopping time of O (ln k/α) for large k and significance level α . In this work, we demonstrate that efficient sequential tests, relying on merging martingales via averaging and products rules, provide improved stopping times, and thus more powerful tests against the null. Using these results, we show that a balanced test can match the Bonferroni rate of O (ln k/α) in the sparse regime (just a few non-null streams) while achieving O (1/k ln 1/α) under dense alternatives (many non-null steams) . We validate our theory through experiments on both synthetic and real-world data.  \nContents  \n1 Introduction 2  \n2 Problem Formulation 3  \n3 A Primer in Sequential Testing 5  \n4 Powerful Global Tests via Merging 7  \n5 Experiments 10  \n6 Conclusion 12  \nA Proofs 16  \nB Useful Lemmas and Inequalities 26  \nC Algorithms 29  \nD Additional Experimental Details & Results 31  \n∗ Corresponding author.  \n†This work is not related to AP’s position at Amazon.  \n1 Introduction  \nMachine learning (ML) systems are increasingly deployed in high-stakes decision-making contexts such as healthcare [30] and criminal justice [26], where they influence critical decisions about individuals. Thus, the development of auditing procedures that continually evaluate a system’s performance has become an essential area of study [22, 5 , 1 , 17 , 34] . Notably, the importance of auditing methods has been emphasized by regulatory and governance bodies, including the U.S. Office of Science and Technology Policy [40], the European Union [6], and the United Nations [33] .  \nStatistical hypothesis testing provides a rigorous framework to perform online auditing of an ML system. For example, suppose we are tasked with auditing a newly deployed state-of-the-art medical foundation model at a hospital. The model can perform zero-shot diagnoses across various imaging modalities, like X-rays, computed tomography, magnetic resonance imaging, and more. Importantly, we receive a continuous stream of data representing the model’s predictions. To audit the model’s performance on a specific modality, such as X-rays, we can define the following null hypothesis:  \nH0 : The model is working as intended on X-rays.  \nMoreover, we can continuously test this null by leveraging sequential hypothesis tests. Unlike traditional hypothesis tests [18], which require collecting a fixed batch of data up to a pre-specified time to obtain a valid p-value, sequential tests allow data collection to stop at arbitrary, datadependent times once sufficient evidence against the null hypothesis has accumulated.  \nEnsuring a system is performing well on one stream is important, but assessing its performance across multiple streams (e.g., demographic groups, geographic locations, etc) is crucial for ensuring fairness, safety, and reliability [32, 20 , 9] . For instance, in the foundation model example, it is critical to ensure the model is performing well on multiple imaging modalities to assure safe patient outcomes. Thus, in this work we consider a multi-stream auditing problem: monitoring k data streams, the goal is to raise an alarm as soon as there is sufficient evidence that the model isn","cbCaijV1VL6faInM","https://ap.wps.com/l/cbCaijV1VL6faInM","pdf",1060274,3,1,33,"English","en",105,"# Introduction\n# Problem Formulation\n# A Primer in Sequential Testing\n# Powerful Global Tests via Merging\n# Experiments\n# Conclusion\n# Proofs","[{\"question\":\"What problem does the paper address in auditing machine learning systems?\",\"answer\":\"It studies how to continuously audit a model using k data streams by raising an alarm as soon as there is sufficient evidence the model is not operating as intended on any stream.\"},{\"question\":\"How is the auditing task modeled statistically?\",\"answer\":\"The task is modeled as sequential hypothesis testing with a global null hypothesis asserting correct behavior across all k streams, with alternative regimes ranging from sparse to dense non-null streams.\"},{\"question\":\"Why does the paper go beyond the Bonferroni-based global sequential test?\",\"answer\":\"The Bonferroni correction provides an expected stopping time of O(ln k/α) but can be unnecessarily large when many streams provide evidence against the null, motivating faster tests via other martingale-merging strategies.\"}]",1784174629,83,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"global-sequential-testing-for-multi-stream-auditing","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/global-sequential-testing-for-multi-stream-auditing/81599/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in auditing machine learning systems?","Question",{"text":75,"@type":76},"It studies how to continuously audit a model using k data streams by raising an alarm as soon as there is sufficient evidence the model is not operating as intended on any stream.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the auditing task modeled statistically?",{"text":80,"@type":76},"The task is modeled as sequential hypothesis testing with a global null hypothesis asserting correct behavior across all k streams, with alternative regimes ranging from sparse to dense non-null streams.",{"name":82,"@type":73,"acceptedAnswer":83},"Why does the paper go beyond the Bonferroni-based global sequential test?",{"text":84,"@type":76},"The Bonferroni correction provides an expected stopping time of O(ln k/α) but can be unnecessarily large when many streams provide evidence against the null, motivating faster tests via other martingale-merging strategies.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]