[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84579-en":3,"doc-seo-84579-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84579,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Auditing Empirical Comparisons in Quantum Software","Empirical quantum-software studies often present “A beats B” claims, but the ordering between baselines depends on benchmark scope, circuit construction, compilation, sampling, backend and noise assumptions, optimizer stochasticity, and resource budgets. Existing testing, benchmarking, and reproducibility assess programs or experiments yet do not directly audit whether the reported comparison is supported by matched evidence in the paper and artifact. CLAIMSTAB-QC is a source-bounded auditing framework that locks comparison design before outcomes and classifies scalar-direction results. Evaluation on 455 claims finds a materialization gap.","Auditing Empirical Comparisons in Quantum  \nSoftware  \nBoshuai Ye∗ , Peng Liang†, Maryam Tavassoli Sabzevari∗ , and Arif Ali Khan∗  \n∗ University of Oulu, Finland  \n[Email:](Email: {boshuai.ye)[ {](Email: {boshuai.ye)[boshuai.ye](Email: {boshuai.ye), maryam.tavassolisabzevari, [arif.khan](arif.khan}@oulu.fi)[}](arif.khan}@oulu.fi)[@oulu.fi](arif.khan}@oulu.fi)  \n†Wuhan University, China  \nEmail: [liangp@whu.edu.cn](liangp@whu.edu.cn)  \narXiv :2607 .005 16v 1 [ cs . SE] 1 Jul 2026  \nAbstract—Empirical quantum-software papers often report comparative claims: one compiler produces smaller circuits, one optimizer is more reliable, or one ansatz achieves better solution quality. These “A beats B” conclusions are not properties of a tool alone; they can change with benchmark scope, circuit construction, compilation, sampling, backend and noise assumptions, optimizer choices, and resource budgets. Existing testing, benchmarking, and reproducibility methods help assess programs, tools, executions, platforms, and experiments, but they do not directly audit whether the reported comparison itself is supported by the evidence exposed in the paper and artifact.  \nWe present CLAIMSTAB-QC, a source-bounded framework for auditing empirical comparisons in quantum software. Given a reported comparison, the framework records the compared baselines, metric, relation, and admissible evidence; locks the comparison design before outcomes are computed; and reports either a scoped relation outcome or an explicit evidence boundary. For strict scalar-directional comparisons, the reported direction is classified as Sustained, Unresolved, or Reversed within the locked audit scope.  \nWe evaluate the framework on 455 comparative claims from 119 quantum-software papers. The central finding is a materialization gap: 175 claims can be represented for audit planning, 79 become scalar-directional planning records, and 53 yield lockable auditor diagnostic designs, but only 8 expose enough matched evidence to audit the original comparison without proxy reconstruction. These 8 records yield 2 Sustained, 4 Unresolved, and 2 Reversed outcomes. Controlled diagnostics over 24 benchmark-relevant quantum-workload comparisons further show that simpler checks can preserve apparent directions whose support weakens under locked audit designs. The results suggest that empirical quantumsoftware sources should report performance orderings together with the matched evidence needed to audit the scope of those orderings.  \nI. INTRODUCTION  \nEmpirical quantum-software papers often report comparative claims: one compiler produces smaller circuits, one optimizer converges faster, one ansatz gives better solution quality, or one backend configuration is more efficient. These claims usually have the form “A beats B” under a stated metric. In quantum software engineering (QSE), however, such orderings are not properties of a technique alone. They are produced by a stack of frameworks, circuit libraries, compilers, optimizers, simulators, cloud platforms, and hardware backends [1]–[9] . The ordering between two baselines can change with benchmark scope, circuit construction, compilation, sampling, backend/noise assumptions, optimizer stochasticity, and resource budget [10],  \n[11] . This makes comparative evidence especially important in QSE: empirical quantum-software studies need to show not only who wins, but also what evidence supports that ordering and where the support ends.  \nRerunning an experiment and auditing a comparison answer different questions. A rerun checks whether one reported result can be reproduced under one set of conditions; a comparison audit checks whether the stated “A beats B” relation is supported by the evidence exposed in the source paper and artifact. Existing methods provide important pieces of this evidence: testing checks programs and platforms [12], benchmarking measures tools and backends [13], reproducibility infrastructure reruns experiments [14], an","cbCaiojjeBSJ9T2X","https://ap.wps.com/l/cbCaiojjeBSJ9T2X","pdf",391301,1,12,"English","en",105,"# Abstract\n# I. Introduction\n## Motivation for comparison auditing\n## Definition of the audit unit and challenges","[{\"question\":\"What problem does CLAIMSTAB-QC address in quantum software research?\",\"answer\":\"It audits whether a reported “A beats B” comparative claim is actually supported by matched evidence exposed in the paper and artifact, rather than relying on reruns or partial components of reproducibility.\"},{\"question\":\"How does CLAIMSTAB-QC perform the auditing of a comparison?\",\"answer\":\"Given a reported comparison, it records the baselines, metric, relation, and admissible evidence, locks the comparison design before computing outcomes, and then reports either a scoped relation result or an explicit evidence boundary.\"},{\"question\":\"What did the evaluation on 455 comparative claims reveal?\",\"answer\":\"A materialization gap occurred: many claims could be planned or partially represented, but only a small fraction exposed enough matched evidence to audit the original comparison without proxy reconstruction, producing a mix of Sustained, Unresolved, and Reversed outcomes.\"}]",1784196901,30,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"auditing-empirical-comparisons-in-quantum-software","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/auditing-empirical-comparisons-in-quantum-software/84579/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does CLAIMSTAB-QC address in quantum software research?","Question",{"text":74,"@type":75},"It audits whether a reported “A beats B” comparative claim is actually supported by matched evidence exposed in the paper and artifact, rather than relying on reruns or partial components of reproducibility.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does CLAIMSTAB-QC perform the auditing of a comparison?",{"text":79,"@type":75},"Given a reported comparison, it records the baselines, metric, relation, and admissible evidence, locks the comparison design before computing outcomes, and then reports either a scoped relation result or an explicit evidence boundary.",{"name":81,"@type":72,"acceptedAnswer":82},"What did the evaluation on 455 comparative claims reveal?",{"text":83,"@type":75},"A materialization gap occurred: many claims could be planned or partially represented, but only a small fraction exposed enough matched evidence to audit the original comparison without proxy reconstruction, producing a mix of Sustained, Unresolved, and Reversed outcomes.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":28,"slug":120},"research-report",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]