[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85111-en":3,"doc-seo-85111-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85111,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","The Illusion of Equivalency Statistical Characterization of Quantization Effects in LLMs","Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet evaluation largely relies on accuracy and perplexity. This work demonstrates that these metrics miss quantization-induced behavioral changes. It introduces correctness agreement, a decision-level overlap metric for correct predictions between a base model and quantized variants. Across 8-bit to 2-bit schemes, behavioral divergence appears under moderate quantization despite preserved task performance. Quantization is analyzed as a structural operator on attention weights, revealing non-linear low-bit breakpoints and greater sensitivity in query/key than value/output projections.","The Illusion of Equivalency: Statistical Characterization of Quantization Effects  \nin LLMs  \nBaha Rababah 12 , Cuneyt Gurcan Akcora3 , Carson K. Leung 1  \n1 Department of Computer Science, University of Manitoba, Canada  \n2 Applied Computer Education Department, Red River College Polytechnic, Canada  \n3 AI Initiative, University of Central Florida, USA  \n[rababahb@myumanitoba.ca](rababahb@myumanitoba.ca), carson.leung@umanitoba.ca, [cuneyt.akcora@ucf.edu](cuneyt.akcora@ucf.edu)  \narXiv :2607 .08734v 1 [ cs .AI] 9 Jul 2026  \nAbstract  \nPost-training quantization is widely used to deploy large language models in resource-constrained settings, yet its evaluation relies almost exclusively on accuracy and perplexity. We show that these metrics fail to capture behavioral changes induced by quantization. We introduce correctness agreement, a decision-level metric that measures overlap in correct predictions between a base model and its quantized variants, independent of absolute accuracy. Across multiple models and quantization schemes from 8-bit to 2-bit, we find that behavioral divergence emerges under moderate quantization even when task performance appears preserved. To explain this effect, we analyze quantization as a structural operator on attention weights and quantify layer-wise distortions using statistical and distributional measures. Our results reveal non-linear breakpoints at low bit-widths and show that query and key projections are consistently more sensitive than value and output projections. These findings expose an illusion of equivalence between base and quantized models and motivate behavioral evaluation beyond conventional performance metrics.  \n1 Introduction  \nRecent advances in large language models (LLMs) have significantly improved text generation and reasoning capabilities [Ziyu et al., 2023], but their growing size introduces substantial challenges for deployment in resource-constrained environments [Husom et al., 2025] . Quantization, which reduces the numerical precision of model parameters and activations, has emerged as an effective approach to lowering memory footprint, accelerating inference, and reducing energy consumption [Jin et al., 2024] . However, while quantization improves efficiency, it can also degrade performance and alter model behavior, particularly under aggressive lowbit precision settings [Dutta et al., 2024] .  \nThe fundamental objective of quantization is to preserve the functional behavior of the original model while achieving computational efficiency. In practice, behavioral preservation extends beyond task-level accuracy to include factual knowledge, reasoning robustness, and consistent stylis-  \ntic and safety-related behaviors. These properties are especially vulnerable to quantization, as small numerical perturbations in model parameters may propagate non-linearly through deep architectures and lead to unexpected behavioral changes [Dutta et al., 2024] .  \nDespite this, existing quantization evaluations [Kurtic et al., 2025; Li et al., 2025; Jin et al., 2024] rely on surface-level metrics such as accuracy, loss, and perplexity. While informative, these metrics do not fully capture whether quantization preserves behavioral consistency between the base model and its quantized variants. In particular, models with similar accuracy or perplexity may still differ substantially in their predictions on individual examples. To address this limitation, we introduce correctness agreement, a behavioral evaluation metric that measures the consistency of correct predictions between the original model and its quantized variants. Unlike aggregate performance metrics, our metric directly assesses whether quantization preserves decision-level behavior and provides a faithful signal of behavioral alignment under precision reduction.  \nBeyond behavioral evaluation, we provide a benchmark analysis of weight distribution shifts across a wide range of quantization levels (8-bit to 2-bit) . This ena","cbCaivsw81nFokVF","https://ap.wps.com/l/cbCaivsw81nFokVF","pdf",495868,1,12,"English","en",105,"# Abstract\n# Introduction\n# Background and Related Work","[{\"question\":\"How does quantization affect attention weights, and which parts are most sensitive?\",\"answer\":\"Quantization is modeled as a structural operator on attention weights, with layer-wise distortions measured via statistical and distributional analysis. Query and key projections are consistently more sensitive than value and output projections, and low-bit widths show non-linear breakpoints.\"}]",1784201167,30,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"the-illusion-of-equivalency-statistical-characterization-of-quantization-effects-in-llms","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/the-illusion-of-equivalency-statistical-characterization-of-quantization-effects-in-llms/85111/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How does quantization affect attention weights, and which parts are most sensitive?","Question",{"text":75,"@type":76},"Quantization is modeled as a structural operator on attention weights, with layer-wise distortions measured via statistical and distributional analysis. Query and key projections are consistently more sensitive than value and output projections, and low-bit widths show non-linear breakpoints.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,114,119,122,126],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":28,"slug":113},"research-report",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":45,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]