[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85120-en":3,"doc-seo-85120-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85120,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Ablation, Statistical Inference, and Validation for KV-Cache Compression","Ablation-driven study compares two KV-cache compression families—TurboQuant (TQ) using randomized Walsh-Hadamard rotation with a data-oblivious Beta Lloyd-Max codebook, and SpectralQuant (SQ) using per-head eigenbases with waterfilling for bit allocation. Both optionally add a 1-bit Johnson–Lindenstrauss residual sketch on keys, values, or both. Three contributions include full ablation across QJL variants, a Kolmogorov–Smirnov based validation separating implementation noise from systematic codec effects, and regime-specific recommendations showing heavy-tailed data breaks eigenbasis methods while structured regimes favor SQ.","arXiv :2607 .09683v1 [ cs .LG] 14 Jun 2026  \nAblation, Statistical Inference, and Validation for  \nKV-Cache Compression  \nPaolo D’Alberto∗ Ashish Sirasao Elliott Delaye  \nRajeev Patwari  \nAdvanced Micro Devices, Inc.  \n{paolo. dalberto, ashish. sirasao, elliott. delaye, [rajeev. patwari](rajeev. patwari}@amd. com)[}](rajeev. patwari}@amd. com)[@amd. com](rajeev. patwari}@amd. com)  \nAbstract  \nWe present a systematic comparative study of two families of KV-cache compression schemes: TurboQuant (TQ), which applies a randomized Walsh-Hadamard rotation and a data-oblivious Beta Lloyd-Max codebook, and SpectralQuant (SQ), which calibrates a per-head eigenbasis and allocates bits via waterfilling. Both families optionally append a 1-bit Johnson-Lindenstrauss (QJL) residual sketch on the key path, the value path, or both.  \nWe make three contributions. First, we perform a full ablation across multiple QJL variants and embedding dimensions, and show that only three schemes are non-dominated: scalar quantization without rotation, WHT rotation with Beta Lloyd-Max codebook, and the latter augmented with QJL on keys.  \nSecond, we introduce a statistical validation methodology for comparing implementations: Python (oracle) and HIP/GPU (production) use different random number generators and matrix operations as explicit experimental variables. We apply the Kolmogorov-Smirnov test to separate systematic codec differences from implementation-induced variance, and identify the K-path as a direct signature of Jensen’sinequality amplifying score variance through the softmax nonlinearity.  \nThird, we compare the final schemes across all regimes and dimensions and derive regime-specific recommendations. Heavy-tailed data is catastrophic for any eigenbasis-based method: sample covariancesare destabilised by outliers, the calibrated basis is systematically misaligned, and no budget increase recovers the loss. On structured regimes, SQ wins when separate K and V eigenbases provide genuine compression; water-filling reduces to uniform allocation throughout. We also characterize the selfcalibrating nature of the effective semantic dimension deff , which adapts to the available calibration budget rather than recovering the true data rank—a property that explains both surprising wins and non-monotone scaling behaviors.  \n1 Introduction  \nTransformer inference at scale is bounded by memory bandwidth: KV-cache access dominates total memory traffic for long-context generation, and reducing cache size directly translates to latency and throughput improvements. Quantization of keys and values is the standard approach, and a growing body of work shows that aggressive quantization to 2–4 bits per element is possible without meaningful accuracy loss [5, 4, 3] .  \nA recurring challenge in evaluating KV-cache compression is the difficulty of attributing observed quality differences to specific algorithmic choices. Evaluations on real large language model (LLM) traffic conflate distributional properties, hardware effects, and algorithmic assumptions, making it hard to understand when and why a method succeeds or fails. This paper introduces a methodology for controlled evaluation of KV-cache quantization schemes: a set of six synthetic statistical regimes, each designed to isolate one structural assumption of the compression pipeline, together with a statistical framework for distinguishing systematic algorithmic differences from implementation noise. The full evaluation is released as an open-source HIP/C++ benchmark targeting AMD GPUs; any new compression scheme can be evaluated against the same regimes by implementing a single scheme interface, without modifying the data generation, metrics, or statistical validation infrastructure.  \n∗ This work was developed in collaboration with Claude (Anthropic) .  \nWe instantiate the methodology on two representative families. TurboQuant (TQ) [5] is data-oblivious: a randomized Walsh-Hadamard transform spreads each token’","cbCaic3zYUSSyBZn","https://ap.wps.com/l/cbCaic3zYUSSyBZn","pdf",2123017,2,1,15,"English","en",105,"# Abstract\n# Introduction\n## Transformer memory bottleneck and KV-cache compression\n## Controlled evaluation methodology\n# Background\n## Multi-head attention formulation","[{\"question\":\"What KV-cache compression schemes are compared in the paper?\",\"answer\":\"The paper compares TurboQuant (TQ) and SpectralQuant (SQ). TQ uses randomized Walsh-Hadamard rotation with a Beta Lloyd-Max codebook, while SQ calibrates a per-head eigenbasis and allocates bits via waterfilling.\"},{\"question\":\"How does the paper validate and compare codec implementations?\",\"answer\":\"It introduces a statistical validation approach where Python (oracle) and HIP/GPU (production) differences are treated as explicit experimental variables. The Kolmogorov–Smirnov test distinguishes systematic differences from implementation-induced variance, highlighting different behavior on key-path versus value-path residual variance.\"},{\"question\":\"Why do eigenbasis-based methods fail on heavy-tailed data?\",\"answer\":\"Heavy-tailed data destabilizes sample covariances through outliers, misaligning the calibrated basis. The paper reports that increasing the budget does not recover the lost performance, making this regime catastrophic for eigenbasis-based methods.\"}]",1784201218,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"ablation-statistical-inference-and-validation-for-kv-cache-compression","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/ablation-statistical-inference-and-validation-for-kv-cache-compression/85120/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What KV-cache compression schemes are compared in the paper?","Question",{"text":75,"@type":76},"The paper compares TurboQuant (TQ) and SpectralQuant (SQ). TQ uses randomized Walsh-Hadamard rotation with a Beta Lloyd-Max codebook, while SQ calibrates a per-head eigenbasis and allocates bits via waterfilling.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper validate and compare codec implementations?",{"text":80,"@type":76},"It introduces a statistical validation approach where Python (oracle) and HIP/GPU (production) differences are treated as explicit experimental variables. The Kolmogorov–Smirnov test distinguishes systematic differences from implementation-induced variance, highlighting different behavior on key-path versus value-path residual variance.",{"name":82,"@type":73,"acceptedAnswer":83},"Why do eigenbasis-based methods fail on heavy-tailed data?",{"text":84,"@type":76},"Heavy-tailed data destabilizes sample covariances through outliers, misaligning the calibrated basis. The paper reports that increasing the budget does not recover the lost performance, making this regime catastrophic for eigenbasis-based methods.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]