[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83406-en":3,"doc-seo-83406-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83406,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","BiSCo-LLM Lookup-Free Binary Spherical Coding for Extreme Low-Bit Large Language Model Compression","Large language models face deployment constraints from memory capacity, weight bandwidth, and checkpoint storage, making extreme low-bit compression a key requirement. Existing low-bit approaches either limit representation near 2 bits per weight or rely on vector-quantization codebooks that add lookup and extra storage accounting. BISCO-LLM introduces a codebook-free binary spherical coding framework using binarized spherical codes, a residual BSQ stage for rate–distortion control without stored codebooks, and category-wise recovery distillation for Transformer modules.","BiSCo-LLM: Lookup-Free Binary Spherical Coding for Extreme Low-Bit Large Language Model  \nCompression  \nYuantian Shao 1,2,†, Peisong Wang2,†,* , Zhilei Liu2 , Chuangyi Li2 , Yuanteng Chen2 , Pengcheng Xie3 , Yiwu Yao3 ,  \nZhihui Wei 1,* , and Jian Cheng2,*  \narXiv :2607 .08643v 1 [ cs .LG] 9 Jul 2026  \nAbstract—Large language models (LLMs) are increasingly constrained by memory capacity, weight bandwidth, and checkpoint storage during deployment. Existing low-bit compression methods mainly follow two directions. Scalar or group-wise quantization is simple and compatible with efficient low-precision kernels, but its representation capacity becomes limited when the target budget approaches 2 bits per weight. Vector-quantized weight compression provides a richer block-level representation, but usually introduces explicit codebooks, index lookup, and additional storage accounting. This paper presents BISCO-LLM, a codebook-free binary spherical coding framework for extreme low-bit LLM weight compression. The proposed pipeline is built on three components. First, local weight chunks are mapped onto a unit hypersphere and binarized into compact spherical codes, so that the main payload is a bit-packed sign stream rather than explicit VQ centroids. Second, a residual BSQ stage encodes the reconstruction error left by the base spherical codec, providing an explicit rate–distortion path without stored codebooks. Third, category-wise recovery distillation is performed after replacing each Transformer module category, reducing themismatch between local weight reconstruction and assembled model behavior. A small 8-bit protected-channel path is used as an auxiliary stabilization mechanism for sensitive channelsand is counted separately from the BSQ payload. The reported storage budget includes binary codes, neural decoders, protectedchannel payloads, LoRA adapters, and metadata. On Qwen3-8B, BISCO-LLM obtains a WikiText-2 perplexity of 10.18 compared with 9.73 for the FP16/BF16 model, and an average downstream accuracy of 68.05 compared with 69.92 over the reported seventask evaluation set. These results indicate that codebook-free spherical coding can preserve model behavior under an extreme low-bit storage budget when residual coding, sensitivity-aware protection, and recovery distillation are jointly considered.  \nIndex Terms—Large language model compression, low-bit quantization, binary spherical coding, codebook-free compression, vector quantization, residual coding, LoRA compensation.  \nI. INTRODUCTION  \nLarge language models (LLMs) have become foundation components for language understanding, reasoning, genera  \ntion, and general-purpose AI services [1]–[3] . Despite their †Yuantian Shao and Peisong Wang contributed equally to this work.  \n*Peisong Wang, Zhihui Wei, and Jian Cheng are corresponding authors.  \n1Yuantian Shao and Zhihui Wei are with Nanjing University of Science and Technology, Nanjing, China.  \n2Yuantian Shao, Peisong Wang, Zhilei Liu, Chuangyi Li, Yuanteng Chen, and Jian Cheng are with the Institute of Automation, Chinese Academy of Sciences, Beijing, China.  \n3Pengcheng Xie and Yiwu Yao are with Huawei, China.  \neffectiveness, the continuous growth of model scale has introduced substantial deployment challenges. Recent technical reports show that LLM deployment targets have moved from several-billion-parameter models to hundreds-of-billions and trillion-scale architectures. DeepSeek-V3 and DeepSeek-R1, for example, adopt a 671B-parameter MoE architecture with 37B activated parameters per token; Kimi K2 scales this regime to a 1T-parameter MoE model with 32B activated parameters; and GLM-4.5 reports 355B total parameters with 32B activated parameters [4]–[7] . These models correspond to terabyte-scale BF16/FP16 checkpoint storage before KV caches, runtime buffers, and serving-system overhead are considered. As a result, efficient inference is constrained not only by peak arithmetic throughput, but also by the amount ","cbCaivgsHBg7HsRP","https://ap.wps.com/l/cbCaivgsHBg7HsRP","pdf",645480,2,1,15,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does BiSCo-LLM address in LLM deployment?\",\"answer\":\"It targets deployment bottlenecks caused by limited memory capacity, weight bandwidth, and checkpoint storage, especially under extreme low-bit weight budgets.\"},{\"question\":\"How does BISCO-LLM avoid using explicit codebooks?\",\"answer\":\"It maps local weight chunks to a unit hypersphere and converts them into compact binary spherical codes, so the payload is a bit-packed sign stream rather than stored VQ centroids.\"},{\"question\":\"What techniques does BISCO-LLM use to improve reconstruction quality and model behavior?\",\"answer\":\"It adds a residual BSQ stage to encode reconstruction error with an explicit rate–distortion path, and applies category-wise recovery distillation after updating each Transformer module category.\"}]",1784187321,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"bisco-llm-lookup-free-binary-spherical-coding-for-extreme-low-bit-large-language-model-compression","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/bisco-llm-lookup-free-binary-spherical-coding-for-extreme-low-bit-large-language-model-compression/83406/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does BiSCo-LLM address in LLM deployment?","Question",{"text":75,"@type":76},"It targets deployment bottlenecks caused by limited memory capacity, weight bandwidth, and checkpoint storage, especially under extreme low-bit weight budgets.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does BISCO-LLM avoid using explicit codebooks?",{"text":80,"@type":76},"It maps local weight chunks to a unit hypersphere and converts them into compact binary spherical codes, so the payload is a bit-packed sign stream rather than stored VQ centroids.",{"name":82,"@type":73,"acceptedAnswer":83},"What techniques does BISCO-LLM use to improve reconstruction quality and model behavior?",{"text":84,"@type":76},"It adds a residual BSQ stage to encode reconstruction error with an explicit rate–distortion path, and applies category-wise recovery distillation after updating each Transformer module category.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]