[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83226-en":3,"doc-seo-83226-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83226,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Hardness of Frequency-Related Queries on Compressed Strings","Compressed indexing targets string queries using space proportional to the compressed representation rather than the original length. Grammar compression augments context-free grammars to support random access and many classic string queries with O(|G|log^{O(1)} n) space and polylogarithmic query time. Frequency-related queries remain a gap, including rank (counting symbol occurrences in a substring) and symbol occurrence (testing presence). New conditional lower bounds show that efficient rank and symbol occurrence queries would imply substantially faster Boolean Matrix Multiplication, and that similar hardness extends to broader frequency tasks under the OV conjecture.","Hardness of Frequency-Related Queries on Compressed Strings  \narXiv :2607 .07366v 1 [ cs .DS] 8 Jul 2026  \nRajat De Stony Brook University, Stony Brook, NY, USA [rde@cs.stonybrook.edu](rde@cs.stonybrook.edu)  \nDominik Kempa∗ Stony Brook University, Stony Brook, NY, USA [kempa@cs.stonybrook.edu](kempa@cs.stonybrook.edu)  \nAbstract  \nCompressed indexing is a recent trend in the design of data structures that aims to support fundamental string queries in space proportional to the size of the data in compressed form. One of the most popular compression frameworks in this field is grammar compression. A length-n string T ∈ Σn (where Σ is any finite set of size up to |Σ| = |T|O(1)) represented using a context-free grammar of size |G| can be augmented to support random access queries (given any i ∈ [1 .. n], return T [i]) in O (|G|logO(1) n) space and O (logO(1) n) time. Numerous other queries, including pattern matching, longest common extension, lexicographical predecessor/successor, Burrows–Wheeler Transform, suffix array, and even suffix tree queries, can also be supported within the same bounds.  \nDespite this progress, one fundamental class of queries has remained elusive: frequencyrelated queries, such as reporting the number of occurrences of a symbol c ∈ Σ in a substring T (b.. e](the so-called rank query), or simply checking whether c occurs in T (b.. e](the symbol occurrence query) . To date, no fully general structure achieving O(|G|logO(1) n) space and O (logO(1) n) query time is known. In this work, we establish new conditional lower bounds for frequency-related problems:  \n• We prove that answering rank and symbol occurrence queries on grammar-compressed texts in polylogarithmic time using a O (|G|logO(1) n)-space structure that is constructible from the input grammar in O (|G|logO(1) n) time would imply an O (n2 logO(1) n)-time algorithm for Boolean Matrix Multiplication (BMM), where the best known algorithms achieve O (n2.371339 ) time. Our result is achieved using a more general lower bound for efficiently answering a batch of rank and symbol occurrence queries.  \n• We generalize the above result, showing that even LZ78-compressed strings cannot support efficient rank queries. Since LZ78 is provably weaker than grammar compression, this yields a stronger result: rank and symbol occurrence queries remain hard for a wider class of compressors. We further show that achieving even additive approximations of rank queries would imply faster BMM algorithms.  \n• After establishing hardness of rank and symbol occurrence queries, we consider a broader class of frequency-related queries and show that, under the popular Orthogonal Vectors (OV) conjecture, other problems, including range distinct counting and range mode frequency queries, also cannot be efficiently supported in compressed space.  \nIn summary, we develop new techniques for reasoning about computation over compressed data, and establish tight connections between compressed indexing and long-standing problems in fine-grained complexity. This sheds new light on compressed indexing by isolating a new class of frequency-related queries whose complexity hinges on known hard problems.  \n1 Introduction  \nText indexing is a classical problem that asks to preprocess a given length-n sequence (text, string) T ∈ Σn over an alphabet Σ , so that we can efficiently answer various queries on T. To date, numerous indexes using O (n) space are known, supporting a wide range of queries with query times typically ranging from  \nO(1) to O (logO(1) n) . These include classical queries such as suffix arrays/trees [MM93 , Wei73 , KLA+ 01],∗ Partially funded by the NSF CAREER Award 2337891 .  \nlongest common extension (LCE) [Wei73 , KK19], pattern matching [M¨02 , BGS17 , BBB+ 14 , MNN20], rank/select [GGV03a, BN14], lexicographical predecessor/successor [ GV00], and many others [Gus97 , Nav14 , Ohl13 , CHL07 , MBCT23] .  \nWhile classical indexes remain fundamental in many applications a","cbCaiktILADDzenp","https://ap.wps.com/l/cbCaiktILADDzenp","pdf",793585,4,1,28,"English","en",105,"# Abstract\n# Introduction\n## Text indexing and compressed indexing overview\n## Grammar compression framework and related repetitiveness measures\n## Hardness focus: frequency-related queries","[{\"question\":\"Which frequency-related string queries are studied as hard on compressed representations?\",\"answer\":\"The document focuses on rank queries that count occurrences of a symbol in a substring and symbol occurrence queries that check whether a symbol appears in a substring.\"},{\"question\":\"What does the main conditional lower bound imply about Boolean Matrix Multiplication?\",\"answer\":\"It states that an index achieving polylogarithmic-time rank and symbol occurrence queries in near-optimal compressed space would imply a faster-than-known algorithm for Boolean Matrix Multiplication.\"},{\"question\":\"Does the hardness result extend beyond grammar-compressed texts?\",\"answer\":\"Yes. The document shows that even LZ78-compressed strings cannot efficiently support rank queries, and hardness further extends to other frequency-related problems such as range distinct counting and range mode frequency under the Orthogonal Vectors conjecture.\"}]",1784186069,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"hardness-of-frequency-related-queries-on-compressed-strings","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/hardness-of-frequency-related-queries-on-compressed-strings/83226/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Which frequency-related string queries are studied as hard on compressed representations?","Question",{"text":75,"@type":76},"The document focuses on rank queries that count occurrences of a symbol in a substring and symbol occurrence queries that check whether a symbol appears in a substring.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the main conditional lower bound imply about Boolean Matrix Multiplication?",{"text":80,"@type":76},"It states that an index achieving polylogarithmic-time rank and symbol occurrence queries in near-optimal compressed space would imply a faster-than-known algorithm for Boolean Matrix Multiplication.",{"name":82,"@type":73,"acceptedAnswer":83},"Does the hardness result extend beyond grammar-compressed texts?",{"text":84,"@type":76},"Yes. The document shows that even LZ78-compressed strings cannot efficiently support rank queries, and hardness further extends to other frequency-related problems such as range distinct counting and range mode frequency under the Orthogonal Vectors conjecture.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]