[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86224-en":3,"doc-seo-86224-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86224,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Beyond Semantic IDs Encoding Business Value Ranking into Document Identifiers for Generative Retrieval","Generative Retrieval (GR) casts retrieval as sequence-to-sequence generation by assigning each document a DocID and decoding it autoregressively, making DocID design central to retrieval quality. Existing DocID methods from discrete semantic representations face collision failures and a critical objective mismatch: they reconstruct semantics while the system is optimized for business conversion. CRID decouples DocID into semantic clustering and business-value ordinal ranking, enabling collision-free identifiers and incremental updates via intra-cluster reranking. Experiments on a 300M-item Taobao corpus show improved top-K hit rate and +1.06% GMV.","Beyond Semantic IDs: Encoding Business-Value Ranking into Document  \nIdentifiers for Generative Retrieval  \nGui Ling, Zhihong Chen, Yu Li, Tong Xiong, Kunhai Lin, Kaixuan Zhang, Yuliang Yan, Dan Ou, Haihong Tang, Bo Zheng  \nTaobao & Tmall Group of Alibaba  \n{ linggui.lg, jhon.czh, ly328242, xiongtong.xt, linkunhai.lkh zhangkaixuan.zkx, yuliang.yyl, oudan.od, piaoxue, [bozheng }@taobao.com](bozheng }@taobao.com)  \narXiv :2607 . 1 1392v 1 [ cs .IR] 13 Jul 2026  \nAbstract  \nGenerative Retrieval (GR) formulates retrieval as a sequence-to-sequence generation task, assigning each document a document identifier (DocID) and retrieving it through autoregressive decoding, making DocID design a critical factor in retrieval quality. However, existing schemes based on discrete representation learning suffer from inherent collision issues and create a mismatch between the DocID’sencoding objective and the system’s business optimization target. To address these limitations, we propose Cluster-Ranked Identifier (CRID), which decouples DocID into semantic clustering and business-value ranking, yielding collision-free identifiers that support incremental updates via intra-cluster reranking. We further introduce an analytical framework that decomposes retrieval gains into personalized preference and statistical prior generalization, revealing how semantic cluster size governs the balance between the two components. Experiments on a 300M-item Taobao e-commerce  \ncorpus show that CRID surpasses the strongest embedding-based retrieval baseline on top-K Hitrate, and delivers +1 .06% GMV in fulltraffic deployment.  \n1 Introduction  \nGenerative Retrieval (GR) has recently emerged as a promising paradigm for information retrieval tasks such as e-commerce search and content search (Chen et al., 2025 ; Pang et al., 2025) . It assigns each document a document identifier (DocID) and retrieves it through autoregressive decoding, making DocID design a fundamental determinant of retrieval performance.  \nExisting DocID schemes span diverse categories including hierarchical IDs (Rajput et al., 2023 ; Chen et al., 2025), parallel IDs (Xing et al., 2025 ; Xu et al., 2025), and other surrogate representations such as text terms (Zhang et al., 2025b) . Among them, the coarse-to-fine hierarchical structure is  \nhighly compatible with the autoregressive generation process of large language models (LLMs), making it the prevailing choice in industrial systems. Most hierarchical schemes rely on discrete representation learning, such as RQ-KMeans (Luo et al., 2025) and RQVAE (Lee et al., 2022), which discretize semantic embeddings into semantic IDs. Such methods suffer from an inherent collision problem, where multiple items may be assigned the same DocID, which becomes especially severe in large-scale corpora. To mitigate collisions, random IDs (Fu et al., 2025) and balancing-based heuristics (Zheng et al., 2024) have been proposed as post-hoc remedies (see Appendix A for a detailed discussion of related work) .  \nBeyond collisions, a more fundamental challenge arises: existing DocID schemes are constructed purely from semantic embeddings, creating a mismatch between the DocID’s encoding objective (semantic reconstruction) and the system’s optimization target (business conversion) . For example, two items in the same semantic cluster may differ by orders of magnitude in conversion rate, yet receive adjacent or even identical DocIDs under purely semantic quantization. In large-scale industrial scenarios such as Taobao e-commerce search, well-optimized embedding-based retrieval (EBR) methods naturally incorporate business signals such as click-through rate and conversion rate through feature-rich training, yet this information is entirely absent from current DocIDs. The gap is amplified by capacity–corpus tension: the candidate pool contains hundreds of millions of items, yet the GR model cannot scale arbitrarily due to latency constraints, making the information efficien","cbCaimMObOuA904G","https://ap.wps.com/l/cbCaimMObOuA904G","pdf",933517,4,1,12,"English","en",105,"# Introduction\n## Generative Retrieval and DocID Design\n## Limitations of Existing DocID Schemes\n## Proposed Approach: CRID\n## Contributions and Analytical Framework","[{\"question\":\"What problem does the paper address with existing DocID schemes in Generative Retrieval?\",\"answer\":\"It highlights two issues: collisions caused by discrete semantic quantization, and a mismatch between DocID encoding goals (semantic reconstruction) and the system’s business optimization objective (conversion).\"},{\"question\":\"How does CRID change DocID construction to improve retrieval?\",\"answer\":\"CRID decouples DocID into semantic clustering and business-value ranking, using ordinal ranks within each cluster to eliminate collisions and to represent business priors directly.\"},{\"question\":\"How are incremental updates handled when new items arrive?\",\"answer\":\"CRID supports incremental updates by reranking items within the affected semantic clusters, avoiding the need for codebook retraining.\"}]",1784209616,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"beyond-semantic-ids-encoding-business-value-ranking-into-document-identifiers-for-generative-retrieval","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/beyond-semantic-ids-encoding-business-value-ranking-into-document-identifiers-for-generative-retrieval/86224/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address with existing DocID schemes in Generative Retrieval?","Question",{"text":75,"@type":76},"It highlights two issues: collisions caused by discrete semantic quantization, and a mismatch between DocID encoding goals (semantic reconstruction) and the system’s business optimization objective (conversion).","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does CRID change DocID construction to improve retrieval?",{"text":80,"@type":76},"CRID decouples DocID into semantic clustering and business-value ranking, using ordinal ranks within each cluster to eliminate collisions and to represent business priors directly.",{"name":82,"@type":73,"acceptedAnswer":83},"How are incremental updates handled when new items arrive?",{"text":84,"@type":76},"CRID supports incremental updates by reranking items within the affected semantic clusters, avoiding the need for codebook retraining.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]