[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85074-en":3,"doc-seo-85074-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85074,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","DaV-Gen End-to-End Generative Retrieval via Draft and Verify","Mainstream industrial information retrieval relies on multi-stage cascade architectures that balance effectiveness and efficiency with a coarse-to-fine retrieval–ranking pipeline, yet inconsistent optimization objectives across stages cause error propagation and limit final quality. End-to-end generative retrieval unifies the pipeline, but autoregressive decoding severely hurts online serving latency and controllability. DaV-Gen introduces a Draft-and-Verify mechanism that refactors search and recommendation into efficient vector-based drafting plus fused verification, enabling speed and precision in a single end-to-end model.","DaV-Gen: End-to-End Generative Retrieval via Draft-and-Verify  \nMeng Zhao , Chunmei Liu , Qinyong Wang HUJING Digital Media & Entertainment Group {[chuming.zm](chuming.zm), meijiang.lcm, [wangqinyong.wqy](wangqinyong.wqy}@alibaba-inc.com)[}](wangqinyong.wqy}@alibaba-inc.com)[@alibaba-inc.com](wangqinyong.wqy}@alibaba-inc.com)  \narXiv :2607 .08365v 1 [ cs .IR] 9 Jul 2026  \nAbstract  \nMainstream industrial information retrieval systems (e.g., search and recommendation) are usually built upon Multi-Stage Cascade Architectures (MCAs), which balance effectiveness and efficiency through a coarse-to-fine “retrieval-ranking”  \npipeline. However, the optimization objectives across different stages are substantially inconsistent, propagating or even amplifying the earlystage errors that ultimately degrade the quality of final results. While emerging end-to-end generative models offer a potential solution by unifying the pipeline, their online serving performance is severely hindered by the auto-regressive process inherited from the standard decoder-only structure.  \nTo bridge this gap, we introduce DaV-Gen, a novel unified solution designed to fundamentally refactor the paradigm for both search and recommendation via a “Draft-and-Verify” mechanism. Inspired by the process used by speculative decoding, our framework redesigns the generation task into two synergistic operations within a single model. During training, the model is concurrently optimized for both candidate drafting and fine-grained verification. This is achieved by a composite loss function that jointly trains the model on two distinct but related objectives: 1) a contrastive loss that structures the embedding space for efficient drafting, and 2) a fusion loss that combines generative likelihood with vector similarity to produce a superior verification score. This integrated training strategy equips the model with dual capabilities. At inference time, it first performs highly efficient vectorbased drafting to generate a candidate set, and then verifies these candidates using the more powerful fused scoring function, thereby achieving both the speed of sparse drafting and the precision of advanced generative models within a unified, end-toend architecture.  \n1 Introduction  \nMost large-scale information retrieval systems, including Search Engines and Recommender Systems, rely on a Multi-  \nStage Cascade Architecture (MCA) [Covington et al., 2016; Zhou et al., 2018] to balance efficacy and efficiency. However, this distinct “retrieve-and-rank” pipeline suffers from objective inconsistency, where misalignment between stagespecific goals leads to error propagation—early misses by the retriever are irreversible, fundamentally limiting final result quality [Zhang et al., 2025; Hron et al., 2021] . To unify this pipeline, recent end-to-end models [Deng et al., 2025] represent items as discrete “Semantic IDs” [Rajput et al., 2023], transforming retrieval into a sequence-to-sequence task. Despite their promise, standard Generative Information Retrieval (GenIR) models are severely hindered by their autoregressive nature, which introduces prohibitive inference latency and lacks deterministic control over recommendation list length.  \nTo bridge the gap between the efficiency of cascades and the unified expressiveness of generative models, we introduce DaV-Gen, a novel framework that refactors the paradigm via a Draft-and-Verify mechanism. Instead of relying onslow token-by-token generation, DaV-Gen reformulates the task into two synergistic operations within a single, end-toend differentiable architecture (Figure 1) . We systematically tackle the inherent trade-offs in current systems by addressing three fundamental design questions:  \n1) How to balance representational efficiency and expressiveness? (Figure 1a) Standard sparse tokens facilitate generation but often lose fine-grained semantics, while dense vectors excel at retrieval but lack structural hierarchy. Our design resolves th","cbCaicDDzj5z90xH","https://ap.wps.com/l/cbCaicDDzj5z90xH","pdf",507294,4,1,"English","en",105,"# Introduction\n## Hybrid Sparse-Dense Representation\n## Unified Collaborative Training\n## Parallel Inference Pipeline\n# Problem: Latency Bottleneck in Generative Retrieval","[{\"question\":\"Why do multi-stage cascade architectures harm final retrieval quality?\",\"answer\":\"Because stages optimize different objectives, misalignment leads to error propagation. Early misses by the retriever are irreversible and constrain final results.\"},{\"question\":\"How does DaV-Gen reduce latency compared with autoregressive generative retrieval?\",\"answer\":\"It replaces slow token-by-token decoding with a Draft-and-Verify paradigm: first drafts candidates using efficient vector-based search, then verifies them using a fused scoring function.\"},{\"question\":\"How are drafting and verification objectives aligned during training?\",\"answer\":\"DaV-Gen uses a collaborative training strategy with a composite loss: contrastive loss for efficient drafting embedding, and a fusion loss that combines generative likelihood with vector similarity to produce verification scores.\"}]",1784200876,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"dav-gen-end-to-end-generative-retrieval-via-draft-and-verify","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":20},"https://docshare.wps.com/document/dav-gen-end-to-end-generative-retrieval-via-draft-and-verify/85074/",{"url":51,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why do multi-stage cascade architectures harm final retrieval quality?","Question",{"text":74,"@type":75},"Because stages optimize different objectives, misalignment leads to error propagation. Early misses by the retriever are irreversible and constrain final results.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does DaV-Gen reduce latency compared with autoregressive generative retrieval?",{"text":79,"@type":75},"It replaces slow token-by-token decoding with a Draft-and-Verify paradigm: first drafts candidates using efficient vector-based search, then verifies them using a fused scoring function.",{"name":81,"@type":72,"acceptedAnswer":82},"How are drafting and verification objectives aligned during training?",{"text":83,"@type":75},"DaV-Gen uses a collaborative training strategy with a composite loss: contrastive loss for efficient drafting embedding, and a fusion loss that combines generative likelihood with vector similarity to produce verification scores.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]