[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85053-en":3,"doc-seo-85053-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85053,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","BACH Bayesian Admixture of Contrastive Heads for Multi-Interest Two-Tower Retrieval","Two-tower retrievers compress each user into a single embedding, limiting performance when user interests are diverse. Multi-interest models use k heads and score items by the maximum inner product, but hard-routing training can cause routing collapse and provides no per-user estimate of interest importance. BACH (Bayesian Admixture of Contrastive Heads) models heads as a per-user mixture fit via variational inference, enabling soft training, per-user interest weights reused at serving, and an efficient codebook variant. Experiments on MovieLens-20M, Taobao, and Netflix improve top-of-ranking retrieval across head counts.","BACH: A Bayesian Admixture of Contrastive Heads for Multi-Interest Two-Tower Retrieval  \nQuoc Phong Nguyen  \nAmazon  \n[qphong@amazon.com](qphong@amazon.com)  \nPaul Albert  \nAmazon  \n[albrtpa@amazon.com](albrtpa@amazon.com)  \nLong Vuong  \nAmazon  \n[longvt@amazon.com](longvt@amazon.com)  \nVuong Le  \nAmazon  \n[levuong@amazon.com](levuong@amazon.com)  \nJulien Monteil  \nAmazon  \n[jul@amazon.com](jul@amazon.com)  \narXiv :2607 .08 107v 1 [ cs .IR] 9 Jul 2026  \nAbstract  \nTwo-tower retrievers compress each user into a single embedding, limiting their ability to serve diverse interests. Multi-interest models give each user several heads scored by a maximum inner product, but their hard-routing training underutilizes heads (routing collapse) and gives no per-user estimate of how much each interest matters for serving. We present BACH (Bayesian Admixture of Contrastive Heads), which casts multi-interest two-tower retrieval as a per-user mixture over the heads, fit by variational inference. The soft mixture trains every head (mitigating collapse), produces a per-user weighting of the interests that is reused at serving, and admits a shared global-codebook variant with precomputable retrieval. On three large-scale benchmarks, MovieLens-20M, Taobao, and Netflix, BACH improvestop-of-ranking retrieval over hard-routing multi-interest and single-vector baselinesat every head count; we further find that scoring every candidate by its best head, consistent with serving, outperforms the usual target-routed training, and that BACH improves further still.  \n1 Introduction  \nTwo-tower retrievers have become the workhorse of large-scale candidate generation in recommender systems and dense information retrieval [6, 14, 26] . The standard formulation encodes a user (or query) and a candidate item with two independent neural networks into a shared d-dimensional embedding space and scores their compatibility by an inner product. Training is typically a sampled-softmax classification of the observed positive item against in-batch [26] or mixed-batch [24] negatives. At serving time, item embeddings are precomputed and indexed for approximate nearest-neighbor (ANN) search, yielding millisecond-scale retrieval over million-item catalogs.  \nA single dense user embedding, however, struggles to represent users whose interests span multiple, weakly related topics: items that satisfy any of the user’s interests must all sit close to a single point in the embedding space, and the resulting retrieval slates tend to over-represent the dominant interest while under-serving the rest. Multi-interest retrievers [5, 17, 19–21] relax this bottleneck by emitting k user embeddings, one per interest, and scoring an item by the maximum inner product across the k heads. This is structurally identical to the multi-vector / late-interaction paradigm in neural information retrieval [15] and yields a strictly more expressive class of user representations. One way to train multi-interest retrievers is hard-routing sampled softmax: only the head with the largest inner product to the observed item receives gradient [5, 17] . This routing has two shortcomings. First, it is prone to a winner-take-all dynamic in which heads that are slightly worse at initialization are selected less frequently, receive fewer gradient updates, and progressively collapse into copies of  \nPreprint.  \nFigure 1: Three formulations of two-tower retrieval. (a) The standard model encodes the user and the item as a single embedding each and scores by inner product. (b) The multi-head extension (Section 3) outputs k user embeddings; the score is the maximum inner product, and gradient flows only through the argmax head. (c) Our variational approach (Section 4) treats the k heads as components of a per-user mixture of softmaxes: each head u(r) defines a soft cluster with mixture weight πu(r) shown above it. The mixture weights are fitted by variational inference and reused at serving time to weight the per-head r","cbCaijsiLgYzOGaF","https://ap.wps.com/l/cbCaijsiLgYzOGaF","pdf",493645,3,1,16,"English","en",105,"# Abstract\n# 1 Introduction\n# 2 Background: Two-Tower","[{\"question\":\"Why do standard two-tower retrievers struggle with multi-interest users?\",\"answer\":\"They represent a user with a single embedding, forcing items satisfying different interests to cluster near one point. This skews retrieval toward the dominant interest and under-serves others.\"},{\"question\":\"What problem does BACH address compared with hard-routing multi-interest models?\",\"answer\":\"Hard-routing training can lead to routing collapse, where heads selected early dominate and others receive fewer updates. It also lacks a per-user estimate of how much each interest/head should matter.\"},{\"question\":\"How does BACH improve training and serving for multi-interest retrieval?\",\"answer\":\"BACH treats the k heads as components of a per-user mixture of softmaxes and learns mixture weights with variational inference, so gradients flow softly to all heads. The per-user weights are reused at serving to weight head-level retrieval.\"}]",1784200669,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"bach-bayesian-admixture-of-contrastive-heads-for-multi-interest-two-tower-retrieval","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/bach-bayesian-admixture-of-contrastive-heads-for-multi-interest-two-tower-retrieval/85053/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do standard two-tower retrievers struggle with multi-interest users?","Question",{"text":75,"@type":76},"They represent a user with a single embedding, forcing items satisfying different interests to cluster near one point. This skews retrieval toward the dominant interest and under-serves others.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What problem does BACH address compared with hard-routing multi-interest models?",{"text":80,"@type":76},"Hard-routing training can lead to routing collapse, where heads selected early dominate and others receive fewer updates. It also lacks a per-user estimate of how much each interest/head should matter.",{"name":82,"@type":73,"acceptedAnswer":83},"How does BACH improve training and serving for multi-interest retrieval?",{"text":84,"@type":76},"BACH treats the k heads as components of a per-user mixture of softmaxes and learns mixture weights with variational inference, so gradients flow softly to all heads. The per-user weights are reused at serving to weight head-level retrieval.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]