[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85331-en":3,"doc-seo-85331-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85331,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Score Only Distillation for Compact Dense Retrieval","Large embedding models boost retrieval quality, yet serving them online is costly. This work investigates whether a compact retriever can mimic teacher ranking behavior using only score vectors, without any access to teacher hidden states. Students are trained on rows from ground-truth positives and negative candidates generated by a dedicated pipeline, with hardnegative mining evaluated separately. A row-centered PairMSE objective enables memory-efficient uniform all-pairs learning, recovering up to 50% of the teacher gap on eight fixed tasks under matched protocols.","Score-Only Distillation for Compact Dense Retrieval  \nKirill Dubovikov  \nMohamed bin Zayed University of Artificial Intelligence Abu Dhabi, United Arab Emirates [kirill.dubovikov@mbzuai.ac.ae](kirill.dubovikov@mbzuai.ac.ae)  \nMartin Takáč  \nMohamed bin Zayed University of Artificial Intelligence Abu Dhabi, United Arab Emirates [martin.takac@mbzuai.ac.ae](martin.takac@mbzuai.ac.ae)  \nSalem Lahlou  \nMohamed bin Zayed University of Artificial Intelligence Abu Dhabi, United Arab Emirates [salem.lahlou@mbzuai.ac.ae](salem.lahlou@mbzuai.ac.ae)  \narXiv :2607 . 1 1465v 1 [ cs .IR] 13 Jul 2026  \nAbstract  \nLarge embedding models improve retrieval quality, but serving large encoders online is expensive. We study whether a compact retriever can learn teacher ranking behavior from score vectors without access to teacher hidden states. The student trains on rows built from ground-truth positives and negative candidates produced by our data generation pipeline; we evaluate student-teacher hardnegative mining separately as an extension. We use a row-centered score-vector objective, a memory-efficient implementation of uniform all-pairs PairMSE loss. On a fixed eight-task evaluation panel, our distillation protocol recovers up to 50% of the base-to-teacher gap. The distilled 0.6B student is 4.7× faster for query encoding and 9.7× faster for document encoding than sequential online teacher fusion. External-transfer performance after distillation remains mixed, so our evidence supports compression of teacher rankings under matched retrieval protocols.  \n1 Introduction  \nDense retrieval systems increasingly use large embedding models [14, 20, 29] . These encoders can provide strong rankings, but their size raises serving cost, index refresh cost, and operational complexity. We ask whether a compact retriever can learn teacher ranking behavior from black-box score vectors alone.  \nIn this setting, the teacher provides query-document scores over candidate rows, while the student receives no teacher embeddings, hidden states, logits over a shared vocabulary, or vector-space alignment. This score-only setting is more restrictive than embeddingalignment distillation, which trains on teacher output embeddingsor proxy teacher representations [13, 26] . Figure 1 summarizes the setting.  \nWe propose a memory-efficient reformulation of uniform allpairs PairMSE loss [19]: it uses the 􀀺 row scores rather than the 􀀺 (􀀺 − 1)/2 explicit pairwise margins. Given student scores and teacher scores over the same candidate row, we center the residual vector between them and minimize its squared norm.  \nOur primary contribution is the training protocol and its empirical characterization. We evaluate on a fixed eight-task panel that separates row-source adaptation from held-out full-corpus retrieval: SciFact [28], NFCorpus [3], and FiQA [17] provide training rows, while ArguAna [27], SciDocs [4], TREC-COVID [24], WebisTouche2020 [2], and Quora [22] are held out from row construction.  \nWe test our approach on two student models with different architectures: Qwen3-Embedding-0.6B [20] and intfloat/e5-large-v2 [29], and observe that the approach works for both models. Teacher scores matter: a label-only contrastive control on the same rows is far below the frozen base, and positive-negative MarginMSE [10] trails centered all-pairs matching.  \nWe also evaluate equal fusion, hard-negative mining, and active row selection as extensions rather than as the main protocol.  \nWe make three contributions:  \n• We define a black-box score-vector distillation protocol for compact retrieval students trained from teacher scores alone.  \n• We propose a row-centered implementation of uniform all-pairs PairMSE and compare it against label-only and positive-negative MarginMSE controls under the same rows, student, and evaluator.  \n• We characterize the protocol’s scope: score-vector distillation improves compact students under matched retrieval protocols, and we provide empirical analysis","cbCaiaKf1bJbF1OF","https://ap.wps.com/l/cbCaiaKf1bJbF1OF","pdf",2625907,1,5,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"What is the core idea of score-only distillation for compact dense retrieval?\",\"answer\":\"The method trains a compact retriever to reproduce a teacher’s ranking behavior using only the teacher’s score vectors. The student does not receive teacher embeddings, hidden states, or representation-space alignment signals.\"},{\"question\":\"How are the training data rows constructed for distillation?\",\"answer\":\"Training rows are built from ground-truth positives and negative candidates produced by a data generation pipeline. This fixed-row setup is used to compare student–teacher behavior under matched retrieval protocols.\"},{\"question\":\"What objective is used to learn from teacher score vectors efficiently?\",\"answer\":\"A row-centered implementation of uniform all-pairs PairMSE loss is used. The formulation centers the residual between student and teacher score vectors over the same candidate row and minimizes its squared norm with a memory-efficient computation.\"}]",1784202546,13,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"score-only-distillation-for-compact-dense-retrieval","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/score-only-distillation-for-compact-dense-retrieval/85331/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the core idea of score-only distillation for compact dense retrieval?","Question",{"text":75,"@type":76},"The method trains a compact retriever to reproduce a teacher’s ranking behavior using only the teacher’s score vectors. The student does not receive teacher embeddings, hidden states, or representation-space alignment signals.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are the training data rows constructed for distillation?",{"text":80,"@type":76},"Training rows are built from ground-truth positives and negative candidates produced by a data generation pipeline. This fixed-row setup is used to compare student–teacher behavior under matched retrieval protocols.",{"name":82,"@type":73,"acceptedAnswer":83},"What objective is used to learn from teacher score vectors efficiently?",{"text":84,"@type":76},"A row-centered implementation of uniform all-pairs PairMSE loss is used. The formulation centers the residual between student and teacher score vectors over the same candidate row and minimizes its squared norm with a memory-efficient computation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":21,"slug":137},19,"General","general"]