[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85380-en":3,"doc-seo-85380-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85380,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Improved Answer Selection with Pre-Trained Word Embeddings","This paper evaluates existing and newly proposed answer selection methods that incorporate pre-trained word embeddings into the ranking stage of question answering systems. Word embeddings capture semantic relatedness between questions and candidate answers, addressing lexical gaps caused by differing word choices. Experiments on three publicly available datasets demonstrate significant improvements over traditional term-frequency retrieval in supervised and unsupervised settings. It also shows that combining embedding features with learning-to-rank approaches matches the effectiveness of state-of-the-art neural networks.","Improved Answer Selection with Pre-Trained Word Embeddings  \nRishav Chakravarti  \nIBM Watson  \n[rchakravarti@us.ibm.com](rchakravarti@us.ibm.com)  \nJiˇ´rı Navr´atil  \nIBM Watson [jiri@us.ibm.com](jiri@us.ibm.com)  \nC´ıcero Nogueira dos Santos  \nAI Foundations, IBM Research [cicerons@us.ibm.com](cicerons@us.ibm.com)  \narXiv : 1708 .04326v1 [ cs .IR] 14 Aug 2017  \nABSTRACT  \n􀂌is paper evaluates existing and newly proposed answer selection methods based on pre-trained word embeddings. Word embeddings are highly eﬀective in various natural language processing tasks and their integration into traditional information retrieval (IR)systems allows for the capture ofsemantic relatedness between questions and answers. Empirical results on three publicly available datasets show signiﬁcant gains over traditional term frequency based approaches in both supervised and unsupervised se􀂊ings. We show that combining these word embedding features with traditional learning-to-rank techniques can achieve similar performance to state-of-the-art neural networks trained for the answer selection task.  \nKEYWORDS  \nword embeddings, rank, re-rank, question answering, answer selection  \n1 INTRODUCTION  \nA core objective of 􀂋estion Answering (QA) systems is to maximize the utility of selected candidate answers with respect to the user’s question. O􀂉en, QA systems implement a solution in two phases:  \n(1) an initial retrieval phase based on term overlap with the (expanded) question  \n(2) a secondary ranking of the candidates to maximize relevance of the top-k answers with respect to the question  \nTraditionally, secondary ranking uses language or relevance modeling approaches that still rely on term overlap statistics between the (expanded) question and the answer text [8, 9] . Reliance on term overlap can suﬀer from lexical gaps between the language used to express the question and the language used in the candidate answer text even when there are semantic matches.  \nWord embeddings such as Word2Vec [10] and GloVe [13] have surfaced as a way to model the semantics behind terms in various natural language processing tasks. We evaluate existing and newly proposed approaches that integrate such pre-trained word embeddings into the ranking phase of the QA pipeline. We show experimental results in both supervised and unsupervised se􀂊ings with comparisons to traditional IR as well as a state-of-the-art deep learning system.  \n􀂌e main contributions of this work are: (1) a thorough comparison of recently proposed methods to integrate word embeddings in the QA pipeline; (2) the proposal of an extension to the Relaxed Word Mover’s Distance method that incorporates matching terms proximity in the answer text; (3) an empirical demonstration that combining traditional features with word embedding based features can boost the result of learning to rank approaches.  \nSection2reviews the existing word embedding based approaches for ranking which we evaluate in Section 5 . In addition, Section 3 proposes additional techniques to address generalizability while maintaining performance. Sections 4 and 5 describe the experimental setup and results, respectively. Finally, in Section 6 we summarize ﬁndings on combining traditional IR techniques with word embedding based features.  \n2 RELATED WORK  \nIn [21], the authors propose using cosine similarity between embeddings for two words as an estimate of translation probability between them. 􀂌e embeddings are pre-trained using a large text corpus using Word2Vec. 􀂌is translation probability is then plugged into a language modeling framework that ranks candidate answers based on the likelihood score of each candidate answer producing the question terms (with some smoothing and collection eﬀects taken into account) . 􀂌eir approach shows some improvements over a traditional language modeling baseline.  \nWork such as [19] similarly incorporate word embeddings into the language model for the query expansion task. Our work here on the ran","cbCaipNGnOecGdz4","https://ap.wps.com/l/cbCaipNGnOecGdz4","pdf",148340,4,1,5,"English","en",105,"# Introduction\n## Contributions\n# Related Work\n## Embedding-based ranking approaches\n## Hybrid models and neural reranking\n# Method Overview\n## Relaxed Word Mover’s Distance extension\n## Min-Max Pooling feature\n# Experiments and Results\n## Experimental setup\n## Performance comparisons\n# Conclusion","[{\"question\":\"为什么在回答选择任务中引入预训练词嵌入？\",\"answer\":\"词嵌入能建模问题与候选答案之间的语义相关性，从而减轻由于表面词汇不同带来的“词汇鸿沟”，即便语义匹配也可能缺少词重叠。\"},{\"question\":\"文中将问题回答流程分为哪些阶段？\",\"answer\":\"系统通常包含两个阶段：先进行基于（扩展）问题与候选答案的词项重叠检索，再对候选答案进行二次排序以最大化前k个答案与问题的相关性。\"},{\"question\":\"把词嵌入特征与传统学习排序方法结合能带来什么效果？\",\"answer\":\"文中表明，将词嵌入特征与传统 learning-to-rank 技术组合，可以提升学习排序结果，并能达到与针对答案选择任务训练的最新神经网络相当的性能。\"}]",1784203044,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"improved-answer-selection-with-pre-trained-word-embeddings","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/improved-answer-selection-with-pre-trained-word-embeddings/85380/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么在回答选择任务中引入预训练词嵌入？","Question",{"text":75,"@type":76},"词嵌入能建模问题与候选答案之间的语义相关性，从而减轻由于表面词汇不同带来的“词汇鸿沟”，即便语义匹配也可能缺少词重叠。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"文中将问题回答流程分为哪些阶段？",{"text":80,"@type":76},"系统通常包含两个阶段：先进行基于（扩展）问题与候选答案的词项重叠检索，再对候选答案进行二次排序以最大化前k个答案与问题的相关性。",{"name":82,"@type":73,"acceptedAnswer":83},"把词嵌入特征与传统学习排序方法结合能带来什么效果？",{"text":84,"@type":76},"文中表明，将词嵌入特征与传统 learning-to-rank 技术组合，可以提升学习排序结果，并能达到与针对答案选择任务训练的最新神经网络相当的性能。","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":22,"slug":137},19,"General","general"]