[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-seo-240016-105":3,"detail-sidebar-cat-1-en-105":80,"doc-detail-240016-en":126},{"code":4,"msg":5,"data":6},0,"ok",{"site_id":7,"language":8,"slug":9,"title":10,"keywords":11,"description":12,"schema_data":13,"social_meta":73,"head_meta":75,"extra_data":77,"updated_unix":79},105,"en","a-comparative-study-of-bing-web-n-gram-language-models-summary","A Comparative Study of Bing Web N-gram Language Models - Summary","","This paper compares Microsoft Web N-gram Language Models (MWNLM) with earlier resources on three web search and natural language processing tasks: search query spelling correction, query reformulation, and statistical machine translation. MWNLM and Microsoft Web N-gram Services make smoothed n-gram probabilities readily accessible, avoiding the training complexity of large language model corpora such as LDC Gigaword and Google Web 1T. Results show MWNLM outperform baseline n-gram models across all tasks, with notable gains from combining multiple web fields and queries using zero count cutoff.",{"@graph":14,"@context":72},[15,34,55],{"@type":16,"itemListElement":17},"BreadcrumbList",[18,23,27,31],{"item":19,"name":20,"@type":21,"position":22},"https://docshare.wps.com","Home","ListItem",1,{"item":24,"name":25,"@type":21,"position":26},"https://docshare.wps.com/template/","Template",2,{"item":28,"name":29,"@type":21,"position":30},"https://docshare.wps.com/template/general/","General",3,{"item":32,"name":10,"@type":21,"position":33},"https://docshare.wps.com/template/a-comparative-study-of-bing-web-n-gram-language-models-summary/240016/",4,{"url":32,"name":10,"@type":35,"image":36,"author":41,"headline":10,"publisher":44,"fileFormat":47,"inLanguage":8,"description":12,"dateModified":48,"datePublished":49,"encodingFormat":47,"isAccessibleForFree":50,"interactionStatistic":51},"DigitalDocument",{"url":37,"@type":38,"width":39,"height":40},"https://docshare.wps.com/thumbnails/a-comparative-study-of-bing-web-n-gram-language-models-summary/240016.png","ImageObject",442,249,{"name":42,"@type":43},"Ethan Miller","Person",{"url":19,"name":45,"@type":46},"DocShare","Organization","application/pdf","2026-09-25","2026-09-11",true,{"@type":52,"interactionType":53,"userInteractionCount":26},"InteractionCounter",{"@type":54},"ViewAction",{"@type":56,"mainEntity":57},"FAQPage",[58,64,68],{"name":59,"@type":60,"acceptedAnswer":61},"What tasks does the paper evaluate using MWNLM?","Question",{"text":62,"@type":63},"The paper evaluates search query spelling correction, query reformulation, and statistical machine translation using MWNLM-based web scale language models.","Answer",{"name":65,"@type":60,"acceptedAnswer":66},"Why are MWNLM and the Microsoft Web N-gram Services easier to use than earlier corpora?",{"text":67,"@type":63},"They provide direct access to smoothed n-gram probabilities via web services, eliminating the need to train models from large text corpora such as LDC Gigaword and Google Web 1T.",{"name":69,"@type":60,"acceptedAnswer":70},"How does MWNLM perform compared with n-gram models trained on other datasets?",{"text":71,"@type":63},"MWNLM outperforms n-gram models trained on the Gigaword corpus and the Google Web 1T N-gram corpus across all three tasks, with significant improvements on spelling correction and reformulation.","https://schema.org",{"og:url":32,"og:type":74,"og:title":10,"og:site_name":45,"og:description":12},"article",{"robots":76,"canonical":32},"index,follow",{"doc_id":78,"site_id":7},240016,1790359600,{"code":4,"msg":81,"data":82},"success",[83,88,93,98,103,108,113,118,123],{"id":84,"doc_module":22,"doc_module_name":25,"category_name":85,"show_sort_weight":86,"slug":87},11,"Presentations",90,"presentations",{"id":89,"doc_module":22,"doc_module_name":25,"category_name":90,"show_sort_weight":91,"slug":92},12,"Resumes",80,"resumes",{"id":94,"doc_module":22,"doc_module_name":25,"category_name":95,"show_sort_weight":96,"slug":97},14,"Invoices",70,"invoices",{"id":99,"doc_module":22,"doc_module_name":25,"category_name":100,"show_sort_weight":101,"slug":102},15,"Posters",60,"posters",{"id":104,"doc_module":22,"doc_module_name":25,"category_name":105,"show_sort_weight":106,"slug":107},16,"Social Media",50,"social-media",{"id":109,"doc_module":22,"doc_module_name":25,"category_name":110,"show_sort_weight":111,"slug":112},17,"Forms",40,"forms",{"id":114,"doc_module":22,"doc_module_name":25,"category_name":115,"show_sort_weight":116,"slug":117},18,"Letters",30,"letters",{"id":119,"doc_module":22,"doc_module_name":25,"category_name":120,"show_sort_weight":121,"slug":122},21,"Paper Templates",5,"papers-templates",{"id":124,"doc_module":22,"doc_module_name":25,"category_name":29,"show_sort_weight":4,"slug":125},158,"general-158",{"code":4,"msg":81,"data":127},{"doc_id":78,"user_id":128,"nickname":42,"user_avatar":129,"doc_module":22,"category_id":124,"category_name":29,"doc_title":10,"doc_description":12,"doc_content":130,"file_id":131,"file_url":132,"file_type":133,"file_size":134,"view_count":30,"is_deleted":4,"is_public":22,"is_downloadable":22,"audit_status":22,"page_count":135,"language":136,"language_code":8,"site_id":7,"html_lang":8,"table_of_contents":137,"faqs":138,"seo_title":139,"seo_description":12,"update_tm":140,"read_time":26},687207017582,"https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d","A Comparative Study of Bing Web N-gram Language Models for Web Search and Natural Language Processing  \nJianfeng Gao, Patrick Nguyen, Xiaolong Li, Chris Thrasher, Mu Li, Kuansan Wang  \nMicrosoft Research  \nOne Microsoft Way  \nRedmond, WA 98052 USA  \n{jfgao; panguyen; xiaolli; cthrash; muli; [kuansanw}@microsoft.com](kuansanw}@microsoft.com)  \nABSTRACT  \nThis paper presents a comparative study of the recently released Microsoft Web N-gram Language Models (MWNLM)1 on three web search and natural language processing tasks: search query spelling correction, query reformulation, and statistical machine translation. MWNLM, as well as the corresponding web services, called Microsoft Web N-gram Services, are much more accessible and easier to use than the previously released text corpora used for large language model training, including the LDC English Gigaword corpus and the Google Web 1T N-gram corpus, because the Microsoft Web N-gram Services provide the access to the smoothed n-gram probabilities based on a set of language models trained from the different text fields from the web documents as well as search queries. Our results show that MWNLM outperform the n-gram models trained on the Gigaword corpus and the Google Web 1T N-gram corpus on all the three tasks. In particular, the significant improvements on search query spelling correction and search query reformulation, resulting from MWNLM, demonstrate the benefit of training multiple language models on different portions of web data and search queries in a principled way with zero count cutoff.  \nCategories and Subject Descriptors  \nH.3.3 [Information Storage and Retrieval]: Information Search and Retrieval;  \nGeneral Terms  \nExperimentation, Measurement  \nKeywords  \nLanguage Model, N-gram, Spelling Correction, Query Reformulation, Statistical Machine Translation  \nPermission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. SIGIR'10, July 19-23, 2010, Geneva, Swizerland.  \nCopyright 2010 ACM 978-1-60558-896-4/10/07…$10.00.  \n1 An earlier version of MWNLM is described in Huang et al.  \n[10] . This paper provides updated information regarding the most recent development of MWNLM.  \n1. INTRODUCTION  \nThe goal of a statistical language model (LM) is to predict the probability of a word string. This is fundamental to a wide variety of web search and natural language processing (NLP) applications. The technique that is still dominating the research communities is the n-gram model, thanks to its simplicity and effectiveness. Let denote a string of L words  \nover a fixed vocabulary. An n-gram language model assigns a probability to according to  \n∏ ( | ) ∏ ( | ) (1)  \nwhere the approximation is based on a Markov assumption that each word depends only upon the immediately preceding n-1 words.  \nN-gram models have been extensively studied in both the research communities and the industry from different perspectives for decades. While the people in the research communities, such as natural language processing, speech and information retrieval, try to figure out a richer and smarter model via better smoothing [6] or capturing more linguistic structures [5]; industry people recently found that simply using more data is far more effective [e.g., 2] . For example, the Google machine translation group trades the mathematical soundness to the scalability of LMs, and use the n-gram models that are not properly normalized but can be efficiently trained on very large amounts of text corpora [2] .  \nIn this paper we strike a better balance between the mathematical soundness and scalability, and demonstrate that it is possible to scale the n-gram LMs wit","cbCaifnkCpRtgidy","https://ap.wps.com/l/cbCaifnkCpRtgidy","pdf",620557,6,"English","# ABSTRACT\n# INTRODUCTION\n## N-gram language model background\n## MWNLM platform and web services\n## Experiments and tasks\n## Model collection statistics","[{\"question\":\"What tasks does the paper evaluate using MWNLM?\",\"answer\":\"The paper evaluates search query spelling correction, query reformulation, and statistical machine translation using MWNLM-based web scale language models.\"},{\"question\":\"Why are MWNLM and the Microsoft Web N-gram Services easier to use than earlier corpora?\",\"answer\":\"They provide direct access to smoothed n-gram probabilities via web services, eliminating the need to train models from large text corpora such as LDC Gigaword and Google Web 1T.\"},{\"question\":\"How does MWNLM perform compared with n-gram models trained on other datasets?\",\"answer\":\"MWNLM outperforms n-gram models trained on the Gigaword corpus and the Google Web 1T N-gram corpus across all three tasks, with significant improvements on spelling correction and reformulation.\"}]","A Comparative Study of Bing Web N-gram Language Models - Summary | PDF",1789156272]