[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-seo-191817-105":3,"detail-sidebar-cat-1-en-105":81,"doc-detail-191817-en":127},{"code":4,"msg":5,"data":6},0,"ok",{"site_id":7,"language":8,"slug":9,"title":10,"keywords":11,"description":12,"schema_data":13,"social_meta":74,"head_meta":76,"extra_data":78,"updated_unix":80},105,"en","automatic-extraction-of-titles-from-general-documents-using-machine-learning","Automatic Extraction of Titles from General Documents using Machine Learning","","A machine learning approach enables automatic title extraction from general documents spanning diverse genres such as presentations, technical papers, brochures, reports, and letters. The method annotates titles in Office files (Word and PowerPoint) to train models and uses formatting signals, especially font size, as primary features. Experiments on intranet data show strong performance, with Word precision/recall of 0.810/0.837 and PowerPoint precision/recall of 0.875/0.895. Cross-domain and cross-language training transfer is also supported, and extracted titles improve document search ranking.",{"@graph":14,"@context":73},[15,34,56],{"@type":16,"itemListElement":17},"BreadcrumbList",[18,23,27,31],{"item":19,"name":20,"@type":21,"position":22},"https://docshare.wps.com","Home","ListItem",1,{"item":24,"name":25,"@type":21,"position":26},"https://docshare.wps.com/template/","Template",2,{"item":28,"name":29,"@type":21,"position":30},"https://docshare.wps.com/template/presentations/","Presentations",3,{"item":32,"name":10,"@type":21,"position":33},"https://docshare.wps.com/template/automatic-extraction-of-titles-from-general-documents-using-machine-learning/191817/",4,{"url":32,"name":10,"@type":35,"image":36,"author":41,"headline":10,"publisher":44,"fileFormat":47,"inLanguage":8,"description":12,"dateModified":48,"datePublished":49,"encodingFormat":47,"isAccessibleForFree":50,"interactionStatistic":51},"DigitalDocument",{"url":37,"@type":38,"width":39,"height":40},"https://docshare.wps.com/thumbnails/automatic-extraction-of-titles-from-general-documents-using-machine-learning/191817.png","ImageObject",442,249,{"name":42,"@type":43},"Adam","Person",{"url":19,"name":45,"@type":46},"DocShare","Organization","application/octet-stream","2026-10-05","2026-09-03",true,{"@type":52,"interactionType":53,"userInteractionCount":55},"InteractionCounter",{"@type":54},"ViewAction",6,{"@type":57,"mainEntity":58},"FAQPage",[59,65,69],{"name":60,"@type":61,"acceptedAnswer":62},"What problem does the paper address?","Question",{"text":63,"@type":64},"It addresses how to automatically extract document titles from general documents, not only from well-structured research papers.","Answer",{"name":66,"@type":61,"acceptedAnswer":67},"How does the proposed method perform title extraction?",{"text":68,"@type":64},"It annotates titles in Word and PowerPoint examples, trains multiple machine learning models, and relies mainly on formatting information such as font size as features.",{"name":70,"@type":61,"acceptedAnswer":71},"What key results and benefits are reported?",{"text":72,"@type":64},"Experiments show accurate title extraction (with reported precision and recall for Word and PowerPoint), models can transfer across domains and languages, and using extracted titles can significantly improve search ranking.","https://schema.org",{"og:url":32,"og:type":75,"og:title":10,"og:site_name":45,"og:description":12},"article",{"robots":77,"canonical":32},"index,follow",{"doc_id":79,"site_id":7},191817,1788410969,{"code":4,"msg":82,"data":83},"success",[84,88,93,98,103,108,113,118,123],{"id":85,"doc_module":22,"doc_module_name":25,"category_name":29,"show_sort_weight":86,"slug":87},11,90,"presentations",{"id":89,"doc_module":22,"doc_module_name":25,"category_name":90,"show_sort_weight":91,"slug":92},12,"Resumes",80,"resumes",{"id":94,"doc_module":22,"doc_module_name":25,"category_name":95,"show_sort_weight":96,"slug":97},14,"Invoices",70,"invoices",{"id":99,"doc_module":22,"doc_module_name":25,"category_name":100,"show_sort_weight":101,"slug":102},15,"Posters",60,"posters",{"id":104,"doc_module":22,"doc_module_name":25,"category_name":105,"show_sort_weight":106,"slug":107},16,"Social Media",50,"social-media",{"id":109,"doc_module":22,"doc_module_name":25,"category_name":110,"show_sort_weight":111,"slug":112},17,"Forms",40,"forms",{"id":114,"doc_module":22,"doc_module_name":25,"category_name":115,"show_sort_weight":116,"slug":117},18,"Letters",30,"letters",{"id":119,"doc_module":22,"doc_module_name":25,"category_name":120,"show_sort_weight":121,"slug":122},21,"Paper Templates",5,"papers-templates",{"id":124,"doc_module":22,"doc_module_name":25,"category_name":125,"show_sort_weight":4,"slug":126},158,"General","general-158",{"code":4,"msg":82,"data":128},{"doc_id":79,"user_id":129,"nickname":42,"user_avatar":130,"doc_module":22,"category_id":85,"category_name":29,"doc_title":10,"doc_description":12,"doc_content":131,"file_id":132,"file_url":133,"file_type":134,"file_size":135,"view_count":55,"is_deleted":4,"is_public":22,"is_downloadable":22,"audit_status":22,"page_count":109,"language":136,"language_code":8,"site_id":7,"html_lang":8,"table_of_contents":137,"faqs":138,"seo_title":139,"seo_description":12,"update_tm":80,"read_time":55},1374404737137,"https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d","Automatic Extraction of Titles from General Documents using Machine Learning\nFeb. 15, 2006\nMSR-TR-2006-17\nMicrosoft Research\nMicrosoft Corporation\nOne Microsoft Way\nRedmond, WA  98052\n\u000fAutomatic Extraction of Titles from General Documents using Machine Learning\nYunhua Hu1\nComputer Science Department\u000bXi’an Jiaotong University\u000bNo 28, Xianning West Road\u000bXi'an, China, 710049\nyunhuahu@mail.xjtu.edu.cn\nDmitriy Meyerzon\nMicrosoft Corporation\u000bOne Microsoft Way\u000bRedmond, WA, \u000bUSA, 98052\ndmitriym@microsoft.com\n\u000e Hang Li, Yunbo Cao\nMicrosoft Research Asia\u000b5F Sigma Center, \u000bNo. 49 Zhichun Road, Haidian, \u000bBeijing, China, 100080\n{hangli,yucao}@microsoft.com\nQinghua Zheng\nComputer Science Department\u000bXi’an Jiaotong University\u000bNo 28, Xianning West Road\u000bXi'an, China, 710049\nqhzheng@mail.xjtu.edu.cn\n\u000eLi Teng\nComputer Science and Engineering, \u000bChinese University of Hong Kong, \u000bShatin, N.T., \u000bHong Kong, China\n\u0013 HYPERLINK \"http://by14fd.bay14.hotmail.msn.com/cgi-bin/compose?curmbox=F000000001&a=76816560e9aacdd215c43639a924648095e57c27454f6c0233a7a7ba135bda76&mailto=1&to=lteng@cse.cuhk.edu.hk&msg=MSG1125992514.13&start=1650911&len=1935&src=&type=x\" \u0014lteng@cse.cuhk.edu.hk\u0015\nABSTRACT\nIn this paper, we propose a machine learning approach to title extraction from general documents. By general documents, we mean documents that can belong to any one of a number of specific genres, including presentations, book chapters, technical papers, brochures, reports, and letters. Previously, methods have been proposed mainly for title extraction from research papers. It has not been clear whether it could be possible to conduct automatic title extraction from general documents. As a case study, we consider extraction from Office including Word and PowerPoint. In our approach, we annotate titles in sample documents (for Word and PowerPoint respectively) and take them as training data, train machine learning models, and perform title extraction using the trained models. Our method is unique in that we mainly utilize formatting information such as font size as features in the models. It turns out that the use of formatting information can lead to quite accurate extraction from general documents. Precision and recall for title extraction from Word are 0.810 and 0.837 respectively, and precision and recall for title extraction from PowerPoint are 0.875 and 0.895 respectively in an experiment on intranet data. Other important new findings in this work include that we can train models in one domain and apply them to other domains, and more surprisingly we can even train models in one language and apply them to other languages. Moreover, we can significantly improve search ranking results in document retrieval by using the extracted titles.\nKeywords: information extraction, metadata extraction, machine learning, search\nINTRODUCTION\nMetadata of documents is useful for many kinds of document processing such as search, browsing, and filtering. Ideally, metadata is defined by the authors of documents and is then used by various systems. However, people seldom define document metadata by themselves, even when they have convenient metadata definition tools (Crystal & Land, 2003). Thus, how to automatically extract metadata from the bodies of documents turns out to be an important research issue.\nMethods for performing the task have been proposed. However, the focus was mainly on extraction from research papers. For instance, Han et al. (Han, Giles, Manavoglu, Zha, Zhang, & Fox, 2003) proposed a machine learning based method to conduct extraction from research papers. They formalized the problem as that of classification and employed Support Vector Machines as the classifier. They mainly used linguistic features in the model.\nIn this paper, we consider metadata extraction from general documents. By general documents, we mean documents that may belong to any one of a number of specific genres. General documents are more widely available in digital libraries, intranets and the intern","cbCaiodnuiC7g02C","https://ap.wps.com/l/cbCaiodnuiC7g02C","doc",588800,"English","# Abstract\n# Introduction\n## Metadata for document processing\n## Title extraction from general documents\n## Machine learning approach and models\n## Research questions and experimental results","[{\"question\":\"What problem does the paper address?\",\"answer\":\"It addresses how to automatically extract document titles from general documents, not only from well-structured research papers.\"},{\"question\":\"How does the proposed method perform title extraction?\",\"answer\":\"It annotates titles in Word and PowerPoint examples, trains multiple machine learning models, and relies mainly on formatting information such as font size as features.\"},{\"question\":\"What key results and benefits are reported?\",\"answer\":\"Experiments show accurate title extraction (with reported precision and recall for Word and PowerPoint), models can transfer across domains and languages, and using extracted titles can significantly improve search ranking.\"}]","Automatic Extraction of Titles from General Documents using Machine Learning | DOC"]