[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127847-en":3,"doc-seo-127847-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},127847,2336474466712,"Maeve","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Machine Learning Approaches to Text Representation using Unlabeled Data","With the rapid expansion in the use of computers for producing digitalized textual documents, the need of automatic systems for organizing and retrieving the information contained in large databases has become essential. In general, information retrieval systems rely on a formal description or representation of documents enabling their automatic processing. The work proposes probabilistic, neural, and multitask machine-learning approaches to enrich document representations by leveraging unlabeled data and learning word usage patterns from context.","View metadata, citation and similar [papers at ](papers at core.ac.uk)[core.ac.uk](papers at core.ac.uk) brought to you by CORE  \nprovided by Infoscience- École polytechnique fédérale de Lausanne  \nR T  \nI D I A P R E S E A R C H R E P O  \nMachine Learning Approaches to Text Representation using  \nUnlabeled Data Mikaela Keller a IDIAP–RR YY-XX  \nMarch 9, 2007  \npublished in Ecole Polytechnique F􀀓ed􀀓erale de Lausanne  \na IDIAP Research Institute, CP 592, 1920 Martigny, Switzerland, [firstname.name@idiap.ch](firstname.name@idiap.ch)  \nIDIAP Research Institute [www.idiap.ch](www.idiap.ch)  \nRue du Simplon 4 P.O. Box 592 1920 Martigny − Switzerland  \nTel: +41 27 721 77 11 Fax: +41 27 721 77 12 Email: [info@idiap.ch](info@idiap.ch)  \nIDIAP Research Report YY-XX  \nMachine Learning Approaches to Text Representation using Unlabeled Data  \nMikaela Keller  \nMarch 9, 2007  \npublished in  \nEcole Polytechnique F􀀓ed􀀓erale de Lausanne  \n\n| \u003Cbr> ✻\u003Cbr> ✐4 ✐5 ✐2 \u003Cbr>  ❄ ❄   ❄    \u003Cbr> ✻ ❄ Header \u003Cbr>  ✻ \u003Cbr>\u003Cbr> ✐6 \u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>Footer ✻ |\n| --- |\n| ~~􀀛 ✐~~1 ~~ ✲~~\u003Cbr>\u003Cbr> |\n\n\n| Margin\u003Cbr>Notes\u003Cbr>~~✐~~9 ~~ ✲~~\u003Cbr>􀀛 | ✻\u003Cbr>Body\u003Cbr>~~􀀛 ~~\u003Cbr>~~􀀛 ~~\u003Cbr>~~􀀛 ~~✐8 ✐11 \u003Cbr>❄\u003Cbr>❄ | ✻\u003Cbr>✐7 \u003Cbr>✲ |\n| --- | --- | --- |\n| ✐10 \u003Cbr>✲\u003Cbr>~~✐~~3 ~~ ✲~~ |  |  |\n\n1 one inch + \\hoffset  \n3 \\evensidemargin = -1pt  \n5 \\headheight = 12pt  \n7 \\textheight = 625pt  \n9 \\marginparsep = 7pt  \n11 \\footskip = 25pt \\hoffset = 0pt \\paperwidth = 597pt  \n2 one inch + \\voffset  \n4 \\topmargin = 0pt  \n6 \\headsep = 18pt  \n8 \\textwidth = 455pt  \n10 \\marginparwidth = 115pt \\marginparpush = 5pt (not shown)\\voffset = 0pt  \n\\paperheight = 845pt  \n✐  \n✐3 ✲􀀛  \n􀀛  \n􀀛 ✐1 ✲  \n✐4   \n❄✻  \n✐5   \n❄  \n✻  \n❄  \n✻  \n✐6   \nHeader  \nBody  \n✻  \n✐2   \n❄  \n✻  \n✐7   \nMargin Notes  \n✐11   \n❄  \n✻  \n✐8   \n✐9 ✲ 􀀛  \n􀀛 10 ✲  \n✲  \n❄ Footer  \n1 one inch + \\hoffset  \n3 \\oddsidemargin = -1pt  \n5 \\headheight = 12pt  \n7 \\textheight = 625pt  \n9 \\marginparsep = 7pt  \n11 \\footskip = 25pt \\hoffset = 0pt \\paperwidth = 597pt  \n2 one inch + \\voffset  \n4 \\topmargin = 0pt  \n6 \\headsep = 18pt  \n8 \\textwidth = 455pt  \n10 \\marginparwidth = 115pt \\marginparpush = 5pt (not shown)\\voffset = 0pt  \n\\paperheight = 845pt  \n4 IDIAP–RR YY-XX  \nR􀀓esum􀀓e  \nAvec l’essor de l’usage des ordinateurs pour la cr􀀓eation de documents textuels digitalis􀀓es, le besoin de syst􀀒emes automatiques pour la recherche et l’organisation de l’information contenue dans de grandes bases de donn􀀓ees est devenu central. En g􀀓en􀀓eral, les syst􀀒emes de recherche d’information s’appuient sur une description formelle (ou repr􀀓esentation) des documents permettant leur traitement automatique.  \nDans la plus commune des repr􀀓esentations, appel􀀓ee sac-de-mots, les documents sont repr􀀓esent􀀓es par l’ensemble des mots les constituant. Deux documents (ou bien un document et une requˆete) sont consid􀀓er􀀓es comme similaires s’ils ont un grand nombre de mots en commun.  \nIl est raisonable de penser que les syst􀀒emes de recherche d’information devraient pouvoir utiliser les grandes quantit􀀓es de donn􀀓ees textuelles disponibles pour “apprendre”, 􀀒a la fa􀀘con des humains, les di􀀋􀀓erents emplois d’un mot en fonction de son contexte. Cette information devrait pouvoir ˆetre utilis􀀓ee pour enrichir la repr􀀓esentation des documents.  \nDans cette th􀀒ese, nous d􀀓eveloppons plusieurs approches originales d’apprentissage automatique quitentent d’atteindre ce but.  \nComme premi􀀒ere approche pour la repr􀀓esentation de documents nous proposons un model􀀓e probabiliste qui suppose que les documents sont tir􀀓es d’un m􀀓elange de distributions sur des “th􀀒emes”, repr􀀓esent􀀓es par une variable cach􀀓ee qui conditionne une distribution multinomiale sur les mots. Simultan􀀓ement, ce mod􀀒ele suppose que les mots sont tir􀀓es d’une distribution sur les “sujets”, repr􀀓esent􀀓es quant 􀀒a eux par une seconde variable cach􀀓ee d􀀓ependante des th􀀒emes.  \nComme deuxi􀀒eme approche, un r􀀓eseau de neuron","cbCainzHJgk7vyKv","https://ap.wps.com/l/cbCainzHJgk7vyKv","pdf",982916,1,101,"English","en",105,"# Abstract\n## Document representation goals\n## Approaches: probabilistic model, neural network, multitask learning","[{\"question\":\"What problem does the thesis address?\",\"answer\":\"It addresses how to automatically organize and retrieve information from large collections of digital textual documents by improving document representations for information retrieval.\"},{\"question\":\"Why does the thesis focus on unlabeled data?\",\"answer\":\"It aims to enrich document representations by using non-labeled data to learn better models of word usage and context.\"},{\"question\":\"What main approaches are proposed?\",\"answer\":\"The thesis develops a probabilistic topic-based model, a neural network approach scoring word suitability in context, and a multitask learning approach that jointly handles retrieval and representation enrichment.\"}]","Machine Learning Approaches to Text Representation using Unlabeled Data | PDF",1785942336,255,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"machine-learning-approaches-to-text-representation-using-unlabeled-data","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-approaches-to-text-representation-using-unlabeled-data/127847/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the thesis address?","Question",{"text":76,"@type":77},"It addresses how to automatically organize and retrieve information from large collections of digital textual documents by improving document representations for information retrieval.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Why does the thesis focus on unlabeled data?",{"text":81,"@type":77},"It aims to enrich document representations by using non-labeled data to learn better models of word usage and context.",{"name":83,"@type":74,"acceptedAnswer":84},"What main approaches are proposed?",{"text":85,"@type":77},"The thesis develops a probabilistic topic-based model, a neural network approach scoring word suitability in context, and a multitask learning approach that jointly handles retrieval and representation enrichment.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]