[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-1-en-105":3,"doc-seo-253863-105":53,"doc-detail-253863-en":126},{"code":4,"msg":5,"data":6},0,"success",[7,14,19,24,29,34,39,44,49],{"id":8,"doc_module":9,"doc_module_name":10,"category_name":11,"show_sort_weight":12,"slug":13},11,1,"Template","Presentations",90,"presentations",{"id":15,"doc_module":9,"doc_module_name":10,"category_name":16,"show_sort_weight":17,"slug":18},12,"Resumes",80,"resumes",{"id":20,"doc_module":9,"doc_module_name":10,"category_name":21,"show_sort_weight":22,"slug":23},14,"Invoices",70,"invoices",{"id":25,"doc_module":9,"doc_module_name":10,"category_name":26,"show_sort_weight":27,"slug":28},15,"Posters",60,"posters",{"id":30,"doc_module":9,"doc_module_name":10,"category_name":31,"show_sort_weight":32,"slug":33},16,"Social Media",50,"social-media",{"id":35,"doc_module":9,"doc_module_name":10,"category_name":36,"show_sort_weight":37,"slug":38},17,"Forms",40,"forms",{"id":40,"doc_module":9,"doc_module_name":10,"category_name":41,"show_sort_weight":42,"slug":43},18,"Letters",30,"letters",{"id":45,"doc_module":9,"doc_module_name":10,"category_name":46,"show_sort_weight":47,"slug":48},21,"Paper Templates",5,"papers-templates",{"id":50,"doc_module":9,"doc_module_name":10,"category_name":51,"show_sort_weight":4,"slug":52},158,"General","general-158",{"code":4,"msg":54,"data":55},"ok",{"site_id":56,"language":57,"slug":58,"title":59,"keywords":60,"description":61,"schema_data":62,"social_meta":119,"head_meta":121,"extra_data":123,"updated_unix":125},105,"en","part-of-speech-tagging-for-twitter-annotation-features-and-experiments","Part-of-Speech Tagging for Twitter - Annotation, Features, and Experiments","","English part-of-speech (POS) tagging for Twitter microblogs is addressed through a dedicated workflow: designing a Twitter tagset, manually annotating 1,827 tweets, engineering Twitter-specific features, and evaluating tagging performance. Results approach 90% accuracy while demonstrating that a carefully designed annotation scheme and feature set—leveraging resources such as tag dictionaries and phonetic normalization—can rapidly enable supervised NLP for new, idiosyncratic social media domains. Data and tools are released to support richer text analysis.",{"@graph":63,"@context":118},[64,80,101],{"@type":65,"itemListElement":66},"BreadcrumbList",[67,71,74,77],{"item":68,"name":69,"@type":70,"position":9},"https://docshare.wps.com","Home","ListItem",{"item":72,"name":10,"@type":70,"position":73},"https://docshare.wps.com/template/",2,{"item":75,"name":51,"@type":70,"position":76},"https://docshare.wps.com/template/general/",3,{"item":78,"name":59,"@type":70,"position":79},"https://docshare.wps.com/template/part-of-speech-tagging-for-twitter-annotation-features-and-experiments/253863/",4,{"url":78,"name":59,"@type":81,"image":82,"author":87,"headline":59,"publisher":90,"fileFormat":93,"inLanguage":57,"description":61,"dateModified":94,"datePublished":95,"encodingFormat":93,"isAccessibleForFree":96,"interactionStatistic":97},"DigitalDocument",{"url":83,"@type":84,"width":85,"height":86},"https://docshare.wps.com/thumbnails/part-of-speech-tagging-for-twitter-annotation-features-and-experiments/253863.png","ImageObject",442,249,{"name":88,"@type":89},"Stanford","Person",{"url":68,"name":91,"@type":92},"DocShare","Organization","application/pdf","2026-09-20","2026-09-13",true,{"@type":98,"interactionType":99,"userInteractionCount":76},"InteractionCounter",{"@type":100},"ViewAction",{"@type":102,"mainEntity":103},"FAQPage",[104,110,114],{"name":105,"@type":106,"acceptedAnswer":107},"What is the main goal of this work on Twitter data?","Question",{"text":108,"@type":109},"To build an English POS tagging system designed specifically for Twitter, including a tailored tagset, annotated data, features, and evaluation results.","Answer",{"name":111,"@type":106,"acceptedAnswer":112},"How was the Twitter dataset for POS tagging prepared?",{"text":113,"@type":109},"A POS tagset was developed and 1,827 tweets were manually annotated using an annotation scheme matched to Twitter’s characteristics.",{"name":115,"@type":106,"acceptedAnswer":116},"Which resources and feature ideas support improved tagging accuracy?",{"text":117,"@type":109},"The feature set captures Twitter-specific properties and exploits existing resources such as tag dictionaries and phonetic normalization.","https://schema.org",{"og:url":78,"og:type":120,"og:title":59,"og:site_name":91,"og:description":61},"article",{"robots":122,"canonical":78},"index,follow",{"doc_id":124,"site_id":56},253863,1789271115,{"code":4,"msg":5,"data":127},{"doc_id":124,"user_id":128,"nickname":88,"user_avatar":129,"doc_module":9,"category_id":50,"category_name":51,"doc_title":59,"doc_description":61,"doc_content":130,"file_id":131,"file_url":132,"file_type":133,"file_size":134,"view_count":76,"is_deleted":4,"is_public":9,"is_downloadable":9,"audit_status":9,"page_count":135,"language":136,"language_code":57,"site_id":56,"html_lang":57,"table_of_contents":137,"faqs":138,"seo_title":139,"seo_description":61,"update_tm":125,"read_time":73},2336477552062,"https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d","Part-of-Speech Tagging for Twitter: Annotation, Features, and Experiments  \nKevin Gimpel, Nathan Schneider, Brendan O'Connor, Dipanjan Das, Daniel Mills, Jacob Eisenstein, Michael Heilman, Dani Yogatama, Jeffrey Flanigan, and Noah A. Smith School of Computer Science, Carnegie Mellon Univeristy, Pittsburgh, PA 15213, USA  \n{kgimpel,nschneid,brenocon,dipanjan,dpmills,  \njacobeis,mheilman,dyogatama,jflanigan,[nasmith](nasmith}@cs.cmu.edu)[}](nasmith}@cs.cmu.edu)[@cs.cmu.edu](nasmith}@cs.cmu.edu)  \nAbstract  \nWe address the problem of part-of-speech tagging for English data from the popular microblogging service Twitter. We develop a tagset, annotate data, develop features, and report tagging results nearing 90% accuracy. The data and tools have been made available to the research community with the goal of enabling richer text analysis of Twitter and related social media data sets.  \n1 Introduction  \nThe growing popularity of social media and usercreated web content is producing enormous quantities of text in electronic form. The popular microblogging service Twitter ([twitter.com](twitter.com)) is one particularly fruitful source of user-created content, and a ﬂurry of recent research has aimed to understand and exploit these data (Ritter et al., 2010; Shariﬁ et al., 2010; Barbosa and Feng, 2010; Asur and Huberman, 2010; O'Connor et al., 2010a; Thelwallet al., 2011) . However, the bulk of this work eschews the standard pipeline of tools which might enable a richer linguistic analysis; such tools are typically trained on newstext and have been shown to perform poorly on Twitter (Finin et al., 2010) .  \nOne of the most fundamental parts of the linguistic pipeline is part-of-speech (POS) tagging, a basic form of syntactic analysis which has countless applications in NLP. Most POS taggers are trained from treebanks in the newswire domain, such as the Wall Street Journal corpus of the Penn Treebank (PTB; Marcus et al., 1993) . Tagging performance degradeson out-of-domain data, and Twitter poses additional challenges due to the conversational nature of the text, the lack of conventional orthography, and 140-character limit of each message (“tweet”) . Figure 1 shows three tweets which illustrate these challenges.  \n(a) @Gunservatively @ obozo ∧ willV goV nutsA when R PA ∧ electsV aD RepublicanA Governor N next P Tue ∧ . , Can V you O say V redistrictingV ? ,  \n(b) Spending V theD dayN withhhP mommmaN ! ,  \n(c) lmao! ... , s/oV toP the D coolA ass N asianA ofﬁcerN 4 P \\#1$ not R runninV myD license N and&  \n\\#2$ not R takinV druN booN toP jail N . , ThankVuO God ∧ . , \\#amen \\#  \nFigure 1: Example tweets with gold annotations. Underlined tokens show tagger improvements due to features detailed in Section 3 (respectively: TAGDICT, METAPH, and DISTSIM) .  \nIn this paper, we produce an English POS tagger that is designed especially for Twitter data. Our contributions are as follows:  \n• we developed a POS tagset for Twitter,  \n• we manually tagged 1,827 tweets,  \n• we developed features for Twitter POS tagging and conducted experiments to evaluate them, and  \n• we provide our annotated corpus and trained POStagger to the research community.  \nBeyond these speciﬁc contributions, we see this work as a case study in how to rapidly engineer a core NLP system for a new and idiosyncratic dataset. This project was accomplished in 200 person-hours spread across 17 people and two months. This was made possible by two things:  \n(1) an annotation scheme that ﬁts the unique characteristics of our data and provides an appropriate level of linguistic detail, and (2) a feature set that captures Twitter-speciﬁc properties and exploits existing resources such as tag dictionaries and phonetic normalization. The success of this approach demonstrates that with careful design, supervised machine learning can be applied to rapidly produce effective language technology in new domains.  \n42  \nProceedings of the 49th Annual Meeting of the Association for Computati","cbCaiobGDlTOi6fQ","https://ap.wps.com/l/cbCaiobGDlTOi6fQ","pdf",184646,6,"English","# Introduction\n## Part-of-speech tagging in NLP pipelines\n## Challenges unique to Twitter\n# Contributions\n## Tagset design and annotation\n## Feature development and evaluation\n## Released corpus and trained tagger\n# Tagset and annotation scheme","[{\"question\":\"What is the main goal of this work on Twitter data?\",\"answer\":\"To build an English POS tagging system designed specifically for Twitter, including a tailored tagset, annotated data, features, and evaluation results.\"},{\"question\":\"How was the Twitter dataset for POS tagging prepared?\",\"answer\":\"A POS tagset was developed and 1,827 tweets were manually annotated using an annotation scheme matched to Twitter’s characteristics.\"},{\"question\":\"Which resources and feature ideas support improved tagging accuracy?\",\"answer\":\"The feature set captures Twitter-specific properties and exploits existing resources such as tag dictionaries and phonetic normalization.\"}]","Part-of-Speech Tagging for Twitter - Annotation, Features, and Experiments | PDF"]