[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-136006-en":3,"doc-seo-136006-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},136006,962085570644,"Melati","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Overview of the Author Identification Task at PAN-2018 Cross-domain Authorship Attribution and Style Change Detection","Authorship identification aims to reveal the authors behind texts and supports domains such as literary research, cyber-security, forensics, and social media analysis. In PAN-2018, two tasks are studied: cross-domain authorship attribution, where known and unknown authorship texts come from different domains, and style change detection, distinguishing single-author from multi-author documents. For attribution, fanfiction texts enable controlled domain study across five languages. For style change detection, a new multi-topic English Q&A dataset is introduced, alongside participant surveys and evaluation results.","Overview of the Author Identiﬁcation Task at PAN-2018 Cross-domain Authorship Attribution and Style Change Detection  \nMike Kestemont, 1 Michael Tschuggnall,2 Efstathios Stamatatos,3 Walter Daelemans, 1 Günther Specht,2 Benno Stein,4 and Martin Potthast5  \n1University of Antwerp, Belgium  \n2University of Innsbruck, Austria  \n3University of the Aegean, Greece  \n4Bauhaus-Universität Weimar, Germany  \n5Leipzig University, Germany  \n[pan@webis.de](pan@webis.de) [http://pan.webis.de](http://pan.webis.de)  \nAbstract Author identiﬁcation attempts to reveal the authors behind texts. It isan emerging area of research associated with applications in literary research, cyber-security, forensics, and social media analysis. In this edition of PAN, we study two task, the novel task of cross-domain authorship attribution, where the texts of known and unknown authorship belong to different domains, and style change detection, where single-author and multi-author texts are to be distinguished. For the former task, we make use of fanﬁction texts, a large part of contemporary ﬁction written by non-professional authors who are inspired by speciﬁc well-known works, to enable us control the domain of texts for the ﬁrst time. We describe a new corpus of fanﬁction texts covering ﬁve languages (English, French, Italian, Polish, and Spanish) . For the latter, a new data set of Q&As covering multiple topics in English is introduced. We received 11 submissions for the cross-domain authorship attribution task and 5 submissions for the style change detection task. A survey of participant methods and analytical evaluation results are presented in this paper.  \n1 Introduction  \nIn recent years, the authenticity of (online) information has attracted much attention, especially in the context of the so-called `fake news' debate in the wake of the US presidential elections. Much emphasis is currently put in various media on the provenance and authenticity of information. In the case of written documents, an important aspect of this sort of provenance criticism relates to authorship: assessing the authenticity of information crucially relates to identifying the original author(s) of these documents. Consequently, one can argue that the development of computational authorship identiﬁcation systems, that can assist humans in various tasks in this domain (journalism, law enforcement, content moderation, etc.), carries great signiﬁcance.  \nQuantitative approaches to tasks like authorship attribution [42], veriﬁcation [26], proﬁling [2] or author clustering [48] rely on the basic assumption that the writing style  \nof documents is somehow quantiﬁed, learned, and used to build prediction models. It  \nis commonly stressed that a unifying goal of the ﬁeld is to develop modeling strategies for texts that focus on style rather than content. Any successful author identiﬁcation system, be it in a attribution setup or in a veriﬁcation setup, must yield robust identiﬁcations across texts in different genres, treating different topics or having different target audiences in mind. Because of this requirement, features such as function words or common character-level n-grams are typically considered valuable characteristics, because they are less strongly tied to the speciﬁc content or genre of texts. Such features nevertheless require relatively long documents to be successful and they typically result in sparse, less useful representations for short documents. As such, one of the ﬁeld's most important goals remains the development of systems that do not overﬁt on the speciﬁc content of training texts and scale well across different text varieties.  \nThis year we focus on so-called fanﬁction, where non-professional authors produce prose ﬁction that is inspired by a well-known author or work. Many fans produce ﬁction across multiple fandoms, raising interesting questions about the stylistic continuity of these authors across these fandoms. Cross-fandom authorship attribution, whi","cbCaiiDE4r12sRbA","https://ap.wps.com/l/cbCaiiDE4r12sRbA","pdf",1059872,1,25,"English","en",105,"# Introduction\n## Authorship and provenance in online information\n## Modeling style for authorship tasks\n## Focus on fanfiction for cross-fandom attribution\n## Style breach vs. PAN 2018 style change detection\n## Building datasets for evaluation","[{\"question\":\"What are the two tasks evaluated in PAN-2018?\",\"answer\":\"PAN-2018 evaluates cross-domain authorship attribution and style change detection. The former separates authorship across different domains, while the latter distinguishes single-author from multi-author documents.\"},{\"question\":\"How does the paper address cross-domain authorship attribution?\",\"answer\":\"It uses fanfiction texts to control the domain of texts for the first time. A new fanfiction corpus is described covering five languages.\"},{\"question\":\"What data is introduced for style change detection?\",\"answer\":\"The paper introduces a new dataset of Q\\u0026As in English across multiple topics. It is created by crawling a Q\\u0026A network and applying cleaning steps to ensure realism and quality.\"}]","Overview of the Author Identification Task at PAN-2018 Cross-domain Authorship Attribution and Style Change Detection | PDF",1787337586,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"overview-of-the-author-identification-task-at-pan-2018-cross-domain-authorship-attribution-and-style-change-detection","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/overview-of-the-author-identification-task-at-pan-2018-cross-domain-authorship-attribution-and-style-change-detection/136006/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-21",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What are the two tasks evaluated in PAN-2018?","Question",{"text":75,"@type":76},"PAN-2018 evaluates cross-domain authorship attribution and style change detection. The former separates authorship across different domains, while the latter distinguishes single-author from multi-author documents.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper address cross-domain authorship attribution?",{"text":80,"@type":76},"It uses fanfiction texts to control the domain of texts for the first time. A new fanfiction corpus is described covering five languages.",{"name":82,"@type":73,"acceptedAnswer":83},"What data is introduced for style change detection?",{"text":84,"@type":76},"The paper introduces a new dataset of Q&As in English across multiple topics. It is created by crawling a Q&A network and applying cleaning steps to ensure realism and quality.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]