[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-216839-en":3,"doc-seo-216839-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},216839,2336475104736,"วิน","https://ap-avatar.wpscdn.com/avatar/22000c4c5e0e5b17e70?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786591360781797222",8,"Research & Report","EACH-USP Ensemble Cross-Domain Authorship Attribution - Notebook for PAN at CLEF 2018","Ensemble cross-domain authorship attribution is presented, combining outputs from three independent classifiers: standard character n-grams, character n-grams with non-diacritic distortion, and word n-grams. The method relies on variable-length n-gram models paired with multinomial logistic regression, selecting the author with the highest predicted probability among the three models. Experiments show performance that generally surpasses the PANCLEF 2018 baseline using fixed-length character n-grams and linear SVM classification.","EACH-USP Ensemble Cross-Domain Authorship Attribution  \nNotebook for PAN at CLEF 2018  \nJosé Eleandro Custódio and Ivandré Paraboni  \nSchool of Arts, Sciences and Humanities (EACH)  \nUniversity of São Paulo (USP)  \nSão Paulo, Brazil  \n{eleandro,[ivandre}@usp.br](ivandre}@usp.br)  \nAbstract. We present an ensemble approach to cross-domain authorship attribution that combines predictions made by three independent classiﬁers, namely, standard char n-grams, char n-grams with non-diacritic distortion and word ngrams. Our proposal relies on variable-length n-gram models and multinomial logistic regression, and selects the prediction of highest probability among the three models as the output for the task. Results generally outperform the PANCLEF 2018 baseline system that makes use of ﬁxed-length char n-grams and linear SVM classiﬁcation.  \n1 Introduction  \nAuthorship attribution (AA) is the computational task of determining the author of a given document from a number of possible candidates [1] . Systems of this kind have a wide range of possible applications, from on-line fraud detection to plagiarism and/or copyright protection. AA is presently a well-established research ﬁeld, and a recurrent topic in the PAN-CLEF shared task series [7,5] .  \nAt PAN-CLEF 2018, a cross-domain authorship attribution task applied to fan ﬁction text has been proposed. In this task, texts written by the same authors in multiple domains were put together, creating a cross-domain setting. The task consists of identifying the author of a given document based on text of a different genre.  \nThe present work describes the results of our own entry in the PAN-CLEF 2018 [2] AA shared task -hereby called the EACH-USP model -using both the baseline system and data provided by the event 1. This consists often individual AA tasks in ﬁve languages (English, French, Italian, Polish and Spanish), being two tasks (with 5 or 20 candidate authors) each.  \nThe rest of this paper is structured as follows. Section 3 describes our main AA approach, and Section 4 describes its evaluation over the PAN-CLEF 2018 AA dataset. Section 5 presents our results and those provided by relevant baseline methods. Finally, Section 6 discusses these results and suggests future work.  \n1 Available from [https://pan.webis.de/clef18/pan18-web/author-identi](https://pan.webis.de/clef18/pan18-web/author-identi)ﬁcation.html  \n2 Related Work  \nThe present work shares similarities with a number of AA studies. Some of these are brieﬂy discussed below.  \nThe work in [9] makes use of text distortion methods intended to preserve only the text structure and style in a cross-domain AA setting. The work focused on the use of word-level information, whereas our current proposal will focus on character-level information.  \nThe work in [8] investigates the role of afﬁxes in the AA task by using char n-gram models for the English language. Similarly, the work in [3] addresses the use of charn-grams models for the Portuguese language, and discusses the role of afﬁx information in the AA task. This is in principle relevant to our current work since the Portuguese language shares a great deal of its structure with Spanish and Italian, which are two of the target languages for the PAN-CLEF 2018 AA task.  \n3 Method  \nCentral to our approach is the idea that the AA task may rely on the combination of different knowledge sources such as lexical preferences, morphological inﬂection, uppercase usage, and text structure, and that different kinds of knowledge may be obtained either from character-based or word-based text models. These alternatives are discussed as follows.  \nWord or content-based models may indicate word usage preferences, and may help distinguish an author from another. However, we notice that a single author may favour certain words in different domains (e.g., ﬁctional versus dialogue text) . Moreover, wordbased models will usually discard punctuation and spaces, which may represent a valuable knowl","cbCaihEjJHtzdUlh","https://ap.wps.com/l/cbCaihEjJHtzdUlh","pdf",379307,1,7,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n# Method\n## Ensemble architecture\n## Text distortion approach\n# Evaluation","[{\"question\":\"What classifiers are combined in the EACH-USP ensemble model?\",\"answer\":\"The approach combines three classifiers: Std.charN (variable-length char n-grams), Dist.charN (variable-length char n-grams with non-diacritic distortion), and Std.wordN (variable-length word n-grams).\"},{\"question\":\"How does the method produce the final author prediction?\",\"answer\":\"Each of the three models generates a prediction, and the system outputs the prediction with the highest probability among them.\"},{\"question\":\"What improvements does the proposal make over the PAN-CLEF 2018 baseline?\",\"answer\":\"It replaces fixed-length character n-grams and linear SVM classification with variable-length n-gram models and multinomial logistic regression, and it uses an ensemble of three independent classifiers to determine the most likely author.\"}]","EACH-USP Ensemble Cross-Domain Authorship Attribution - Notebook for PAN at CLEF 2018 | PDF",1788827949,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"each-usp-ensemble-cross-domain-authorship-attribution-notebook-for-pan-at-clef-2018","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/each-usp-ensemble-cross-domain-authorship-attribution-notebook-for-pan-at-clef-2018/216839/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-08",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What classifiers are combined in the EACH-USP ensemble model?","Question",{"text":75,"@type":76},"The approach combines three classifiers: Std.charN (variable-length char n-grams), Dist.charN (variable-length char n-grams with non-diacritic distortion), and Std.wordN (variable-length word n-grams).","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the method produce the final author prediction?",{"text":80,"@type":76},"Each of the three models generates a prediction, and the system outputs the prediction with the highest probability among them.",{"name":82,"@type":73,"acceptedAnswer":83},"What improvements does the proposal make over the PAN-CLEF 2018 baseline?",{"text":84,"@type":76},"It replaces fixed-length character n-grams and linear SVM classification with variable-length n-gram models and multinomial logistic regression, and it uses an ensemble of three independent classifiers to determine the most likely author.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]