[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118708-en":3,"doc-seo-118708-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118708,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Translator Attribution for Arabic Using Machine Learning - Accepted Version Article","Given a set of Arabic target-language documents and their associated translators, the translator attribution task identifies which translator produced which text and estimates translator-specific style. The work builds a dedicated corpus of translations of well-known books into Arabic, then applies preprocessing such as cleaning irrelevant material, morphological segmentation, and devocalisation. Stylometric markers are derived from the 100 most frequent words and morphologically segmented function words, followed by supervised and unsupervised machine learning with a cluster-to-author index.","Please cite the Published Version  \nMohamed, Emad, Sarwar, Raheem  and Mostafa, Sayed (2023) Translator attribution for Arabic using machine learning. Digital Scholarship in the Humanities, 38 (2) . pp. 658-666. ISSN 2055- 7671  \nDOI: [https://doi.org/10.1093/llc/fqac054](https://doi.org/10.1093/llc/fqac054)  \nPublisher: Oxford University Press (OUP)  \nVersion: Accepted Version  \nDownloaded from: [https://e-space.mmu.ac.uk/630544/](https://e-space.mmu.ac.uk/630544/)  \nUsage rights:  In Copyright  \nAdditional Information: This is a pre-copyedited, author-produced version of an article accepted for publication in Digital Scholarship in the Humanities following peer review. The version of record Emad Mohamed, Raheem Sarwar, Sayed Mostafa, Translator attribution for Arabic using machine learning, Digital Scholarship in the Humanities, 2022;, fqac054 is available online at: [https://academic.oup.com/dsh/advance-article/doi/10.1093/llc/fqac054/6760698,](https://academic.oup.com/dsh/advance-article/doi/10.1093/llc/fqac054/6760698,)  \n[https://doi.org/10.1093/llc/fqac054](https://doi.org/10.1093/llc/fqac054)  \nEnquiries:  \nIf you have questions about this document, contact [openresearch@mmu.ac.uk. Please](openresearch@mmu.ac.uk. Please) include the URL of the record in e-space. If you believe that your, or a third party's rights have been compromised through this document please see our Take Down policy (available from [https://www.mmu.ac.uk/library/using-the-library/policies-and-guidelines](https://www.mmu.ac.uk/library/using-the-library/policies-and-guidelines))  \nTranslator Attribution for Arabic Using Machine  \nLearning  \nEmad Mohamed 1 , Raheem Sarwar ∗2 , and Sayed Mostafa3  \n1Research Group in Computational Linguistics, University of Wolverhampton, United Kingdom  \n2Department of Operations, Technology, Events and Hospitality Management, Manchester Metropolitan University, Manchester, M15 6BH, United Kingdom  \n3Department of Mathematics & Statistics, North Carolina A&T State University, USA  \nOctober 19, 2022  \nAbstract  \nGiven a set of target language documents and their translators, the translator attribution task aims at identifying which translator translated which documents. The attribution and the identification of the translator’s style could contribute to fields including translation studies, digital humanities, and forensic linguistics. To conduct this investigation, firstly, we develop a new corpus containing the translations of world-famous books into Arabic. We then pre-process the books in our corpus which mainly involves cleaning irrelevant material, morphological segmentation analysis of words and devocalisation. After pre-processing the books, we propose to use 100 most frequent words and/or morphologically segmented function words as writing style markers of the translators (i.e., stylometric features) to differentiate between translations of different translators. After the completion of features extraction process, we applied several supervised and unsupervised machine learning algorithms along with our novel cluster-to-author index to perform this task. We found that the translators are not invisible, and morphological analysis  \n∗ Corresponding Author ([raheem.bwl@gmail.com](raheem.bwl@gmail.com))  \nmay not be more useful than just using the 100 most frequent words as features. The SVM Linear Kernel algorithm reported 99% classification accuracy. Similar findings were reported by the unsupervised machine learning methods, namely, K-Mean Clustering and Hierarchical Clustering.  \n1 Introduction  \nAuthorship Attribution is a sub-discipline of Computational Linguistics that tries to determine whether an anonymous document was written by one of several candidate authors. The field maintains that each author has a specific linguistic fingerprint (or author print) and that computational tools can discover this unique style. Authorship attribution is usually performed by extracting writing style features from the t","cbCaiqGj0O0MXstI","https://ap.wps.com/l/cbCaiqGj0O0MXstI","pdf",380538,1,17,"English","en",105,"# Abstract\n# Introduction\n## Authorship vs. translation attribution\n## Translator invisibility and motivation","[{\"question\":\"What problem does the paper address?\",\"answer\":\"The paper studies translator attribution: given target-language documents and candidate translators, it identifies which translator produced which document and captures translator style.\"},{\"question\":\"How is the dataset built and preprocessed?\",\"answer\":\"It constructs a new corpus of world-famous books translated into Arabic, then preprocesses texts by cleaning irrelevant material, performing morphological segmentation, and applying devocalisation.\"},{\"question\":\"Which features and machine learning methods are used?\",\"answer\":\"It uses the 100 most frequent words and/or morphologically segmented function words as stylometric features, then applies supervised and unsupervised machine learning, including SVM (linear kernel) and clustering methods.\"}]","Translator Attribution for Arabic Using Machine Learning - Accepted Version Article | PDF",1785719850,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"translator-attribution-for-arabic-using-machine-learning-accepted-version-article","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/translator-attribution-for-arabic-using-machine-learning-accepted-version-article/118708/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04","2026-08-03",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper address?","Question",{"text":76,"@type":77},"The paper studies translator attribution: given target-language documents and candidate translators, it identifies which translator produced which document and captures translator style.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is the dataset built and preprocessed?",{"text":81,"@type":77},"It constructs a new corpus of world-famous books translated into Arabic, then preprocesses texts by cleaning irrelevant material, performing morphological segmentation, and applying devocalisation.",{"name":83,"@type":74,"acceptedAnswer":84},"Which features and machine learning methods are used?",{"text":85,"@type":77},"It uses the 100 most frequent words and/or morphologically segmented function words as stylometric features, then applies supervised and unsupervised machine learning, including SVM (linear kernel) and clustering methods.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]