[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120697-en":3,"doc-seo-120697-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120697,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1785132997149421697",8,"Research & Report","Distinguishing Human Generated Text From ChatGPT Generated Text Using Machine Learning - Research Summary","ChatGPT生成文本与人类撰写文本高度相似，带来误导新闻、信息操纵及法律与伦理风险。本研究提出一种基于机器学习的检测方法，利用TF-IDF向量化与极端随机树分类器，将模型对ChatGPT文本进行识别，并对共计11种机器学习与深度学习算法进行对比评估。在Kaggle含10,000条样本的数据集上，模型在GPT-3.5语料中取得77%的准确率，为文本来源鉴别提供可行技术路径。","Distinguishing Human Generated Text From ChatGPT Generated Text Using Machine Learning  \nNiful Islam 1 , Debopom Sutradhar1 , Humaira Noor1 , Jarin Tasnim Raya2 , Monowara Tabassum Maisha2 , Dewan Md Farid 1  \n1Department of CSE, United International University (UIU), Bangladesh  \n2Department of CSE, University of Asia Pacific (UAP), Bangladesh  \nEmail : {nislam201057, [dsutradhar201046](dsutradhar201046}@bscse.uiu.ac.bd)[}](dsutradhar201046}@bscse.uiu.ac.bd)[@bscse.uiu.ac.bd](dsutradhar201046}@bscse.uiu.ac.bd), [hnoor222007@mscse.uiu.ac.bd](hnoor222007@mscse.uiu.ac.bd),  \n{20101002, [20101001](20101001}@uap-bd.edu)[}](20101001}@uap-bd.edu)[@uap-bd.edu](20101001}@uap-bd.edu), [dewanfarid@cse.uiu.ac.bd](dewanfarid@cse.uiu.ac.bd)  \narXiv :2306 .01761v1 [ cs .CL] 26 May 2023  \nAbstract—ChatGPT is a conversational artificial intelligence that is a member of the generative pre-trained transformer of the large language model family. This text generative model was finetuned by both supervised learning and reinforcement learning so that it can produce text documents that seem to be written by natural intelligence. Although there are numerous advantages of this generative model, it comes with some reasonable concerns as well. This paper presents a machine learning-based solution that can identify the ChatGPT delivered text from the human written text along with the comparative analysis of a total of 11 machine learning and deep learning algorithms in the classification process. We have tested the proposed model on a Kaggle dataset consisting of 10,000 texts out of which 5,204 texts were written by humans and collected from news and social media. On the corpus generated by GPT-3.5, the proposed algorithm presents an accuracy of 77% .  \nIndex Terms—ChatCPT, Classification, Generative AI, NLP, Tokenization  \nI. INTRODUCTION  \nThe emergence of generative AI models are rapidly changing our way of communication. It is being extensively used in content creation, arts and design, healthcare and many more. Although these models, especially conversational AI models like ChatGPT, have the potential of revolutionizing the society, they also come with some possible dangers. Oneof the biggest concerns is that, it can produce false news or spread misinformation [1] [2] . Since, the AI-generated texts are almost identical to the human-generated texts, the model can be used to manipulate individuals or organizations in various ways. There are some legal and ethical concerns of using generative AI. Since these models are trained on large datasets, there might remain some bias in a particular sector. For decision making, if an individual employs these models, it may lead to discriminatory attitudes. Furthermore, students may rely too heavily on the AI generated tools, which could damage their critical thinking and communication ability, negatively impacting their academic and professional life [3] . For newly developed problems, it requires new solutions. As these models are trained on old data, asking solutions for advanced challenges might provide misleading solutions [4] . In addition, incorporating conversational AI into the system could result in low user satisfaction. Therefore, it requires a  \nsystem for identifying human generated text and AI generated text.  \nNatural Language Processing (NLP) is a rapidly growing field of study that works on understanding human language. NLP gives machines the ability to learn human language by turning it into numerical data [5] . With the increasing number of digital texts, the need of NLP is growing rapidly [6] . In recent years, NlP has provided a large scale analysis and management of text data making sentiment analysis, emotion detection and other complicated tasks possible. Furthermore, with the help of NLP, it is possible to detect mental illness at early stage and provide treatment [7] . Previously the training of NLP models were slow and inefficient [8] . Especially after 2017, the innovation of transfo","cbCaii1jAizlTkh4","https://ap.wps.com/l/cbCaii1jAizlTkh4","pdf",617937,1,6,"English","en",105,"# Abstract\n# Introduction\n## Generative AI 与风险概述\n## NLP与Transformer相关背景\n# Proposed Detection Approach\n## TF-IDF向量化\n## 极端随机树分类器\n## 算法对比与数据预处理影响","[{\"question\":\"本文要解决的核心问题是什么？\",\"answer\":\"如何区分ChatGPT生成文本与人类撰写文本，降低误导信息与相关风险。\"},{\"question\":\"模型采用了哪些主要技术流程？\",\"answer\":\"使用TF-IDF向量化句子，再用极端随机树分类器进行文本分类。\"},{\"question\":\"实验使用的数据集规模与表现如何？\",\"answer\":\"在Kaggle数据集上共10,000条文本（人类5,204条），在GPT-3.5生成语料上准确率达到77%。\"}]","Distinguishing Human Generated Text From ChatGPT Generated Text Using Machine Learning - Research Summary | PDF",1785731595,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"distinguishing-human-generated-text-from-chatgpt-generated-text-using-machine-learning-research-summary","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/distinguishing-human-generated-text-from-chatgpt-generated-text-using-machine-learning-research-summary/120697/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"本文要解决的核心问题是什么？","Question",{"text":75,"@type":76},"如何区分ChatGPT生成文本与人类撰写文本，降低误导信息与相关风险。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"模型采用了哪些主要技术流程？",{"text":80,"@type":76},"使用TF-IDF向量化句子，再用极端随机树分类器进行文本分类。",{"name":82,"@type":73,"acceptedAnswer":83},"实验使用的数据集规模与表现如何？",{"text":84,"@type":76},"在Kaggle数据集上共10,000条文本（人类5,204条），在GPT-3.5生成语料上准确率达到77%。","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]