[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-156697-en":3,"doc-seo-156697-105":30,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},156697,687207024478,"Mia  ","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",4,"Exam","Effects of Generation Model on Detecting AI-generated Essays in a Writing Test","Various detectors have been developed to identify AI-generated essays using labeled datasets of human-written and AI-generated responses, often reporting strong performance. In practical testing, however, essays may be produced by different generation models than those used for detector training. This study evaluates how generation model choice affects detector performance, comparing GPT-3.5 and GPT-4 generation on standardized English writing items and multiple detector architectures and training strategies.","Effects of Generation Model on Detecting AI-generated Essays in a Writing  \nTest  \nJiyun Zu and Michael Fauss and Chen Li  \nEducational Testing Service, Princeton, NJ  \nCorrespondence: [jzu@ets.org](jzu@ets.org)  \nAbstract  \nVarious detectors have been developed to detect AI-generated essays using labeled datasets of human-written and AI-generated essays, with many reporting high detection accuracy. In real-world settings, essays may be generated by models different from those used to train the detectors. This study examined the effects of generation model on detector performance. We focused on two generation models – GPT- 3.5 and GPT-4 – and used writing items from a standardized English proficiency test. Eight detectors were built and evaluated. Six were trained on three training sets (human-written essays combined with either GPT-3.5-generated essays, or GPT-4-generated essays, or both) using two training approaches (feature-based machine learning and fine-tuning RoBERTa), and the remaining two were ensembled detectors. Results showed that a) fine-tuned detectors outperformed feature-based machine learning detectors on all studied metrics; b) detectors trained with essays generated from only one model were more likely to misclassify essays generated by the other model as human-written essays (false negatives), but did not misclassify more human-written essays as AI-generated (false positives); c) the ensembled fine-tuned RoBERTa detector had fewer false positives, but slightly more false negatives than detectors trained with essays generated by both models.  \n1 Introduction  \nGenerative artificial intelligence (AI) tools, such as ChatGPT, Copilot, and Gemini, have become increasingly capable at generating human-like text and are now more accessible. In education, AI has great potential at enhancing teaching, learning, and assessments (U.S. Department of Education, Office of Educational Technology, 2023) . At the sametime, there are also concerns about the misuse of AI in writing tasks (Lund et al., 2025) . Writing assignments are routinely given in both K-12 and  \nhigher education. There are also many standardized writing tests designed to measure test takers writing proficiency, such as the ACT writing test, the Graduate Record Examinations (GRE) writing test, and the Writing assessment program (WrAP) for grades 3-12 students. These tests require test takers to write essays independently. If some test takers use generative AI tools to write essays and use these essays as their own, the validity and fairness of the writing assessment are compromised.  \nTo address concerns about AI-generated text, many detectors have been developed to identify such content. For example, Grammarly (Grammarly Inc., 2025), Scribbr (Scribbr, 2025), and GPTZero (GPTZero, 2025) provide online tools that allow users to enter text and then output an estimated percentage of the text being AI-generated, although documentation on how they were trained is generally unpublished. Several research studies reported the training and evaluation of custom-built AI-generated essay detectors. For example, Yanet al. (2023) generated essays using GPT-3 for four writing items from a large-scale assessment. Using these essays and real human test takers’ essays, the authors trained two detectors: one using supervised machine learning (ML) approach and the other by fine-tuning the pre-trained language model RoBERTa (Liu et al., 2019) . The detection accuracy on a holdout test set was respectively 96% and 99.75% for these two detectors. Jiang et al. (2024) studied the accuracy and potential bias in detecting ChatGPT-generated essays. Using 10,000 essays generated by ChatGPT and 10,000 essays written by real test takers for 50 GRE writing items, the authors trained detectors using supervised ML with linguistic features extracted by e-rater (Attali and Burstein, 2006) and GPT-2-based perplexity features. Detection accuracy of the best performing detector was nearly 100% ","cbCaiqaZRKqwKoYi","https://ap.wps.com/l/cbCaiqaZRKqwKoYi","pdf",217632,1,7,"English","en",105,"# Abstract\n# 1 Introduction\n## Background on AI writing tools and assessment validity\n## Existing detectors and prior research\n## Research motivation: cross-model detector generalization\n## Study design and objectives","[{\"question\":\"Why might AI-essay detectors perform differently in real-world writing tests?\",\"answer\":\"Because test-takers may generate essays with models different from the ones used to train the detectors, changing the linguistic patterns the detectors learned.\"},{\"question\":\"How did the study evaluate the effect of generation models on detection?\",\"answer\":\"It focused on essays generated by GPT-3.5 and GPT-4, using writing items from a standardized English proficiency test, and evaluated multiple detectors trained and ensembled with different strategies.\"},{\"question\":\"What were the main findings about detector training approaches?\",\"answer\":\"Fine-tuned RoBERTa-based detectors outperformed feature-based machine learning detectors, and detectors trained on a single generation model tended to produce more false negatives for essays from the other model.\"}]","Effects of Generation Model on Detecting AI-generated Essays in a Writing Test | PDF",1787965076,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":28},"effects-of-generation-model-on-detecting-ai-generated-essays-in-a-writing-test","",{"@graph":36,"@context":84},[37,53,67],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/exam/",3,{"item":52,"name":13,"@type":43,"position":11},"https://docshare.wps.com/document/effects-of-generation-model-on-detecting-ai-generated-essays-in-a-writing-test/156697/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-29",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why might AI-essay detectors perform differently in real-world writing tests?","Question",{"text":74,"@type":75},"Because test-takers may generate essays with models different from the ones used to train the detectors, changing the linguistic patterns the detectors learned.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How did the study evaluate the effect of generation models on detection?",{"text":79,"@type":75},"It focused on essays generated by GPT-3.5 and GPT-4, using writing items from a standardized English proficiency test, and evaluated multiple detectors trained and ensembled with different strategies.",{"name":81,"@type":72,"acceptedAnswer":82},"What were the main findings about detector training approaches?",{"text":83,"@type":75},"Fine-tuned RoBERTa-based detectors outperformed feature-based machine learning detectors, and detectors trained on a single generation model tended to produce more false negatives for essays from the other model.","https://schema.org",{"og:url":52,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,103,108,113,117,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":101,"slug":102},70,"exam",{"id":104,"doc_module":4,"doc_module_name":46,"category_name":105,"show_sort_weight":106,"slug":107},5,"Comic",60,"comic",{"id":109,"doc_module":4,"doc_module_name":46,"category_name":110,"show_sort_weight":111,"slug":112},6,"Technology",50,"technology",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":114,"show_sort_weight":115,"slug":116},"Healthcare",40,"healthcare",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":119,"show_sort_weight":120,"slug":121},8,"Research & Report",30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":104,"slug":137},19,"General","general"]