[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81736-en":3,"doc-seo-81736-105":30,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81736,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Prompting GPT-5 on Scrum Certification Questions An Empirical Accuracy Study","Large Language Models are increasingly used in Agile software development for documentation, coaching, and training, but reliability on normative topics remains uncertain. This study evaluates how prompt techniques change factual accuracy for GPT-5 answers to Professional Scrum Master (PSM)-aligned certification questions. A dataset of 993 validated items is answered with zero-shot, chain-of-thought, and source-citation prompting. All methods exceed 85% accuracy; citation prompting performs best at 89.1%, with errors clustering around Scrum Guide misalignment and out-of-scope or outdated interpretations.","Prompting GPT-5 on Scrum Certification Questions:  \nAn Empirical Accuracy Study  \nMirko Perkusich2 , Danyllo Albuquerque2 , Joo Paiva2 , Robson Vilar2 , Emanuel Dantas2 , Ademar Franc¸a de Sousa Neto2 , Rohit Gheyi 1 , Kyller Gorgnio2 , and Angelo Perkusich2  \n1 Federal University of Campina Grande, Brazil  \n2 ISE Group, VIRTUS/UFCG, Campina Grande, Brazil  \nCorresponding author: [emanuel.dantas@virtus.ufcg.edu.br](emanuel.dantas@virtus.ufcg.edu.br)  \narXiv :2607 .00049v1 [ cs . SE] 29 Jun 2026  \nAbstract—Large Language Models (LLMs) are increasingly used in Agile Software Development for documentation, coaching, and training. As practitioners adopt these tools to prepare for certifications such as Professional Scrum Master (PSM), a key question is whether LLMs can reliably reason about Scrum, a framework with normative, well-defined rules described in the Scrum Guide (2020). This paper examines how different prompt techniques affect the factual accuracy of LLM responses to Scrum certification-style questions. A dataset of 993 validated PSM-aligned questions was answered by GPT-5 using three techniques: zero-shot, chain-of-thought, and with-source citation. All prompts achieved certification-level accuracy above 85%, with the citation-based variant performing best (89.1%) and yielding the lowest error rate. Correct answers concentrated in well-defined topics, such as Definition of Done, Events, and Product Backlog Management, and in single-answer multiplechoice items, while multi-select questions and more interpretive areas, such as Scrum Team and Product Value, were less stable. Among questions where at least one prompt failed (16.2%), errors clustered into misalignment with the Scrum Guide (28%), content outside its scope (34%), and outdated or biased interpretations (38%). Overall, prompt techniques produced modest but consistent improvements, particularly in reducing misinterpretation and version drift, supporting more reliable use of LLMs in Agile learning and certification preparation.  \nIndex Terms—Scrum, large language models, prompt engineering, certification assessment, and empirical study  \nI. INTRODUCTION  \nAgile Software Development has become the dominant approach to building and evolving software systems across industries [1] . Its principles of collaboration, adaptability, and continuous improvement have shaped how teams design, build, and deliver software [2] . As adoption has grown, so has the demand for formal education and certification programs that ensure a shared understanding of Agile values and frameworks [3] . Certifications such as Professional Scrum Master (PSM) and Professional Scrum Product Owner (PSPO) from [Scrum.org](Scrum.org) are widely used to assess practitioners’ mastery of the Scrum framework and its underlying principles [4] .  \nScrum, as codified in the Scrum Guide (2020), defines normative and unambiguous rules governing accountabilities, events, and artifacts. This clarity makes Scrum-based certification exams a rigorous and measurable context for assessing knowledge consistency and factual understanding. Yet, learning and preparing for such certifications remains challenging, particularly for newcomers who lack access to experienced coaches or mentoring communities [4] .  \nRecently, Large Language Models (LLMs), such as ChatGPT, Gemini, and DeepSeek, have emerged as promising tools to support education and training [5], [6] . They can generate explanations, quizzes, and feedback in natural language, potentially democratizing access to learning support. However, the reliability of their responses, especially when applied to normative, rule-based content such as Scrum, remains poorly understood. LLMs are known to produce plausible yet inaccurate statements, a phenomenon often referred to as“hallucination,” and their outputs are highly sensitive to the way prompts are formulated [7], [8] . These limitations raise important concerns when such models are used in training orevaluative cont","cbCaienwJV52QbjF","https://ap.wps.com/l/cbCaienwJV52QbjF","pdf",206697,5,1,6,"English","en",105,"# Introduction\n## Research aim and contributions\n## Prompt techniques and evaluation design","[{\"question\":\"Where do errors tend to cluster when at least one prompt fails?\",\"answer\":\"Errors cluster around misalignment with the Scrum Guide (28%), content outside its scope (34%), and outdated or biased interpretations (38%).\"}]",1784175736,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":28},"prompting-gpt-5-on-scrum-certification-questions-an-empirical-accuracy-study","",{"@graph":36,"@context":78},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/prompting-gpt-5-on-scrum-certification-questions-an-empirical-accuracy-study/81736/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"Where do errors tend to cluster when at least one prompt fails?","Question",{"text":76,"@type":77},"Errors cluster around misalignment with the Scrum Guide (28%), content outside its scope (34%), and outdated or biased interpretations (38%).","Answer","https://schema.org",{"og:url":52,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":85},[86,90,94,98,102,106,111,114,119,122,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":20,"slug":129},19,"General","general"]