[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81735-en":3,"doc-seo-81735-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81735,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns","Large Language Models (LLMs) are increasingly used for exam- and certification-style question answering, where retrieving, interpreting, and applying domain knowledge can be evaluated in a controlled setting. In Software Engineering, such tasks matter because correctness depends on strict, normative definitions and rule-based terminology. This paper empirically compares GPT-5 mini, Gemini 3 Flash, and DeepSeek Chat 3.2 on 993 PSM I-aligned Scrum questions, testing accuracy, intra-model stability, topic/format effects, and recurring error patterns in incorrect answers.","Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns  \nRobson Alves Vilar, Emanuel Dantas Filho, Ademar França de Sousa Neto, Mirko Perkusich, Danyllo Wagner Albuquerque, João Paiva, Kyller Gorgônio, Angelo Perkusich  \nVIRTUS Research, Development and Innovation Center, Federal University of Campina Grande  \nCampina Grande, PB, Brazil  \narXiv :2607 .00048v1 [ cs . SE] 29 Jun 2026  \nAbstract  \nLarge Language Models (LLMs) are increasingly used in exam-and certification-style question answering tasks, where their ability to retrieve, interpret, and apply domain-specific knowledge can be systematically assessed. In Software Engineering, such settings are particularly relevant when questions depend on strict adherence to normative definitions, roles, artifacts, and rules. This paper evaluates the performance of three contemporary LLMs, GPT-5 mini, Gemini 3 Flash, and DeepSeek Chat 3.2, in answering 993 Scrum certification-style questions aligned with the Professional Scrum Master I (PSM I) assessment format. We evaluated the models under three prompting strategies (zero-shot, chain-of-thought, and sourcegrounded), with repeated executions to assess intra-model stability. We also analyzed performance across Scrum topics and question formats, complemented by a qualitative analysis of recurring error patterns in incorrect answers. Results revealed clear differences among models, with Gemini 3 Flash achieving the highest accuracy, followed by GPT-5 mini and DeepSeek Chat 3.2, while intra-model variability remained low across all conditions. By question format, the models achieved the highest accuracy on single-answer multiple-choice items, whereas multi-select and True/False questions were more error-prone. By topic, performance was more consistent in normatively explicit areas such as Artifacts, Empiricism, and Product Value, but more fragile in Scrum Values, Self-Managing Teams, and Stakeholders & Customers. The qualitative analysis showed that errors were systematic rather than random, involving overgeneralization, restrictive wording, compound distractors, and conflicts between common market interpretations and strict Scrum definitions.  \nKeywords  \nLarge Language Models; Scrum; Software Engineering; Empirical Evaluation; Knowledge Assessment; Error Analysis  \n1 Introduction  \nAgile Software Development has become a dominant approach to building and evolving software systems across industries. Its principles of collaboration, adaptation, and continuous improvement have shaped how software teams organize work, deliver value, and respond to change [14] . As Agile adoption has grown, so has the demand for education and certification mechanisms that support a shared understanding of Agile values, frameworks, and practices [5] . Scrum is a prominent example of this movement. Certifications such as the Professional Scrum Master (PSM) and Professional  \nScrum Product Owner are widely used to assess practitioners’ understanding of the Scrum framework and its underlying principles [13] .  \nScrum is particularly well-suited to certification-style assessment because it is grounded in a concise yet normative body of knowledge. The Scrum Guide (2020) specifies accountabilities, events, artifacts, commitments, and rules that practitioners are expected to interpret consistently. As a result, Scrum certification-style questions provide a rigorous and measurable setting for assessing domainspecific knowledge whose correctness depends not only on semantic plausibility but also on strict adherence to normative definitions, role boundaries, and rule-based terminology.  \nIn parallel, Large Language Models (LLMs) have been increasingly investigated as question-answering systems in exam-and certification-style settings, where their ability to retrieve, interpret, and apply domain-specific knowledge can be assessed in a controlled manner [9, 22] . Recent studies have examined LLM performance in r","cbCaikGXD9MKlvwG","https://ap.wps.com/l/cbCaikGXD9MKlvwG","pdf",918375,3,1,11,"English","en",105,"# Introduction\n# Methodology\n## Prompting strategies\n## Experimental setup and evaluation\n# Results\n## Accuracy and stability\n## Performance by topic and question format\n# Error pattern analysis\n## Qualitative error characterization\n# Discussion\n# Conclusion","[{\"question\":\"What question formats and Scrum topics were most error-prone?\",\"answer\":\"Single-answer multiple-choice items yielded the highest accuracy, while multi-select and True/False questions were more error-prone. Performance was more consistent in areas such as Artifacts, Empiricism, and Product Value, but less stable in Scrum Values, Self-Managing Teams, and Stakeholders \\u0026 Customers.\"}]",1784175736,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"comparing-large-language-models-on-scrum-certification-style-questions-accuracy-stability-and-error-patterns","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/comparing-large-language-models-on-scrum-certification-style-questions-accuracy-stability-and-error-patterns/81735/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What question formats and Scrum topics were most error-prone?","Question",{"text":75,"@type":76},"Single-answer multiple-choice items yielded the highest accuracy, while multi-select and True/False questions were more error-prone. Performance was more consistent in areas such as Artifacts, Empiricism, and Product Value, but less stable in Scrum Values, Self-Managing Teams, and Stakeholders & Customers.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]