[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84042-en":3,"doc-seo-84042-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84042,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","Large Language Models Have Unreliable Understanding of Software Engineering Terminology","Large Language Models (LLMs) are increasingly used in software engineering, but their true ability to understand standardized software engineering terminology remains insufficiently studied. This work evaluates how well state-of-the-art LLMs verify definitions from ISO/IEC/IEEE 24765:2017 Vocabulary for Systems and Software Engineering. Models are prompted with correct definitions and with systematically falsified ones (semantic and structural). Classification accuracy and the coherence of generated reasoning tokens are measured, revealing a rejection bias against correct definitions despite strong detection of falsifications.","Large Language Models Have Unreliable Understanding of Software Engineering  \nTerminology  \n1st Huzaifa Ejaz Faculty of Computer Science and Mathematics University of Passau Passau, Germany [ejaz01@ads.uni-passau.de](ejaz01@ads.uni-passau.de)  \n2nd Fabian C. Pe˜na   \nFaculty of Computer Science and Mathematics University of Passau Passau, Germany [fabiancamilo.penalozano@uni-passau.de](fabiancamilo.penalozano@uni-passau.de)  \n3rd Steffen Herbold  Faculty of Computer Science and Mathematics University of Passau Passau, Germany [steffen.herbold@uni-passau.de](steffen.herbold@uni-passau.de)  \narXiv :2607 .06004v 1 [ cs . SE] 7 Jul 2026  \nAbstract—Large Language Models (LLMs) are increasingly used in software engineering (SE), yet there is no systematic study that determines to which degree these LLMs actually understand standardized SE terminology. Lack of such understanding can lead to miscommunication and misunderstanding, both by LLMs consuming text but also by human-developers acting on LLMgenerated text. Within this paper, we investigate to which degree state-of-the-art LLMs are able to identify whether definitions from the ISO/IEC/IEEE 24765:2017 Systems and Software Engineering—Vocabulary are correct. We prompt LLMs both with correct definitions, as well as systematically falsified definitions. The falsifications are both semantic (substitution of key terms) and structural (removing critical information). We measure both classification accuracy and whether reasoning tokens generated by the LLMs make sense with respect to understanding the definition. While most LLMs detect falsified definitions with high accuracy, they also reject many correct definitions, indicating a systematic rejection bias rather than genuine discriminative understanding. Explicit reasoning does not consistently improve results and may even hinder performance through over-thinking. Our work demonstrates that while the performance of LLMs (including their agentic use) in many SE tasks is impressive, there are still fundamental issues to understand how this will impact SE, including the consistent use of terminology.  \nIndex Terms—Large language models, software engineering terminology, software engineering language understanding  \nI. INTRODUCTION  \nLarge Language Models (LLMs) have become deeply integrated into the software engineering (SE) practice. For example, a systematic literature review by Hou et al. [1], analyzing 395 studies, found LLMs applied across 85 distinct SE tasks spanning the entire software development lifecycle, from requirements engineering and code generation to testing, maintenance, and documentation. The rapid adoption of these LLMs reflects their demonstrated ability to process and generate both natural language and code with increasing fluency [2], [3] . The strong and ever-increasing performance  \nof frontier LLMs in benchmarks such as SWE-bench [4] or Terminal-bench [5] notwithstanding, we still lack a principled understanding of the SE knowledge that leads to this success.  \nWithin this paper, we study one specific, low-level SE capability: understanding the correctness of definitions from the SE context. This Natural Language Understanding (NLU) task is critical, since SE is a domain that relies heavily on precisely defined concepts to facilitate communication between different actors [6] . Concepts such as verification and validation, or fault, failure, and error, are not interchangeable: each refers to a distinct phenomenon with specific implications for how software is developed, tested, and assessed. ALLM that conflates these terms, or accepts a subtly incorrect definition as valid, is also likely to misuse such terms later. Such misuse could silently lead to a multitude of issues, e.g., misunderstanding inputs and, therefore, generating faulty outputs, but also generating software artifacts that misuse terminology, possibly leading to communication issues with downstream actors (both LLM and human), that consume these","cbCaigRd6IlvqhXf","https://ap.wps.com/l/cbCaigRd6IlvqhXf","pdf",894083,5,1,11,"English","en",105,"# Introduction\n## Research Question and Hypotheses\n## Ground Truth and Definition Falsification","[{\"question\":\"What problem does the paper address about LLMs in software engineering?\",\"answer\":\"The paper addresses the lack of systematic evidence on whether LLMs truly understand standardized software engineering terminology, which can cause miscommunication by both LLMs and human developers using LLM-generated text.\"},{\"question\":\"How does the study test LLM understanding of software engineering definitions?\",\"answer\":\"It prompts LLMs with definitions from ISO/IEC/IEEE 24765:2017 and also with falsified definitions. Falsifications are created through semantic substitutions and structural removal of critical information, then evaluation covers classification accuracy and the consistency of reasoning tokens.\"},{\"question\":\"What key finding about LLM performance is reported?\",\"answer\":\"While most LLMs detect falsified definitions with high accuracy, they also reject many correct definitions, indicating a systematic rejection bias rather than genuine discriminative understanding.\"}]",1784192200,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"large-language-models-have-unreliable-understanding-of-software-engineering-terminology","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/large-language-models-have-unreliable-understanding-of-software-engineering-terminology/84042/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper address about LLMs in software engineering?","Question",{"text":76,"@type":77},"The paper addresses the lack of systematic evidence on whether LLMs truly understand standardized software engineering terminology, which can cause miscommunication by both LLMs and human developers using LLM-generated text.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the study test LLM understanding of software engineering definitions?",{"text":81,"@type":77},"It prompts LLMs with definitions from ISO/IEC/IEEE 24765:2017 and also with falsified definitions. Falsifications are created through semantic substitutions and structural removal of critical information, then evaluation covers classification accuracy and the consistency of reasoning tokens.",{"name":83,"@type":74,"acceptedAnswer":84},"What key finding about LLM performance is reported?",{"text":85,"@type":77},"While most LLMs detect falsified definitions with high accuracy, they also reject many correct definitions, indicating a systematic rejection bias rather than genuine discriminative understanding.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]