[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84784-en":3,"doc-seo-84784-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84784,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","LLM-Based Test Oracles: Source-of-Authority Taxonomy—A Systematic Literature Review","Large language models increasingly generate test oracles—components that judge whether observed behavior is correct—yet a clear account of where their authority comes from remains missing. This work presents a PRISMA 2020-guided systematic literature review, screening 2,436 records down to 54 included studies. It analyzes authority sources, oracle forms, and adjudication mechanisms across domains, languages, models, and adaptation strategies. Specification-derived authority covers about half the studies, while others yield verdicts without specifications, and failure modes are examined.","LLM-Based Test Oracles: Source-of-Authority Taxonomy—A Systematic Literature Review  \nAli Hassaan Mughal iD , Member, IEEE, and Muhammad Bilal iD  \narXiv :2607 .0503 1v 1 [ cs . SE] 6 Jul 2026  \nAbstract—Large language models (LLMs) are increasingly used to produce test oracles, the part of a test that decides whether observed behavior is correct. Yet a clear account of where these oracles draw their authority is missing. Prior secondary studies organize the area by oracle form or by LLM technique. None organizes it by the source of the verdict’s authority, the property that governs how far a verdict can be trusted. This article presents a systematic literature review, conducted and reported under the PRISMA 2020 guidelines. From 2,436 records, an LLM pre-filter followed by independent dual human screening (reviewer agreement, a Cohen’s kappa of 0.79) and full-text assessment yielded 54 included studies. We analyze these along three axes: the source of an oracle’s authority, the form it takes, and the mechanism that adjudicates it. We characterize the landscape of domains, languages, models, and adaptation strategies. Specification-derived authority, though the most common single source, covers about half of the studies (28 of 54). The remaining 26 reach a verdict with no specification at all. The source of authority and the adjudication mechanism cross-cut: the same source is checked by several mechanisms and one mechanism serves several sources, so a label such as LLMas-a-judge names a mechanism rather than a basis for trust. We further report how these oracles are evaluated and how they fail, and read the sparse and empty regions of the taxonomy asa research agenda. The protocol, search query, and per-study coding sheet are released as supplementary material.  \nIndex Terms—Large language models, LLM-as-a-judge, oracle problem, software quality assurance, software testing, systematic literature review, test automation, test oracle.  \nI. INTRODUCTION  \nSOFTWARE now runs in almost every part of daily life,  \nand how much we can trust it depends largely on how well we test it. A test feeds an input to a program, runs it, and then asks one question: is the result correct? The part of the test that answers that question is the test oracle. A test is only as good as its oracle.  \nDeciding what counts as correct is hard. For many programs there is no easy way to know the right answer in advance. This difficulty has a name, the oracle problem. Barr et al. [1] surveyed it before large language models (LLMs) were in wide use. They organized the known ways to obtain an oracle and showed how often a strong oracle is simply missing.  \nThis work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible.  \nThis work received no external funding. The authors declare no conflicts of interest.  \nA. H. Mughal is an Independent Researcher, USA (e-mail: alihassaan[mughal.work@gmail.com](mughal.work@gmail.com)).  \nM. Bilal is with the Technical University of Munich, Germany (e-mail: [m.bilal@tum.de](m.bilal@tum.de)). He is the corresponding author.  \nLLMs have changed the picture. An LLM can now write an oracle for a program [2]–[4], or act as the oracle itself by judging the output directly [5]–[7] . This is a fast-growing area [8], [9]: most of the reviewed studies appeared in 2025 or later.  \nThe literature describes these new oracles. It characterizes them, however, by a feature that does not, by itself, capture how far to trust the verdict. Papers sort oracles by their shape, such as an assertion [3] or a metamorphic relation [10], or by the technique used, such as “LLM as a judge” [5], [7] . They rarely ask where the verdict’s authority comes from, yet that is what sets the ceiling on how much to trust it. Two oracles can look identical yet rest on different ground. One assertion may encode a written specification; another, with the same syntax,","cbCaif6nRBQMrujA","https://ap.wps.com/l/cbCaif6nRBQMrujA","pdf",593326,3,1,15,"English","en",105,"# Introduction\n# Background: The Oracle Problem in Plain Terms\n## What a Test Oracle Is\n## LLMs and Oracle Authority","[{\"question\":\"What problem does the review address about LLM-generated test oracles?\",\"answer\":\"The review targets the missing explanation of where an LLM test oracle’s verdict authority comes from and how far that authority can be trusted. It argues that identical-looking oracles can rely on different foundations for correctness.\"},{\"question\":\"How was the systematic literature review conducted?\",\"answer\":\"The study followed PRISMA 2020 guidelines. It screened 2,436 records using an LLM pre-filter and independent dual human screening, resulting in 54 included studies after full-text assessment.\"},{\"question\":\"What does the taxonomy reveal about the source and adjudication mechanism of oracle verdicts?\",\"answer\":\"It shows that specification-derived authority is the most common source (28 of 54 studies), while 26 verdicts reach a decision without specification. The source of authority and the adjudication mechanism cross-cut, with the same source checked by multiple mechanisms and one mechanism serving multiple sources.\"}]",1784198213,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"llm-based-test-oracles-source-of-authority-taxonomya-systematic-literature-review","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/llm-based-test-oracles-source-of-authority-taxonomya-systematic-literature-review/84784/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the review address about LLM-generated test oracles?","Question",{"text":75,"@type":76},"The review targets the missing explanation of where an LLM test oracle’s verdict authority comes from and how far that authority can be trusted. It argues that identical-looking oracles can rely on different foundations for correctness.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How was the systematic literature review conducted?",{"text":80,"@type":76},"The study followed PRISMA 2020 guidelines. It screened 2,436 records using an LLM pre-filter and independent dual human screening, resulting in 54 included studies after full-text assessment.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the taxonomy reveal about the source and adjudication mechanism of oracle verdicts?",{"text":84,"@type":76},"It shows that specification-derived authority is the most common source (28 of 54 studies), while 26 verdicts reach a decision without specification. The source of authority and the adjudication mechanism cross-cut, with the same source checked by multiple mechanisms and one mechanism serving multiple sources.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]