[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86558-en":3,"doc-seo-86558-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86558,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Fail-Aware and Explainable Test Oracle Prediction","Fault detection depends on test oracles, yet effective oracle construction remains difficult. While learning-based methods can generate assertions, syntactic correctness often fails to expose real bugs. FOCAL instead trains a model to predict whether a test prefix passes or fails, using labeled pairs of prefixes and methods. It applies fail-emphasizing losses and produces statement-level behavioral evidence. Compared with SEER, FOCAL improves failure-case performance on unseen projects and adds grounded explanations validated by behavioral checks, supporting failure-aware discriminative oracle prediction as a complement to existing test generation approaches.","Fail-Aware and Explainable Test Oracle Prediction  \nYue Zhao, Binish Tanveer, Jelena Zdravkovic  \nDepartment of Computer and Systems Sciences, Stockholm University, Stockholm, Sweden {yue.zhao, binish.tanveer, [jelenaz](jelenaz}@dsv.su.se)[}](jelenaz}@dsv.su.se)[@dsv.su.se](jelenaz}@dsv.su.se)  \narXiv :2607 . 1 1342v 1 [ cs . SE] 13 Jul 2026  \nAbstract—Despite their central role in fault detection, test oracles remain challenging to construct effectively. Recent learningbased methods address this challenge by automatically generating test assertions, yet even if syntactically correct, they are often ineffective in revealing bugs. Rather than generating assertions, this study explores a different approach by training a model to directly predict whether a given test preﬁx passes or fails.  \nWe present FOCAL, an emerging code LLM-based discriminative oracle predictor. It learns from labeled pairs of test preﬁxesand methods under test, employs losses that emphasize failing cases during training, and grounds its predictions in statementlevel behavioral evidence. Compared with the baseline method SEER, we substantially improve performance on failing cases for unseen projects and provide richer explanations.  \nA preliminary evaluation on fault-detection benchmarks and automated test-generation artifacts shows that our approach is highly accurate within its training distribution and substantially improves failure detection on previously unseen projects where prior discriminative oracles collapse. Moreover, the highlighted statements are supported by behavioral explanation checks. These early results suggest that fail-aware discriminative oracle prediction can complement existing approaches such as fuzzing, search-based testing, and LLM-based test generation. These techniques produce test preﬁxes at scale but often lack faultoriented oracles. In future work, FOCAL could take generated test preﬁxes and attach fault-aware predicted oracles to them, turning high-volume input generation into executable tests that are more likely to expose semantic failures.  \nIndex Terms—Unit testing, test oracle problem, code LLM, discriminative oracle prediction, AI4SE.  \nI. INTRODUCTION  \nAs a cornerstone of modern software quality assurance, unit testing is widely used to check the correctness of individual program units and provide fast, localized feedback during development [1], [2] . A unit test typically combines a test preﬁx with a test oracle. The preﬁx executes the setup and stimulus code for the Method Under Test (MUT), and the test oracle determines whether the observed behavior conforms to the required speciﬁcation [3] . Automated test preﬁx generation has become practical through tools such as EvoSuite [4] and Randoop [5] . In contrast, test oracle construction remains oneof the most persistent and costly challenges, widely recognized as the oracle problem [6], [7] .  \nAutomated unit testing constructs oracles through different strategies. Fuzzing tools [8] use runtime signals such as crashes, hangs, and sanitizer violations as implicit oracles to capture visible runtime failures. Search-based tools record observed behavior as regression assertions, which detect future behavioral drift and take current behavior as the reference [9] . Neural-based assertion generation tools [10], [11] train models  \nto synthesize assert statements that emphasize textual similarity to developer-written assertions. LLM-based and hybrid approaches [12] produce readable oracles following common patterns. More precise judgments can be obtained when speciﬁcations, contracts, or reference implementations are available and reliable [13], but such conditions are uncommon in many real-world projects. These approaches provide useful oracle signals, but they still leave a key gap in our target setting. Given a test preﬁx without an oracle and an MUT, there is often no explicit judgment of whether the observed behavior reveals a fault.  \nSEER [14], a new oracle constructi","cbCaij8l2kqVQfgS","https://ap.wps.com/l/cbCaij8l2kqVQfgS","pdf",338788,5,1,6,"English","en",105,"# Introduction\n## Background: unit testing and the oracle problem\n## Automated oracle construction approaches\n## Fail-aware evaluation motivation\n## Contributions and method overview","[{\"question\":\"What problem does the paper address in unit testing?\",\"answer\":\"The paper targets the difficulty of constructing effective test oracles, a long-standing and costly issue known as the oracle problem. It emphasizes that many learned approaches may not translate aggregate accuracy into meaningful failure detection.\"},{\"question\":\"How does FOCAL differ from assertion-generation methods and from SEER?\",\"answer\":\"FOCAL predicts PASS/FAIL for a test prefix and the method under test directly, instead of generating assertion statements. It is fail-aware and focuses training on failing cases, whereas SEER’s overall accuracy does not provide strong sensitivity to failures on unseen projects.\"},{\"question\":\"How are FOCAL’s predictions explained and validated?\",\"answer\":\"FOCAL provides statement-level explanations for FAIL predictions using Integrated Gradients. It evaluates the highlighted evidence with behavioral checks such as deletion, keep-only, and counterfactual tests to ground the explanation in observable behavior.\"}]",1784212621,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"fail-aware-and-explainable-test-oracle-prediction","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/fail-aware-and-explainable-test-oracle-prediction/86558/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper address in unit testing?","Question",{"text":76,"@type":77},"The paper targets the difficulty of constructing effective test oracles, a long-standing and costly issue known as the oracle problem. It emphasizes that many learned approaches may not translate aggregate accuracy into meaningful failure detection.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does FOCAL differ from assertion-generation methods and from SEER?",{"text":81,"@type":77},"FOCAL predicts PASS/FAIL for a test prefix and the method under test directly, instead of generating assertion statements. It is fail-aware and focuses training on failing cases, whereas SEER’s overall accuracy does not provide strong sensitivity to failures on unseen projects.",{"name":83,"@type":74,"acceptedAnswer":84},"How are FOCAL’s predictions explained and validated?",{"text":85,"@type":77},"FOCAL provides statement-level explanations for FAIL predictions using Integrated Gradients. It evaluates the highlighted evidence with behavioral checks such as deletion, keep-only, and counterfactual tests to ground the explanation in observable behavior.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]