[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120462-en":3,"doc-seo-120462-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120462,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Specifying and Fuzzing Machine-Learning Models - Doctoral Thesis","Machine Learning (ML) models increasingly operate in safety-critical settings where failures can cause catastrophic consequences. While prior work targets properties such as robustness and fairness, specifying and checking general functional-correctness for ML models remains difficult. This thesis proposes testing-inspired tools and methods to specify and fuzz ML artifacts for functional correctness, including techniques for testing action-policy reliability, metamorphic-relations-based test oracles, and the NOMOS specification language with automated testing frameworks across multiple ML domains.","Specifying and Fuzzing Machine-Learning Models  \nThesis approved by the Department of Computer Science University of Kaiserslautern-Landau for the award of the Doctoral Degree Doctor of Engineering (Dr.-Ing.)  \nto  \nHasan Ferit Enis¸er  \nDate of Defense: Dean:  \nReviewers:  \n13 / 06 / 2025  \nProf. Dr. Christoph Garth Prof. Dr. Maria Christakis Prof. Dr. Rupak Majumdar  \nProf. Dr. Joerg Hoffmann  \nDE-386  \nTo my children...  \nAbstract  \nMachine Learning (ML) models are increasingly prevalent in safety-critical systems, from self-driving cars to aviation, where failures can result in catastrophic outcomes. While researchers have addressed speciﬁc properties like robustness and fairness, specifying and checking general functional-correctness expectations from ML models remains challenging. This thesis introduces novel tools and approaches inspired by software testing concepts to specify and fuzz ML artifacts for their functional correctness. Software testing identiﬁes bugs by running programs with given inputs, facing two main challenges: generating test inputs and ﬁnding test oracles. Fuzzing is a widely adopted method for generating test inputs, while speciﬁcations address the oracle problem. These techniques and concepts have proven effective and crucial for validating software reliability. We tailor these methods to assess ML model reliability. One of the biggest recent advancements in machine learning has been in solving sequential decision-making problems where agents learn action policies. In this thesis, we devise techniques to test action policies’ reliability. Beyond checking if policies lead to undesirable outcomes, we address: how can we identify undesirable yet avoidable outcomes? We present novel test oracles based on metamorphic relations and develop the ! →fuzz framework to identify bugs in action policies. Next, we formalize metamorphic relations as k-safety properties, or hyperproperties, describing relationships between multiple input-output pairs. We show hyperproperties can specify functional correctness across various ML domains. To express these, we create NOMOS, a declarative, domain-agnostic speciﬁcation language with an automated testing framework. We demonstrate its effectiveness in ﬁnding bugs across various domains including image classiﬁcation, sentiment analysis, and speech recognition. We also extend NOMOS to accommodate code translation models. Overall, this thesis contributes to the ﬁeld by providing a speciﬁcation language and novel automated testing frameworks to validate the reliability and safety of ML models which are now prevalent in our daily lives.  \nKurzfassung  \nMachine Learning (ML) Modelle sind zunehmend in sicherheitskritischen Systemenverbreitet, von selbstfahrenden Autos bis zur Luftfahrt, wo Ausflle zu katastrophalen Folgen f¨uhren knnen. Whrend Forscher speziﬁsche Eigenschaften wie Robustheit und Fairness behandelt haben, bleibt die Speziﬁkation und ¨Uberpr¨ufung allgemeiner funktionaler Korrektheitserwartungen von ML-Modellen eine Herausforderung. Diese Arbeit stellt neuartige Werkzeuge und Anstze vor, die von Konzepten des SoftwareTestens inspiriert sind, um ML-Artefakte bez¨uglich ihrer funktionalen Korrektheit zuspeziﬁzieren und zu fuzzen. Software-Testing identiﬁziert Fehler durch das Ausf¨uhren von Programmen mit gegebenen Eingaben und steht vor zwei Hauptherausforderungen: der Generierung von Testeingaben und dem Finden von Test-Orakeln. Fuzzing ist eine weit verbreitete Methode zur Generierung von Testeingaben, whrend Speziﬁkationendas Orakel-Problem adressieren. Diese Techniken und Konzepte haben sich als effektiv und entscheidend f¨ur die Validierung der Software-Zuverlssigkeit erwiesen. Wir passen diese Methoden an, um die Zuverlssigkeit von ML-Modellen zu bewerten. Eine der grßten j¨ungsten Fortschritte im maschinellen Lernen war die Lsung sequenzieller Entscheidungsprobleme, bei denen Agenten Handlungsrichtlinien lernen. In dieser Arbeit entwickeln wir Techniken zur ","cbCais5repZDCFg2","https://ap.wps.com/l/cbCais5repZDCFg2","pdf",3570344,1,128,"English","en",105,"# Abstract\n## Motivation and challenge in functional correctness\n## Testing-inspired specification and fuzzing for ML\n## Action-policy reliability and test oracles\n## Metamorphic relations as hyperproperties (k-safety)\n## NOMOS: declarative specification language and automated testing\n## Applications and extensions","[{\"question\":\"Why is functional correctness for ML models difficult to specify and check?\",\"answer\":\"Researchers have addressed properties like robustness and fairness, but general functional-correctness expectations are challenging to formalize and verify for ML models, especially in safety-critical use.\"},{\"question\":\"How does the thesis connect software testing concepts to ML model validation?\",\"answer\":\"It adapts software-testing ideas—test-input generation and test oracles—to ML by using fuzzing for inputs and specifications to address the oracle problem for functional correctness.\"},{\"question\":\"What new techniques are proposed for testing action policies’ reliability?\",\"answer\":\"The thesis introduces test oracles based on metamorphic relations and presents the ! →fuzz framework to identify bugs in action policies, including ways to find undesirable yet avoidable outcomes.\"}]","Specifying and Fuzzing Machine-Learning Models - Doctoral Thesis | PDF",1785730223,323,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"specifying-and-fuzzing-machine-learning-models-doctoral-thesis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/specifying-and-fuzzing-machine-learning-models-doctoral-thesis/120462/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is functional correctness for ML models difficult to specify and check?","Question",{"text":75,"@type":76},"Researchers have addressed properties like robustness and fairness, but general functional-correctness expectations are challenging to formalize and verify for ML models, especially in safety-critical use.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the thesis connect software testing concepts to ML model validation?",{"text":80,"@type":76},"It adapts software-testing ideas—test-input generation and test oracles—to ML by using fuzzing for inputs and specifications to address the oracle problem for functional correctness.",{"name":82,"@type":73,"acceptedAnswer":83},"What new techniques are proposed for testing action policies’ reliability?",{"text":84,"@type":76},"The thesis introduces test oracles based on metamorphic relations and presents the ! →fuzz framework to identify bugs in action policies, including ways to find undesirable yet avoidable outcomes.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]