[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125153-en":3,"doc-seo-125153-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125153,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Automated Trustworthiness Testing for Machine Learning Classifiers","Machine Learning (ML) is used in high-stakes domains, making it essential to assess trustworthiness beyond predictive accuracy. This work targets explanations produced by explainable techniques such as LIME and SHAP, where the plausibility of explanation content strongly affects confidence in model decisions. It introduces TOWER, an automated approach that builds trustworthiness oracles for text classifiers by using word embeddings to test whether explanation words are semantically related to the predicted class. Experiments analyze unsupervised configurations using noisy, untrustworthy models and evaluate on a human-labeled dataset.","Automated Trustworthiness Testing for Machine Learning Classifiers  \nSteven Cho  \nUniversity of Auckland Auckland, New Zealand [scho518@aucklanduni.ac.nz](scho518@aucklanduni.ac.nz)  \nSeaton Cousins-Baxter  \nUniversity of Auckland Auckland, New Zealand [scou766@aucklanduni.ac.nz](scou766@aucklanduni.ac.nz)  \nStefano Ruberto JRC European Commission  \nIspra, Italy stefano.ruberto@ec.europa.eu  \nValerio Terragni  \nUniversity of Auckland Auckland, New Zealand [v.terragni@auckland.ac.nz](v.terragni@auckland.ac.nz)  \narXiv :2406 .0525 1v 1 [ cs .LG] 7 Jun 2024  \nAbstract—Machine Learning (ML) has become an integral part of our society, commonly used in critical domains such as finance, healthcare, and transportation. Therefore, it is crucial to evaluate not only whether ML models make correct predictions but also whether they do so for the correct reasons, ensuring our trust that will perform well on unseen data. This concept is known as trustworthiness in ML. Recently, explainable techniques (e.g., LIME, SHAP) have been developed to interpret the decisionmaking processes of ML models, providing explanations for their predictions (e.g., words in the input that influenced the prediction the most). Assessing the plausibility of these explanations can enhance our confidence in the models’ trustworthiness. However, current approaches typically rely on human judgment to determine the plausibility of these explanations.  \nThis paper proposes TOWER, the first technique to automatically create trustworthiness oracles that determine whether text classifier predictions are trustworthy. It leverages word embeddings to automatically evaluate the trustworthiness of a model-agnostic text classifiers based on the outputs of explanatory techniques. Our hypothesis is that a prediction is trustworthy if the words in its explanation are semantically related to the predicted class.  \nWe perform unsupervised learning with untrustworthy models obtained from noisy data to find the optimal configuration of TOWER. We then evaluated TOWER on a human-labeled trustworthiness dataset that we created. The results show that TOWER can detect a decrease in trustworthiness as noise increases, but is not effective when evaluated against the humanlabeled dataset. Our initial experiments suggest that our hypothesis is valid and promising, but further research is needed to better understand the relationship between explanations and trustworthiness issues.  \nIndex Terms—Machine Learning, Machine Learning Testing, Test Oracle Problem, Trustworthiness, Explainability, Software Testing, Text Classifications, Word Embeddings  \nI. INTRODUCTION  \nMachine Learning (ML) has become integral to many areas of society. With its prevalence, it is crucial to determine whether we can trust ML models. Trust goes beyond correctness. An ML model can achieve high accuracy and make correct predictions yet still be untrustworthy. This occurs when the reasons behind its correct predictions are flawed, making the model unreliable for unseen data.  \nOne way to measure trustworthiness is through understanding the decision-making process of ML models [5], [7], [15] . While there are ML classifiers that are inherently explainable (e.g., decision trees), most ML classifiers (e.g.,(deep) neural  \nFig. 1: Example of “untrustworthy” prediction from LIME paper [25] . Algorithm 2 based the prediction on irrelevant words: “Posting”,“Host”,“Re”, and “Nntp”  \nnetworks) cannot directly explain their decisions in a way that a human would understand. However, the area of Explainable Machine Learning (XML) [26] studies techniques to explain the predictions of any classifier, as long as it has interpretable inputs (i.e., text, numbers, or images) .  \nOne of the first and most representative XML technique is LIME [25] . It produces textual or visual artifacts for a qualitative understanding of the relationship between the test input instance’s components (e.g., words in a text) and the model’s prediction. For ex","cbCaiuTHAN4FZX4d","https://ap.wps.com/l/cbCaiuTHAN4FZX4d","pdf",1018953,1,10,"English","en",105,"# Introduction\n## Trustworthiness beyond correctness\n## Explainable ML and XML techniques\n## LIME as a representative approach\n## Motivation for automated trustworthiness oracles\n# TOWER Approach\n## Building trustworthiness oracles\n## Word embedding relatedness hypothesis\n# Experiments and Evaluation\n## Learning from noisy untrustworthy models\n## Evaluation on a human-labeled dataset\n## Findings and implications","[{\"question\":\"What does “trustworthiness” mean for machine learning classifiers in this paper?\",\"answer\":\"Trustworthiness means that a model’s correct predictions rely on the right reasons, so the decision rationale remains reliable for unseen data.\"},{\"question\":\"How does TOWER determine whether a text classifier’s prediction is trustworthy?\",\"answer\":\"TOWER uses word embeddings to automatically check whether words in the explanation are semantically related to the predicted class.\"},{\"question\":\"What were the main outcomes of the TOWER experiments?\",\"answer\":\"Results indicate TOWER can detect trustworthiness decreases as noise increases, but it was not effective when evaluated against the human-labeled trustworthiness dataset.\"}]","Automated Trustworthiness Testing for Machine Learning Classifiers | PDF",1785897012,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"automated-trustworthiness-testing-for-machine-learning-classifiers","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/automated-trustworthiness-testing-for-machine-learning-classifiers/125153/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does “trustworthiness” mean for machine learning classifiers in this paper?","Question",{"text":75,"@type":76},"Trustworthiness means that a model’s correct predictions rely on the right reasons, so the decision rationale remains reliable for unseen data.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does TOWER determine whether a text classifier’s prediction is trustworthy?",{"text":80,"@type":76},"TOWER uses word embeddings to automatically check whether words in the explanation are semantically related to the predicted class.",{"name":82,"@type":73,"acceptedAnswer":83},"What were the main outcomes of the TOWER experiments?",{"text":84,"@type":76},"Results indicate TOWER can detect trustworthiness decreases as noise increases, but it was not effective when evaluated against the human-labeled trustworthiness dataset.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]