[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119621-en":3,"doc-seo-119621-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119621,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Quantitative analysis of Machine Learning model performance - and the need to consider explainability - Technical Talk","Quantitative analysis of machine learning model performance highlights how far generative and multimodal systems can match human-level capabilities across tasks such as summarization, translation, factual question answering, code generation, image captioning, and speech recognition. Benchmark results illustrate strong performance on several academic and reasoning metrics, while cited research warns of fragility in mathematical reasoning as questions become more complex. The talk also quantifies business value of accuracy improvements and motivates explainability and richer evaluation metrics beyond accuracy, including precision, recall, and F1, to reveal true model behavior.","Quantitative analysis of Machine Learning model performance and the need to consider explainability  \nVishnu S. Pendyala, Ph.D.  \nSan Jose State University  \nTo cite this presentation: Pendyala, V.S. (2024)“Quantitative analysis of Machine Learning model performance and the need to consider explainability”. IEEE Computer Society, Santa Clara Valley Chapter  \nTechnical Talk, December 30, 2024  \n©Vishnu S. Pendyala This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License  \n\n| Task | Performance Compared to Humans |\n| --- | --- |\n| Text Summarization | Can achieve similar quality |\n| Machine Translation | Near human-quality for some languages |\n| Question Answering on Factual Topics | Can perform well on factual topics with large datasets |\n| Code Generation | Can generate some basic code |\n| Image Captioning | Can generate accurate descriptions of images |\n| Speech Recognition | Achieves high accuracy in controlled environments |\n\nTHE UNREASONABLE EFFECTIVENESS OF GENERATIVE AI’S  \nMULTIMODAL ABILITIES  \n©Vishnu S. Pendyala This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License  \n©Vishnu S. Pendyala This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License  \nGPT-4o: Academic Benchmarks  \n\n| Benchmark | Score | Interpretation |\n| --- | --- | --- |\n| MMLU | 88.7 | High level of understanding across a wide range of academic subjects, comparable to undergraduates or even a graduates in a general field. |\n| GPQA | 53.6 | Moderate proficiency in handling complex, nuanced questions, which aligns with the capabilities of an undergraduate. |\n| MATH | 76.6 | Strong mathematical abilities, akin to a student with an undergraduate degree\u003Cbr>in mathematics or a related field. |\n| HumanEval | 90.2 | Excellent programming skills, similar to those of a highly proficient software engineer or computer science graduate. |\n| MGSM | 90.5 | Exceptional proficiency in solving grade school level math problems across\u003Cbr>multiple languages. |\n| DROP | 83.4 | Strong reading comprehension and reasoning abilities, comparable to undergraduates well-prepared for graduate-level work. |\n\nSource: [https://community.openai.com/t/education-level-interpretation-of-gpt-4os-benchmarks/763947](https://community.openai.com/t/education-level-interpretation-of-gpt-4os-benchmarks/763947)  \nBut with a caveat…  \nMirzadeh, et al. \"Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models. \" arXiv preprint  \n“we investigate the fragility of mathematical reasoning in these models and demonstrate that their performance significantly deteriorates as the number of clauses in a question increases. We hypothesize that this decline is due to the fact that current LLMs are not capable of genuine logical reasoning; instead, they attempt to replicate the reasoning steps observed in their training data”  \n©Vishnu S. Pendyala This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License  \nWhy care for accuracy?  \nEvery decimal % pays!  \nSource: [https://www.datarobot.com/customers/steward-health-care/](https://www.datarobot.com/customers/steward-health-care/)  \n“Just a 1% reduction in registered nurses' hours paid per patient day netted $2 million in savings per year, for just eight of the 38 hospitals in Steward’s network”  \n“Reducing patient length of stay by 0.1% results in savings of over $10 million per year”  \n©Vishnu S. Pendyala This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License  \n©Vishnu S. Pendyala This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License  \ncompare  \nApproaches to evaluating Classifiersk-fold Cross-Validation  \nThis Photo by Unknown Author is licensed under CC BY-NC   \nThis Photo by Unknown Author is licensed under CC BY  \nThis Photo by Unknown Author is licensed under CC BY-SA ","cbCaies4twhp5CAS","https://ap.wps.com/l/cbCaies4twhp5CAS","pdf",6624157,1,31,"English","en",105,"# Model capabilities across tasks\n## Human comparison and observed strengths\n# Benchmark evidence for generative AI\n## GPT-4o academic benchmarks\n## Caveat: limitations in mathematical reasoning\n# Why accuracy matters\n## Cost impact of small error reductions\n# Classifier evaluation metrics\n## Precision and recall trade-offs\n## F1 score and need for additional metrics\n# Practical example: spam filtering\n## Coldmail scenario and metric interpretation","[{\"question\":\"Which machine learning tasks does the presentation compare against human-like performance?\",\"answer\":\"It compares text summarization, machine translation, factual question answering, code generation, image captioning, and speech recognition, indicating strong results in several controlled or data-rich scenarios.\"},{\"question\":\"What do GPT-4o academic benchmarks show in the talk?\",\"answer\":\"The talk reports scores on MMLU, GPQA, MATH, HumanEval, MGSM, and DROP, interpreting them as varying degrees of understanding, proficiency, and reasoning comparable to different education levels.\"},{\"question\":\"Why does the presentation argue that accuracy alone is insufficient for evaluating classifiers?\",\"answer\":\"A spam-filtering example shows that a model can achieve high accuracy while failing to provide meaningful precision or recall, so the “true picture” requires additional metrics like precision, recall, and F1.\"}]","Quantitative analysis of Machine Learning model performance - and the need to consider explainability - Technical Talk | PDF",1785725344,78,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"quantitative-analysis-of-machine-learning-model-performance-and-the-need-to-consider-explainability-technical-talk","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/quantitative-analysis-of-machine-learning-model-performance-and-the-need-to-consider-explainability-technical-talk/119621/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Which machine learning tasks does the presentation compare against human-like performance?","Question",{"text":75,"@type":76},"It compares text summarization, machine translation, factual question answering, code generation, image captioning, and speech recognition, indicating strong results in several controlled or data-rich scenarios.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What do GPT-4o academic benchmarks show in the talk?",{"text":80,"@type":76},"The talk reports scores on MMLU, GPQA, MATH, HumanEval, MGSM, and DROP, interpreting them as varying degrees of understanding, proficiency, and reasoning comparable to different education levels.",{"name":82,"@type":73,"acceptedAnswer":83},"Why does the presentation argue that accuracy alone is insufficient for evaluating classifiers?",{"text":84,"@type":76},"A spam-filtering example shows that a model can achieve high accuracy while failing to provide meaningful precision or recall, so the “true picture” requires additional metrics like precision, recall, and F1.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]