[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127225-en":3,"doc-seo-127225-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},127225,2336475104042,"Skyler","https://ap-avatar.wpscdn.com/avatar/22000c4c32af1715be0?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786537525561427321",8,"Research & Report","SYSTEM, METHOD, AND COMPUTER PROGRAM PRODUCT FOR EVALUATING GENERATIVE MACHINE LEARNING MODELS - Article style description","The disclosure presents systems, methods, and computer program products for evaluating generative machine learning models to address deficiencies in transparency of generated answers. It targets assessment of confidence, correctness, and completeness, including hallucination detection and verification against ground-truth answers. A scoring evaluator model compares generated answers to ground-truth using a grading scale with at least five scores spanning accuracy, honesty, and completeness, enabling threshold-based rejection and selection of answer subsets for downstream use.","Technical Disclosure Commons  \n\n| Defensive Publications Series |\n| --- |\n| 03 Jun 2025\u003Cbr>SYSTEM, METHOD, AND COMPUTER PROGRAM PRODUCT FOR EVALUATING GENERATIVE MACHINE LEARNING MODELS\u003Cbr>Yang Wang VISA\u003Cbr>Alberto Garcia Hernandez VISA\u003Cbr>Roman Kyslyi\u003Cbr>VISA\u003Cbr>Nicholas Kersting VISA\u003Cbr>Ajit Vilasrao Patil\u003Cbr>VISA\u003Cbr>See next page for additional authors\u003Cbr>Follow this and additional works at: [https://www.tdcommons.org/dpubs_series](https://www.tdcommons.org/dpubs_series) |\n\nRecommended Citation  \nWang, Yang; Hernandez, Alberto Garcia; Kyslyi, Roman; Kersting, Nicholas; Patil, Ajit Vilasrao; and Dutta, Ranjan, \"SYSTEM, METHOD, AND COMPUTER PROGRAM PRODUCT FOR EVALUATING GENERATIVE MACHINE LEARNING MODELS\", Technical Disclosure Commons,(June 03, 2025)  \n[https://www.tdcommons.org/dpubs_series/8187](https://www.tdcommons.org/dpubs_series/8187)  \nThis work is licensed under a Creative Commons Attribution 4.0 License.  \nThis Article is brought to you for free and open access by Technical Disclosure Commons. It has been accepted for inclusion in Defensive Publications Series by an authorized administrator of Technical Disclosure Commons.  \nInventor(s)  \nYang Wang, Alberto Garcia Hernandez, Roman Kyslyi, Nicholas Kersting, Ajit Vilasrao Patil, and Ranjan Dutta  \nThis article is available at Technical Disclosure Commons: [https://www.tdcommons.org/dpubs_series/8187](https://www.tdcommons.org/dpubs_series/8187)  \nTITLE: “SYSTEM, METHOD, AND COMPUTER PROGRAM PRODUCT FOR EVALUATING GENERATIVE MACHINE  \nLEARNING MODELS”  \nVISA  \nYANG WANG  \nALBERTO GARCIA HERNANDEZ  \nROMAN KYSLYI  \nNICHOLAS KERSTING  \nAJIT VILASRAO PATIL  \nRANJAN DUTTA  \nPublished by Technical Disclosure Commons, 2025 2  \nTECHNICAL FIELD  \n[0001] The present disclosure relates generally to artificial intelligence, and, in non-limiting embodiments or aspects, to systems, methods, and computer program products for evaluating  \ngenerative machine learning models.  \nBACKGROUND  \n[0002] Generative machine learning models may suffer from a lack of transparency related to the confidence, correctness, and completeness of generated answers. Even for machine learning models that employ retrieval-augmented generation (RAG), the hallucination offacts in generated answers may create errors, which may be exacerbated when such generated answers are relied upon for downstream tasks. Users of generative machine learning models may prefer receiving answers that admit not knowing the answer to a query, rather than receiving a seemingly confident and incorrect answer, but users may not have the means to assess model honesty. Moreover, users of generative machine learning models may prefer receiving fully accurate and complete answers, rather than receiving partly accurate or incomplete answers, but users may not have the means to assess whether the entirety of generated answers are accurate, or whether answers are missing  \nimportant details.  \n[0003] There is a need in the art for a technical solution that can efficiently identify hallucination in generative machine learning models, as well as assess generated answers for  \naccuracy and missing information.  \nSUMMARY  \n[0004] Accordingly, provided are improved systems, methods, and computer program products for evaluating generative machine learning models.  \n[0005] According to non-limiting embodiments or aspects, provided is a system for evaluating generative machine learning models. The system includes at least one processor. The at least one processor is configured to determine a plurality of queries and determine a plurality of groundtruth answers associated with the plurality of queries. The at least one processor is also configured to generate a plurality of generated answers based on the plurality of queries using at least one  \n[https://www.tdcommons.org/dpubs_series/8187](https://www.tdcommons.org/dpubs_series/8187) 3  \nmachine learning model that is being evaluated. The at least one processor is further configured to input the","cbCaiq0LCvUvtgPW","https://ap.wps.com/l/cbCaiq0LCvUvtgPW","pdf",351401,1,20,"English","en",105,"# Technical Field\n# Background\n# Summary\n## System for Evaluating Models\n## Computer-Implemented Method","[{\"question\":\"What problem does the disclosure address in generative machine learning models?\",\"answer\":\"It addresses limited transparency in generated answers, including missing or incorrect information and hallucinations that can lead to downstream errors.\"},{\"question\":\"How does the proposed approach evaluate generated answers?\",\"answer\":\"It generates answers for multiple queries, compares them to associated ground-truth answers using an evaluator model, and produces scores based on a grading scale covering accuracy, honesty, and completeness.\"},{\"question\":\"What happens after scoring is computed?\",\"answer\":\"Answers with scores failing a predetermined threshold are rejected, while those meeting the threshold are provided as a selected subset of generated answers.\"}]","SYSTEM, METHOD, AND COMPUTER PROGRAM PRODUCT FOR EVALUATING GENERATIVE MACHINE LEARNING MODELS - Article style description | PDF",1785937622,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"system-method-and-computer-program-product-for-evaluating-generative-machine-learning-models-article-style-description","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/system-method-and-computer-program-product-for-evaluating-generative-machine-learning-models-article-style-description/127225/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-22","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the disclosure address in generative machine learning models?","Question",{"text":76,"@type":77},"It addresses limited transparency in generated answers, including missing or incorrect information and hallucinations that can lead to downstream errors.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the proposed approach evaluate generated answers?",{"text":81,"@type":77},"It generates answers for multiple queries, compares them to associated ground-truth answers using an evaluator model, and produces scores based on a grading scale covering accuracy, honesty, and completeness.",{"name":83,"@type":74,"acceptedAnswer":84},"What happens after scoring is computed?",{"text":85,"@type":77},"Answers with scores failing a predetermined threshold are rejected, while those meeting the threshold are provided as a selected subset of generated answers.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":29,"slug":114},6,"Technology","technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":21,"slug":126},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":21,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]