[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119513-en":3,"doc-seo-119513-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119513,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Dataset Identifiers in Machine Learning Model Signature for Lineage Tracing","This disclosure introduces a method for embedding fingerprints of training datasets into machine learning model checkpoints and signatures to enable auditable dataset lineage tracing. Each authorized training dataset is registered with a unique identifier that records dataset size, number of examples, and provenance. The dataset fingerprint is carried forward to derived models, with additional fingerprints appended when new datasets are used in later training stages. A checkpoint-saving mechanism cross-validates recorded dataset size against fingerprints, failing on mismatches, to support compliance auditing and reduce legal and financial risks.","Technical Disclosure Commons  \nDefensive Publications Series  \n30 Jul 2025  \nDataset Identifiers in Machine Learning Model Signature for Lineage Tracing  \nLenord Melvix Joseph Stephen Max  \nFollow this and additional works at: [https://www.tdcommons.org/dpubs_series](https://www.tdcommons.org/dpubs_series)  \nRecommended Citation  \nMax, Lenord Melvix Joseph Stephen, \"Dataset Identifiers in Machine Learning Model Signature for Lineage Tracing\", Technical Disclosure Commons,(July 30, 2025)  \n[https://www.tdcommons.org/dpubs_series/8408](https://www.tdcommons.org/dpubs_series/8408)  \nThis work is licensed under a Creative Commons Attribution 4.0 License.  \nThis Article is brought to you for free and open access by Technical Disclosure Commons. It has been accepted for inclusion in Defensive Publications Series by an authorized administrator of Technical Disclosure Commons.  \nDataset Identifiers in Machine Learning Model Signature for Lineage Tracing  \nABSTRACT  \nThis disclosure introduces a method for embedding fingerprints of training datasets used for model training into machine learning model checkpoints and signatures. This allows for auditable lineage tracing of datasets used in model training, even for models that are derived from a base model (or part ofa sequence) . Each authorized dataset is registered with a unique identifier that includes its size, number of examples, and provenance. The dataset fingerprint is carried over to subsequent models, along with new fingerprints for any additional datasets used for training of the subsequent models. A checkpoint saving system cross-validates training dataset size against recorded fingerprints, failing if there's a mismatch. This technique enables auditing models for compliance. The techniques provide a systematic and auditable way to  \ntrack dataset usage, thereby enabling greater trust and reliability in released models. The  \ntechniques help organizations that build and provide machine learning models to maintain a  \nrecord of dataset usage, mitigating legal and financial risks.  \nKEYWORDS  \n● Model training ● Model checkpoint  \n● LLM training ● Model lineage  \n● LLM audit ● Derived model  \n● Machine learning audit ● Dataset augmentation  \n● Training dataset ● Unauthorized data use  \n● Dataset fingerprinting  \nPublished by Technical Disclosure Commons, 2025 2  \nBACKGROUND  \nLarge language models (LLMs) find a wide variety of applications. There are several  \nwell-known providers of LLMs via chatbot and/or API interfaces. Additionally, many  \norganizations also customize or release their own variants of LLMs trained on massive web and  \ndigital datasets.  \nThe datasets used for training such models can sometimes include material that the organization building the model may not have legitimate access to-this can occur inadvertently or due to unauthorized use. However, once a model has been trained and released, proving the breach of data rights (in building of the model) is extremely difficult. During training,  \nparameters of the model-which may run into millions-are adjusted based on the training  \ndataset. Information from the training dataset is implicitly deeply embedded into the model,  \nrepresented in a high-dimensional space.  \nDetection of unauthorized use of training data is particularly difficult when the unauthorized data-use is deep inside a lineage of models that are sequentially trained one on top of another.  \nApplication developers can license or purchase access to LLMs from a model provider to build end-user applications on top ofthe LLM. In such cases, the application developer may face legal and/or financial risk for use of the models from the model provider if the model provider has made inappropriate use of data during model training. To prevent such a situation, the application developer may insist on the model provider to warrant that the training dataset was used with appropriate permissions and in compliance with applicable law (e.g., copyright) .  \nW","cbCaivkixCGeZMr0","https://ap.wps.com/l/cbCaivkixCGeZMr0","pdf",180445,1,6,"English","en",105,"# Abstract\n# Background\n## Risks of unauthorized training data\n## Limitations of voluntary disclosures and data cards\n# Description\n## Dataset identifiers and provenance\n## Fingerprints carried across model lineage\n## Checkpoint cross-validation for compliance","[{\"question\":\"How does the method enable auditable lineage tracing of training datasets?\",\"answer\":\"It embeds dataset fingerprints and dataset identifiers into model checkpoints and signatures during training, then carries them forward to derived models so lineage can be audited across sequential model derivations.\"},{\"question\":\"What information is included in a dataset identifier?\",\"answer\":\"The dataset identifier includes dataset size, number of examples, and metadata describing the source and provenance of the dataset.\"},{\"question\":\"How does the checkpoint saving system detect inconsistencies?\",\"answer\":\"When saving a checkpoint, it cross-validates the recorded training dataset size against the stored fingerprints and fails if there is a mismatch, preventing unverifiable checkpoint exports.\"}]","Dataset Identifiers in Machine Learning Model Signature for Lineage Tracing | PDF",1785724724,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"dataset-identifiers-in-machine-learning-model-signature-for-lineage-tracing","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/dataset-identifiers-in-machine-learning-model-signature-for-lineage-tracing/119513/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the method enable auditable lineage tracing of training datasets?","Question",{"text":75,"@type":76},"It embeds dataset fingerprints and dataset identifiers into model checkpoints and signatures during training, then carries them forward to derived models so lineage can be audited across sequential model derivations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What information is included in a dataset identifier?",{"text":80,"@type":76},"The dataset identifier includes dataset size, number of examples, and metadata describing the source and provenance of the dataset.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the checkpoint saving system detect inconsistencies?",{"text":84,"@type":76},"When saving a checkpoint, it cross-validates the recorded training dataset size against the stored fingerprints and fails if there is a mismatch, preventing unverifiable checkpoint exports.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]