[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117775-en":3,"doc-seo-117775-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117775,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","An Investigation of Licensing of Datasets for Machine Learning Based on the GQM Model - Systematic Assessment Using GQM Method","Dataset licensing is a persistent challenge in machine learning development. Publicly available datasets are often sourced from the web, which can mean many images are not commercially usable, while training teams frequently ignore whether the dataset license permits the intended use. As a result, dataset licensing remains incomplete and difficult to evaluate for commercial compliance. An investigation of two collection datasets shows that most current datasets lack licenses, preventing determination of commercial availability. The study applies the GQM method to systematically analyze licensing issues, design measurable questions, and assess 311 GitHub repositories and 42 datasets to identify potential licensing violations.","An investigation of licensing of datasets for machine learning based on the GQM model  \nJunyu Chen  \nNagoya University Nagoya, Japan  \n[chen.junyu.v3@s.mail.nagoya-u.ac.jp](chen.junyu.v3@s.mail.nagoya-u.ac.jp)  \nNorihiro Yoshida  \nRitsumeikan University Kusatsu, Japan [norihiro@fc.ritsumei.ac.jp](norihiro@fc.ritsumei.ac.jp)  \nHiroaki Takada  \nNagoya University Nagoya, Japan [hiro@ertl.jp](hiro@ertl.jp)  \narXiv :2303 . 13735v1 [ cs . SE] 24 Mar 2023  \nAbstract—Dataset licensing is currently an issue in the development of machine learning systems. And in the development of machine learning systems, the most widely used are publicly available datasets. However, since the images in the publicly available dataset are mainly obtained from the Internet, some images are not commercially available. Furthermore, developers of machine learning systems do not often care about the license of the dataset when training machine learning models with it. In summary, the licensing of datasets for machine learning systems is in a state of incompleteness in all aspects at this stage.  \nOur investigation of two collection datasets revealed that most of the current datasets lacked licenses, and the lack of licenses made it impossible to determine the commercial availability of the datasets. Therefore, we decided to take a more scientiﬁc and systematic approach to investigate the licensing of datasets and the licensing of machine learning systems that use the dataset to make it easier and more compliant for future developers of machine learning systems.  \nThis paper applies the GQM method, which is a framework originally used for systematic metrics and analysis in the ﬁeld of software engineering, to investigate the licensing of datasets and the licensing of machine learning systems. We designed two questions, “Are there licensing issues in the machine learning dataset?” and “Is it easy to use machine learning datasets following the licensing?” as the ﬁnal goal of this investigation, and designed 7 related questions and 12 quantiﬁable metrics to investigate and answer. To solve the ﬁnal goal, we investigated 311 repositories related to machine learning systems on GitHub, with a total of 42 datasets used by these systems. By investigating the licensing of these repositories and datasets, we found that in the present environment of lack of well-managed dataset licensing, developers of machine learning systems are exposed to potential licensing violations if they want to use publicly available datasets.  \nIndex Terms—Dataset, Machine learning, Licensing conﬂict, Software engineering  \nI. INTRODUCTION  \nIn this section, we will introduce the background of the current research on licensing of machine learning datasets. Next, we will introduce our target based on this background and our contribution.  \nA. Background  \nThe development of machine learning systems in the last decade has been exponential. And in machine learning systems, the essential thing is the models trained based on datasets. Machine learning datasets are collections of data used to train and evaluate machine learning models. So as machine  \nlearning development grows and becomes huge in scalability, the use of datasets becomes increasingly popular. Therefore, the collection and development of datasets are particularly needed. However, as the ﬁeld of machine learning has shifted to approaches with more signiﬁcant data requirements over the past decade, the skilled and methodical annotation applied in early dataset collection practices was considered slow and consuming. Today's data collection tends to shift to unconstrained data, which has led to more and more data collection from the web [14] . These datasets can come from a variety of sources, such as government agencies, research institutions, and private companies.  \nSince traditional data collection and annotation are considered to be slow and costly, most developers of datasets use script-like tools to collect data directly from the","cbCairFvr3QcgLZ2","https://ap.wps.com/l/cbCairFvr3QcgLZ2","pdf",1049402,1,17,"English","en",105,"# Introduction\n## Background\n## Legal foundations for dataset licensing\n# Research Method and Questions\n## GQM-based investigation design\n## Research questions and metrics\n# Data Collection and Study Scope\n## GitHub repositories and used datasets\n# Findings and Implications\n## Licensing gaps and commercial uncertainty\n## Risk of potential licensing violations","[{\"question\":\"Why is dataset licensing a key issue in machine learning systems?\",\"answer\":\"Publicly available datasets are often obtained from the internet, so parts of them may not be commercially available. Developers also frequently train models without carefully checking license terms, which leads to incomplete licensing information and uncertainty about compliant use.\"},{\"question\":\"What did the investigation of collection datasets reveal?\",\"answer\":\"Most current datasets lacked clear licenses. Without license information, it becomes impossible to determine whether the datasets are commercially available for downstream machine learning use.\"},{\"question\":\"How does the paper apply the GQM method to study licensing?\",\"answer\":\"The paper uses GQM to define the final goals as specific questions about licensing issues in datasets and whether licensing makes datasets easy to use. It then derives seven related questions and twelve quantifiable metrics, and evaluates 311 GitHub repositories using 42 datasets.\"}]","An Investigation of Licensing of Datasets for Machine Learning Based on the GQM Model - Systematic Assessment Using GQM Method | PDF",1785679486,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"an-investigation-of-licensing-of-datasets-for-machine-learning-based-on-the-gqm-model-systematic-assessment-using-gqm-method","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/an-investigation-of-licensing-of-datasets-for-machine-learning-based-on-the-gqm-model-systematic-assessment-using-gqm-method/117775/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is dataset licensing a key issue in machine learning systems?","Question",{"text":75,"@type":76},"Publicly available datasets are often obtained from the internet, so parts of them may not be commercially available. Developers also frequently train models without carefully checking license terms, which leads to incomplete licensing information and uncertainty about compliant use.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What did the investigation of collection datasets reveal?",{"text":80,"@type":76},"Most current datasets lacked clear licenses. Without license information, it becomes impossible to determine whether the datasets are commercially available for downstream machine learning use.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the paper apply the GQM method to study licensing?",{"text":84,"@type":76},"The paper uses GQM to define the final goals as specific questions about licensing issues in datasets and whether licensing makes datasets easy to use. It then derives seven related questions and twelve quantifiable metrics, and evaluates 311 GitHub repositories using 42 datasets.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]