[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118816-en":3,"doc-seo-118816-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118816,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Open Data for Machine Learning - Bachelor thesis version","This thesis investigates how large language models, specifically GPT-3.5, can enhance the quality and usefulness of datasets in Open Data portals. It addresses the gap left by metadata standards such as DCAT, which do not capture creation context and related dimensions required for ML fairness and safety. By extracting structured information from accompanying natural-language dataset documentation, including Data Management Plans and User Guides, the work evaluates prompting strategies and demonstrates promising results for generating usable structured outputs.","This is the published version of the bachelor thesis:  \nFernández Alvarez, Raul; Giner Miguelez, Joan, dir. Open Data for Machine Learning. 2023. (Enginyeria Informàtica)  \nThis version is available at [https://ddd.uab.cat/record/280712](https://ddd.uab.cat/record/280712)[ ](https://ddd.uab.cat/record/280712)under the terms of the  license  \nOpen Data for Machine Learning  \nRaúl Fenández Álvarez  \nResum— Aquest treball explora el potencial dels grans models de llenguatge (LLMs), concretamentel GPT-3.5, en millorar la qualitat de les dades en portals de dades obertes. Estudis recents a la comunitat de Machine Learning (ML) així com Iniciatives Legislatives com l'European AI ACT, apunten a la necessitat de documentar els datasets usat per entrenar models de ML en un seguit de dimensions per garantir la seva equitat i seguretat. En aquestes iniciatives s'hi destaca la importa de documentar el context de creació de dades, així com els equips i infraestructura que han participat en la col·lecció i anotació dedades. En el cas dels Open Data portals, els estàndard de metadades com DCAT, no ofereixen suport per anotar d'aquesta informació i aquesta, en cas de ser-hi, només la podem trobar a la documentació adjunta dels dataset en format de text natural.  \nEn aquest treball s'explora l'ús de LLM per extreure de forma estructurada d'aquesta informació dela documentació dels datasets. Amb aquest fi, s'ha identificat els tipus de documentació presents susceptibles de funcionar amb el mètode proposat i s'ha explorat diferents estratègies de prompting per optimitzar l'ús LLM. Els resultats d'aquest estudi mostren bon resultat en format de documentació estructuradade dades presents als Open Data portals, com el Data Mangement Plans (DMP), i obren possibilitat adesenvolupar eines i mètodes per millorar la qualitat de les dades en aquests portals.  \nParaules clau—Grans Models de Llenguatge, Open Data portal, Metadades, Data Management Plan (DMP), User Guide, Prompting Strategies, Extracció de Dades.  \nAbstract—This work explores the potential of large language models (LLMs), specifically GPT-3 .5, in improving the quality of data in Open Data portals. Recent studies in the machine learning (ML) community and legislative initiatives like the European AI ACT emphasize the need to document the datasets used to train ML models across various dimensions to ensure their fairness and safety. These initiatives highlight the importance of documenting the data creation context, as well as the teams and infrastructure involved in data collection and annotation. In the case of Open Data portals, metadata standards like DCAT do not provide support for annotating this information, and if present, it can only be found in the accompanying documentation of the dataset in natural language format.  \nThis work explores the use of LLMs to extract this information from the documentation of datasets in a structured manner. To this end, the types of susceptible documentation present for the proposed method have been identified, and different prompting strategies have been explored to optimize the use of LLMs. The results of this study demonstrate good performance in generating structured documentation of data present in Open Data portals, such as Data Management Plans (DMPs), and open up possibilities for developing tools and methods to improve the quality of data in these portals.  \nIndex Terms—Large Language Models, Open Data portals, Metadata, Data Management Plan (DMP), User Guide, Prompting Strategies, Data Extraction  \n  ◆    \n1 INTRODUCTION  \nS ociety  \nmeans  \nhas evolved to be more data-driven, which that data is produced and used in order to  \nmake most decisions in life. The advancements in the field of AI have led to these applications having an  \n• E-mail de contacte: [raul.fernandeza@uab.cat](raul.fernandeza@uab.cat)  \n• Menció realitzada: Tecnologies de la Informació  \n• Treball tutoritzat per: Joan Giner-Miguelez  \n• Curs 2022/23  \nincreasingly signif","cbCaigrtvR2rh8eN","https://ap.wps.com/l/cbCaigrtvR2rh8eN","pdf",638971,1,10,"English","en",105,"# Introduction\n## Motivation and data quality challenges\n## Need for dataset creation context\n## Open Data standards and limitations\n# Related approaches and problem formulation\n## Legislative and community requirements\n## Target document types in portals\n# Method and experimentation\n## Prompting strategies for structured extraction\n## Dataset selection and document extraction\n# Results and implications\n## Structured documentation generation performance\n## Opportunities for tooling to improve data quality","[{\"question\":\"What problem does the thesis focus on for Open Data portals?\",\"answer\":\"It focuses on the difficulty of capturing dataset creation context for ML use, since standards like DCAT do not support annotating that information and it often exists only in natural-language documentation.\"},{\"question\":\"How are large language models used in the proposed approach?\",\"answer\":\"The thesis uses LLMs such as GPT-3.5 to extract required dimensions in a structured manner from dataset documentation provided by Open Data portals.\"},{\"question\":\"Which types of documentation are examined for extraction?\",\"answer\":\"The work identifies two key documentation types commonly found in portals: Data Management Plans (DMPs) and User Guides, and then applies prompting strategies to extract structured data from them.\"}]","Open Data for Machine Learning - Bachelor thesis version | PDF",1785720421,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"open-data-for-machine-learning-bachelor-thesis-version","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/open-data-for-machine-learning-bachelor-thesis-version/118816/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the thesis focus on for Open Data portals?","Question",{"text":75,"@type":76},"It focuses on the difficulty of capturing dataset creation context for ML use, since standards like DCAT do not support annotating that information and it often exists only in natural-language documentation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are large language models used in the proposed approach?",{"text":80,"@type":76},"The thesis uses LLMs such as GPT-3.5 to extract required dimensions in a structured manner from dataset documentation provided by Open Data portals.",{"name":82,"@type":73,"acceptedAnswer":83},"Which types of documentation are examined for extraction?",{"text":84,"@type":76},"The work identifies two key documentation types commonly found in portals: Data Management Plans (DMPs) and User Guides, and then applies prompting strategies to extract structured data from them.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]