[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120517-en":3,"doc-seo-120517-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120517,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Extraction of Research Objectives, Machine Learning Model Names, and Dataset Names from Academic Papers and Analysis of Their Interrelationships Using LLM and Network - analysis","Machine learning adoption across industries requires mapping tasks to suitable models and datasets, yet this is costly because it demands both machine-learning and domain expertise. A proposed methodology extracts research objectives, machine learning methods, and dataset names from academic papers, then analyzes their interrelationships using LLM-based extraction, embedding-based synonym clustering, and network clustering over co-occurrence graphs. Using Llama3, expression extraction achieves F-scores above 0.8 across categories, and financial-domain benchmarking demonstrates practical value, including insights into recent ESG-related datasets.","arXiv :2408 . 12097v1 [ cs .LG] 22 Aug 2024  \nExtraction of Research Objectives, Machine Learning Model Names, and Dataset Names from Academic Papers and Analysis of Their Interrelationships Using LLM and Network  \nAnalysis  \nS. Nishio, H. Nonaka, N. Tsuchiya, A. Migita, Y. Banno∗  \nT. Hayashi† H. Sakaji‡ T. Sakumoto§ K. Watabe¶ August 23, 2024  \nAbstract  \nMachine learning is widely utilized across various industries. Identifying the appropriate machine learning models and datasets for specific tasks is crucial for the effective industrial application of machine learning. However, this requires expertise in both machine learning and the relevant domain, leading to a high learning cost. Therefore, research focused on extracting combinations of tasks, machine learning models, and datasets from academic papers is critically important, as it can facilitate the automatic recommendation of suitable methods. Conventional information extraction methods from academic papers have been limited to identifying machine learning models and other entities as named entities. To address this issue, this study proposes a methodology extracting tasks, machine learning methods, and dataset names from scientific papers and analyzing the relationships between these information by using LLM, embedding model, and network clustering. The proposed method’s expression extraction performance, when using Llama3, achieves an F-score exceeding 0.8 across various categories, confirming its practical utility. Benchmarking results on financial domain papers have demonstrated the effectiveness of this method, providing insights into the use of the latest datasets, including those related to ESG (Environmental, Social, and Governance) data.  \n∗ Aichi Institute of Technology †University of Tokyo ‡Hokkaido University  \n§ Nagaoka University of Technology ¶ Saitama University  \n1 Introduction  \nThe use of machine learning for analysis has rapidly spread in recent years and is now employed in various fields such as services and finance [ACMENH21][Nea18][YN21][KWS+ 21] . In this context, selecting the appropriate data and machine learning methods to solve specific problems requires not only knowledge of machine learning but also domain knowledge of the field being analyzed. Therefore, establishing methods to support decision-making has become urgent. There isan increasing amount of research focused on extracting technical terms, such as machine learning methods and dataset names, from academic papers and patent documents to aid in decision-making. For example, an early study before the widespread adoption of deep learning by [NKS+ 12] focused on extracting the technologies used in literature. However, studies conducted before the advent of deep learning faced performance issues and had practical limitations. Recently, models leveraging deep learning have emerged to improve performance. Studies such as [HMPM21] and [YYZ+ 23] have proposed methods to extract machine learning methods and dataset names from academic papers using language models. However, to “select the appropriate data and machine learning methods for problem-solving,” it is necessary to comprehensively analyze the relationships between research objectives, datasets, and machine learning methods, rather than merely extracting them. Furthermore, synonymous terms, such as “SVM”and “Support Vector Machine,” need to be recognized as the same expression to avoid separate analyses, which could lead to inconvenient statistical trends. Therefore, methods for semantic aggregation must also be employed. In this study, we propose a method that utilizes the large language model Llama2 to extract research objectives, machine learning methods, and dataset names from individual papers. The extracted expressions are then aggregated based on synonym relationships using the embedding model E5 . Additionally, we analyze the relationships between objectives, machine learning methods, and datasets using network clustering bas","cbCaiuBil2FKFM5l","https://ap.wps.com/l/cbCaiuBil2FKFM5l","pdf",645480,1,10,"English","en",105,"# Abstract\n# Introduction\n## Research motivation and challenges\n## Proposed approach and evaluation context\n# Methodology\n## Overview of the proposed method\n## Extraction using Llama","[{\"question\":\"What problem does the study address in applying machine learning to real tasks?\",\"answer\":\"It addresses the difficulty of selecting appropriate machine learning models and datasets for specific tasks, which requires expertise in both machine learning and the target domain.\"},{\"question\":\"How does the proposed method extract information from academic papers?\",\"answer\":\"It uses Llama-based prompts to extract research objectives, machine learning methods, and dataset names from paper text, then aggregates synonymous expressions using embeddings.\"},{\"question\":\"How are relationships between objectives, models, and datasets analyzed?\",\"answer\":\"The method builds a co-occurrence graph from the papers, treats clusters as nodes, and applies network clustering to reveal interrelationships.\"}]","Extraction of Research Objectives, Machine Learning Model Names, and Dataset Names from Academic Papers and Analysis of Their Interrelationships Using LLM and Network - analysis | PDF",1785730452,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"extraction-of-research-objectives-machine-learning-model-names-and-dataset-names-from-academic-papers-and-analysis-of-their-interrelationships-using-llm-and-network-analysis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/extraction-of-research-objectives-machine-learning-model-names-and-dataset-names-from-academic-papers-and-analysis-of-their-interrelationships-using-llm-and-network-analysis/120517/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the study address in applying machine learning to real tasks?","Question",{"text":75,"@type":76},"It addresses the difficulty of selecting appropriate machine learning models and datasets for specific tasks, which requires expertise in both machine learning and the target domain.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method extract information from academic papers?",{"text":80,"@type":76},"It uses Llama-based prompts to extract research objectives, machine learning methods, and dataset names from paper text, then aggregates synonymous expressions using embeddings.",{"name":82,"@type":73,"acceptedAnswer":83},"How are relationships between objectives, models, and datasets analyzed?",{"text":84,"@type":76},"The method builds a co-occurrence graph from the papers, treats clusters as nodes, and applies network clustering to reveal interrelationships.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]