[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119024-en":3,"doc-seo-119024-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},119024,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","Data Extraction, Transformation, and Loading Process Automation for Algorithmic Trading Machine Learning Modelling and Performance Optimization","A data warehouse streamlines machine learning workflows by enabling efficient preparation of data for fast analysis and modelling. The paper reviews existing approaches to automating the Data Extraction, Transformation, and Loading (ETL) process for algorithmic trading and positions Data Warehouses and future Data Lakes as enablers for time- and performance-critical research requirements. It focuses on architectures that support rapid data processing and improves operational reliability through ETL automation aligned with stock-price prediction tasks.","Data Extraction, Transformation, and Loading Process Automation for Algorithmic Trading Machine Learning Modelling and Performance  \nOptimization  \narXiv :2312 . 12774v1 [ cs .DC] 20 Dec 2023  \nNassi Ebadifard  \nComputer Science Okanagan College Kelowna, Canada 0009-0002-9087-5259  \nAjitesh Parihar  \nComputer Science Okanagan College Kelowna, Canada 0009-0001-3162-1470  \nYoury Khmelevsky  \nComputer Science Okanagan College Kelowna, Canada 0000-0002-6837-3490  \nGatan Hains  \nLACL Universite´ Paris-Est Crteil, France 0000-0002-1687-8091  \nAlbert Wong  \nMathematics and Statistics Langara College Vancouver, Canada 0000-0002-0669-4352  \nFrank Zhang  \nSchool of Computing University of the Fraser Valley Abbotsford, Canada 0000-0001-7570-9805  \nAbstract—A data warehouse efficiently prepares data for effective and fast data analysis and modelling using machine learning algorithms. This paper discusses existing solutions for the Data Extraction, Transformation, and Loading (ETL) process and automation for algorithmic trading algorithms. Integrating the Data Warehouses and, in the future, the Data Lakes with the Machine Learning Algorithms gives enormous opportunities in research when performance and data processing time become critical non-functional requirements.  \nIndex Terms—Stock Price Predictions, Exogenous variables, Support Vector Regression, Multilayer Perceptron, Random Forest, XGBoost, Machine Learning, Algorithmic Trading.  \nI. INTRODUCTION  \nThe popularity of machine learning models and algorithms has been increasing greatly over the past decade. They will continue to be used even more, especially with the tremendous success of Large Language Models (LLM) in data analysis [1] . Professionals like researchers and analysts incorporated machine learning into daily lives. From corporations to individuals, the use of machine learning is applicable in a wide range of settings [2] .  \nTo improve the performance of the machine learning algorithms for stock price forecasting, we decided to implement two layers of the databases: (1) an OLTPDBMS system for the data collection and (2) following data migration and transformation into a data warehouse (DW) employing an efficient ETL process.  \nWe would like to acknowledge and thank the Post Degree Diploma program in Data Analytics and the Work on Campus programs at Langara College and the Computer Science Department at Okanagan College for supporting our research.  \nThe advancement of automation in the ETL process aims to reduce human errors and streamline the process. This is usually done using tools and software to automate data extraction, cleansing, and transformation tasks. Automation can also schedule and monitor ETL jobs, ensuring they run on time and produce accurate results.  \nFirst, this paper analyzes existing works related to Algorithmic Trading data collection and the following transformation into a DW. Then, we describe our prototype design and development for automatic data collection, which involves using various tools and technologies. The main contributions of this paper are (1) automation of data collection in an OLTP database in a cloud and (2) automated ETL process for the data collection, transformation and loading in a DBMS and then transfer into DW (3) ML performance improvement in the domain of Algorithmic Trading Machine Learning Modelling. The following steps will relate to the data pre-aggregation, drastically improving ML performance (see more information in our previous works here [3]–[31]) .  \nFor the 1st testing prototype in the Cloud, we used Oracle APEX to test data collection and cleaning. The 2nd system’s prototype was built using Compute Canada research resources (now this is The Digital Research Alliance of Canada [32]) . We used a Python Programming language and a PostgreSQL OLTP database system in the Cloud to gather publicly available data from commercial companies to test the effectiveness of ML models. But we found that the programming environmen","cbCaivGKqnJHTjeA","https://ap.wps.com/l/cbCaivGKqnJHTjeA","pdf",306126,1,"English","en",105,"# Introduction\n## System architecture and database layers\n# Existing Works\n## Machine learning for stock price prediction\n## ETL/ELT approaches for data warehouses\n# Prototype Design and Development\n## Cloud data collection and cleaning\n## Tooling and databases used\n# Contributions and Performance Optimization\n## Automated collection and ETL pipeline\n## Machine learning modelling improvements","[{\"question\":\"What problem does the proposed approach address in algorithmic trading machine learning?\",\"answer\":\"It targets the efficiency and automation of the data preparation pipeline so models can run with critical performance and reduced processing time.\"},{\"question\":\"How is the data pipeline structured in the proposed system?\",\"answer\":\"The approach uses two database layers: an OLTPDBMS for data collection, followed by migration, transformation, and loading into a data warehouse via an automated ETL process.\"},{\"question\":\"What is automated in the ETL and data collection workflow?\",\"answer\":\"The workflow automates data extraction, cleansing, transformation, scheduling, monitoring of ETL jobs, and the collection steps needed to support machine learning modelling.\"}]","Data Extraction, Transformation, and Loading Process Automation for Algorithmic Trading Machine Learning Modelling and Performance Optimization | PDF",1785721974,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"data-extraction-transformation-and-loading-process-automation-for-algorithmic-trading-machine-learning-modelling-and-performance-optimization","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/data-extraction-transformation-and-loading-process-automation-for-algorithmic-trading-machine-learning-modelling-and-performance-optimization/119024/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":22,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does the proposed approach address in algorithmic trading machine learning?","Question",{"text":74,"@type":75},"It targets the efficiency and automation of the data preparation pipeline so models can run with critical performance and reduced processing time.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How is the data pipeline structured in the proposed system?",{"text":79,"@type":75},"The approach uses two database layers: an OLTPDBMS for data collection, followed by migration, transformation, and loading into a data warehouse via an automated ETL process.",{"name":81,"@type":72,"acceptedAnswer":82},"What is automated in the ETL and data collection workflow?",{"text":83,"@type":75},"The workflow automates data extraction, cleansing, transformation, scheduling, monitoring of ETL jobs, and the collection steps needed to support machine learning modelling.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]