[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124894-en":3,"doc-seo-124894-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124894,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Couler - Unified Machine Learning Workflow - Optimization in Cloud","Machine Learning workflows often become complex, resource-intensive, and time-consuming to build and optimize, especially when users must adapt to many different workflow engines and their distinct APIs. Expanding ML workflows across heterogeneous data infrastructure can further increase workload scale and deployment costs. COULER addresses these gaps by generating workflows from natural-language descriptions, integrating LLMs for automated workflow creation, and exposing a unified programming interface. It also improves efficiency via multi-stage caching, large-scale auto-parallelization, and automatic hyperparameter tuning, reducing redundant computation and enhancing fault tolerance in deep learning training.","Couler: Unified Machine Learning Workflow  \nOptimization in Cloud  \nXiaoda Wang§ , Yuan Tang†, Tengda Guo§ , Bo Sang∗ Jingji Wu∗ , Jian Sha∗ , Ke Zhang∗ , Jiang Qian ∥ , Mingjie Tang§  \n∗ Ant Group †Red Hat, Inc. ∥ Snap, Inc § Sichuan University  \n{wangxiaoda, [guotengda](guotengda}@stu.scu.edu.cn)[}](guotengda}@stu.scu.edu.cn)[@stu.scu.edu.cn](guotengda}@stu.scu.edu.cn), {tangrock, [terrytangyuan](terrytangyuan}@gmail.com)[}](terrytangyuan}@gmail.com)[@gmail.com](terrytangyuan}@gmail.com),  \n{b.sang, jingji.wjw, shajian, [yingzi.zk](yingzi.zk}@antgroup.com)[}](yingzi.zk}@antgroup.com)[@antgroup.com](yingzi.zk}@antgroup.com)  \narXiv :2403 .07608v1 [ cs .DB] 12 Mar 2024  \nAbstract—Machine Learning (ML) has become ubiquitous, fueling data-driven applications across various organizations. Contrary to the traditional perception of ML in research, ML workflows can be complex, resource-intensive, and timeconsuming. Expanding an ML workflow to encompass a wider range of data infrastructure and data types may lead to larger workloads and increased deployment costs. Currently, numerous workflow engines are available (with over ten being widely recognized). This variety poses a challenge for end-users in terms of mastering different engine APIs. While efforts have primarily focused on optimizing ML Operations (MLOps) for a specific workflow engine, current methods largely overlook workflow optimization across different engines.  \nIn this work, we design and implement COULER , a system designed for unified ML workflow optimization in the cloud. Our main insight lies in the ability to generate an ML workflow using natural language (NL) descriptions. We integrate Large Language Models (LLMs) into workflow generation, and provide a unified programming interface for various workflow engines. This approach alleviates the need to understand various workflow engines’ APIs. Moreover, COULER enhances workflow computation efficiency by introducing automated caching at multiple stages, enabling large workflow auto-parallelization and automatic hyperparameters tuning. These enhancements minimize redundant computational costs and improve fault tolerance during deep learning workflow training. COULER is extensively deployed in real-world production scenarios at ANT GROUP , handling approximately 22k workflows daily, and has successfully improved the CPU/Memory utilization by more than 15% and the workflow completion rate by around 17% .  \nIndex Terms—Machine Learning Workflow, LLM, Cloud  \nI. INTRODUCTION  \nA workflow, commonly known as a data pipeline, entails a sequence of steps that process raw data from various sources, directing it to a destination for both storage and analysis. Similarly, an ML workflow streamlines the comprehensive MLOps workflow, spanning data acquisition, exploratory data analysis (EDA), data augmentation, model creation, and deployment. Post-deployment, this ML workflow facilitates reproducibility, tracking, and monitoring. Such workflows enhance the efficiency and management of the entire model lifecycle, leading to accelerated usability and streamlined deployment [27],[42] . To automate and oversee these workflows, ML orchestration tools are deployed, offering an intuitive and collaborative interface.  \nMingjie Tang, Yuan Tang, and Qian Jiang performed most of this work while at Ant Group.  \nFig. 1: An example of a financial company’s journey in leveraging machine learning to predict market trends.  \nExample. Refer to the example in Figure 1, a financial company aims to predict market trends using ML models. Initially, developers are tasked with selecting a workflow engine from a variety of available options, such as Argo, Airflow, Dolphin Scheduler, MetaFlow or Kubeflow Pipeline etc. Then, end users need to dedicate time to mastering the programming API of specific workflow engines. Upon defining the workflow, the first step entails data preprocessing. Subsequently, three models are evaluated, and the most promising model","cbCaiolLDbrsLLLo","https://ap.wps.com/l/cbCaiolLDbrsLLLo","pdf",1499900,1,22,"English","en",105,"# Introduction\n## Workflow orchestration and challenges\n# System overview\n## Natural-language-driven workflow generation\n## Unified interface across workflow engines\n# Optimization in cloud\n## Multi-stage caching\n## Auto-parallelization and hyperparameter tuning\n# Deployment impact","[{\"question\":\"What core problem does COULER target in cloud ML workflow engineering?\",\"answer\":\"COULER targets the difficulty of building and optimizing ML workflows across many workflow engines with different APIs, which increases user effort, time, and deployment cost.\"},{\"question\":\"How does COULER generate an ML workflow?\",\"answer\":\"COULER generates the ML workflow from natural-language descriptions using large language models (LLMs), reducing the need to manually learn each engine’s programming interface.\"},{\"question\":\"What mechanisms does COULER use to optimize workflow efficiency?\",\"answer\":\"COULER enhances efficiency with automated caching at multiple stages, enabling large workflow auto-parallelization and automatic hyperparameter tuning to cut redundant computation and improve fault tolerance.\"}]","Couler - Unified Machine Learning Workflow - Optimization in Cloud | PDF",1785895266,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"couler-unified-machine-learning-workflow-optimization-in-cloud","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/couler-unified-machine-learning-workflow-optimization-in-cloud/124894/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What core problem does COULER target in cloud ML workflow engineering?","Question",{"text":75,"@type":76},"COULER targets the difficulty of building and optimizing ML workflows across many workflow engines with different APIs, which increases user effort, time, and deployment cost.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does COULER generate an ML workflow?",{"text":80,"@type":76},"COULER generates the ML workflow from natural-language descriptions using large language models (LLMs), reducing the need to manually learn each engine’s programming interface.",{"name":82,"@type":73,"acceptedAnswer":83},"What mechanisms does COULER use to optimize workflow efficiency?",{"text":84,"@type":76},"COULER enhances efficiency with automated caching at multiple stages, enabling large workflow auto-parallelization and automatic hyperparameter tuning to cut redundant computation and improve fault tolerance.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]