[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83050-en":3,"doc-seo-83050-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83050,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design","Large language models generate BPMN process models from natural-language descriptions, but supervised fine-tuning restricts outputs to training-set patterns. Reinforcement learning can surpass this limit with external quality measures, yet reward design for multi-dimensional quality has remained unclear. This study systematically examines reward composition for RL-based process model generation using two LLM families (Llama 3.1 8B, Qwen 2.5 14B) across 48 configurations and 38 automated evaluation metrics spanning syntactic, pragmatic, and semantic quality.","Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design  \nAlexander Rombacha,b,∗ , Chantale Lauera,b and Nijat Mehdiyeva,b  \na German Research Center for Artificial Intelligence (DFKI), Campus D3 2, Saarbrücken, 66123, Saarland, Germany b Saarland University, Campus D3 2, Saarbrücken, 66123, Saarland, Germany  \narXiv :2607 .06 175v 1 [ cs .CL] 7 Jul 2026  \nARTICLE INFO  \nKeywords:  \nBusiness Process Modeling Large Language Models Supervised Finetuning Reinforcement Learning  \nAB STRACT  \nLarge language models (LLMs) can generate BPMN process models from natural-language descriptions, yet supervised fine-tuning (SFT) limits their output quality to the patterns present in the training data. Reinforcement learning (RL) can optimize beyond this ceiling using external quality measures, but how the reward function should be designed when quality is multi-dimensional remains unexplored. We present a systematic investigation of reward function design for RL-based process model generation, training two LLM families (Llama 3.1 8B, Qwen 2.5 14B) under 48 configurations using Group Sequence Policy Optimization with rewards derived from an automated evaluation framework comprising 38 metrics across syntactic, pragmatic, and semantic quality. Three findings emerge. First, RL significantly improves pragmatic and syntactic quality while preserving semantic fidelity, reducing output variability by more than sixfold. Second, equal reward weighting consistently outperforms targeted weighting: emphasizing a specific dimension fails to improve it and can collapse the model into a low-quality mode. Third, design choices interact with model architecture in non-trivial ways: the invalidity penalty is essential for one model but irrelevant for the other, and SFT initialization is indispensable for one architecture but counterproductive for another. These results demonstrate that reward composition is a primary determinant of optimization outcomes, with effects as large as the decision to apply RL itself. The findings generalize to any structured generation task where quality is assessed along multiple automated dimensions. We release our implementation and experimental code at [https://github.com/chlauer99/RL_for_process_modeling](https://github.com/chlauer99/RL_for_process_modeling).  \n1. Introduction  \nThe creation of Business Process Model and Notation (BPMN) models is a complex task requiring both domain knowledge and proficiency in modeling conventions. As organizations increasingly seek to document, analyze, and automate their workflows, the demand for process models has outpaced the availability of trained modelers, creating a bottleneck that limits the scalability of business process management initiatives [19] . Large language models (LLMs) offer a promising path to addressing this bottleneck by enabling the generation of process models directly from natural-language descriptions, thereby lowering the expertise barrier and accelerating the modeling lifecycle [20] .  \nRecent work has demonstrated that LLMs can indeed produce BPMN-like structures from text, with interactive systems supporting iterative refinement [22, 26] and compact intermediate representations improving generation robustness [5, 21] . Systematic benchmarking has further shown that open-source LLMs achieve competitive syntactic and pragmatic quality on BPMN generation tasks, though semantic fidelity and output validity remain significant challenges [27] . Supervised fine-tuning (SFT) with parameterefficient methods such as LoRA [17] has been shown to improve structural correctness and validity, enabling smaller models to outperform larger untuned baselines. However,  \n∗Corresponding author  \n alexander_michael. [rombach@uni-saarland.de](rombach@uni-saarland.de) (A. Rombach); [chantale.lauer@dfki.de](chantale.lauer@dfki.de) (C. Lauer); [nijat.mehdiyev@dfki.de](nijat.mehdiyev@dfki.de) (N. Mehdiyev)  \nORCID(s): 0000-0002-91","cbCaidCN27iGCLh6","https://ap.wps.com/l/cbCaidCN27iGCLh6","pdf",683629,4,1,21,"English","en",105,"# Introduction\n# Reward Function Design for RL-Based Process Modeling\n## Evaluation Framework and Multi-Dimensional Quality","[{\"question\":\"Why does supervised fine-tuning (SFT) limit LLM-generated BPMN quality?\",\"answer\":\"SFT imitates patterns in the training data, so it cannot push the model beyond the demonstrations toward consistently better outputs under new criteria.\"},{\"question\":\"How does reinforcement learning (RL) improve BPMN process model generation here?\",\"answer\":\"RL optimizes the model directly against external automated quality measures derived from an evaluation framework, enabling improvements over SFT baselines.\"},{\"question\":\"What does the study find about reward function weighting and model behavior?\",\"answer\":\"Equal reward weighting outperforms targeted weighting; emphasizing a specific dimension may fail to improve it and can drive the model into a low-quality mode, with reward components interacting with model architecture in non-trivial ways.\"}]",1784184871,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"improving-llm-generated-process-model-quality-through-reinforcement-learning-the-role-of-reward-function-design","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/improving-llm-generated-process-model-quality-through-reinforcement-learning-the-role-of-reward-function-design/83050/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does supervised fine-tuning (SFT) limit LLM-generated BPMN quality?","Question",{"text":75,"@type":76},"SFT imitates patterns in the training data, so it cannot push the model beyond the demonstrations toward consistently better outputs under new criteria.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does reinforcement learning (RL) improve BPMN process model generation here?",{"text":80,"@type":76},"RL optimizes the model directly against external automated quality measures derived from an evaluation framework, enabling improvements over SFT baselines.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the study find about reward function weighting and model behavior?",{"text":84,"@type":76},"Equal reward weighting outperforms targeted weighting; emphasizing a specific dimension may fail to improve it and can drive the model into a low-quality mode, with reward components interacting with model architecture in non-trivial ways.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]