[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83907-en":3,"doc-seo-83907-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83907,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","On the Risk of Coding Before Testing: An Empirical Study on LLM-Based Test Generation Workflow","Large Language Models (LLMs) are increasingly used to generate both source code and corresponding test suites in software engineering workflows. This dual capability underpins test-first and agentic paradigms, yet it assumes that generated tests are independent, reliable oracles. The paper challenges this assumption by studying error propagation: faults in generated code are replicated in generated test artifacts, yielding mutually consistent but incorrect implementations and masking defects. Experiments quantify reduced fault detection effectiveness, show persistence across prompting strategies, and analyze effects in multi-step workflows.","On the risk of coding before testing: An empirical study on LLM-based test generation workflow  \nMichael Konstantinou  \nSnT, University of Luxembourg Luxembourg [michael.konstantinou@uni.lu](michael.konstantinou@uni.lu)  \nFlorian Tambon  \nSnT, University of Luxembourg Luxembourg [florian.tambon@uni.lu](florian.tambon@uni.lu)  \nMike Papadakis  \nSnT, University of Luxembourg Luxembourg [michail.papadakis@uni.lu](michail.papadakis@uni.lu)  \narXiv :2607 .05 139v 1 [ cs . SE] 6 Jul 2026  \nAbstract—Large Language Models (LLMs) are increasingly used in software engineering workflows to automatically generate both source code and corresponding test suites. This dual capability has enabled emerging development paradigms, including test-first and agentic workflows, where a single model is responsible for producing and validating implementations. However, these approaches implicitly assume that generated tests act as independent and reliable oracles—a fundamental requirement for effective software testing. In this paper, we challenge this assumption and investigate whether LLM-generated code biases the generation of subsequent tests. We introduce and empirically study the phenomenon of error propagation, where faults present in generated code are systematically replicated in the associated test artifacts. This leads to cases where incorrect implementationsand tests are mutually consistent, thereby masking defects rather than revealing them. We evaluate this effect across a range of programming tasks and agentic workflows, analyzing the consistency between generated code and test assertions, with particular focus on scenarios of aligned failures. Our study examines (i) whether erroneous code artifacts bias test generation,(ii) whether such bias persists under different prompting strategies, including chain-of-thought reasoning, and (iii) how errors propagate across multi-step workflows in which intermediate outputs are reused as context. The results show that error propagation is both prevalent and impactful: generating tests after faulty code significantly reduces fault detection effectiveness compared to generating tests independently (14% vs. 25%). These findings highlight a fundamental limitation of current workflows, where lack of independence between generated artifacts undermines the reliability of automated testing. Furthermore, our results exposea previously underexplored threat to validity in empirical studies that rely on coupled generation pipelines.  \nIndex Terms—LLM-based test generation, error propagation  \nI. INTRODUCTION  \nLarge Language Models (LLMs) are increasingly integrated into software engineering workflows, where they are used not only to generate production code but also to synthesize accompanying test suites. This dual capability has fueled optimistic visions of highly automated development pipelines, in which both implementation and validation are delegated to the same generative system. In particular, recent toolchains promote test-first or co-evolutionary workflows, where LLMs generate candidate tests and subsequently produce code that satisfies them, potentially accelerating development while reducing human effort. However, this paradigm implicitly assumes that LLM-generated tests provide an independent and reliable oracle for assessing correctness. In traditional software testing,  \neffectiveness critically depends on the independence between the system under test and the test oracle [1], [2] . When both artifacts are produced by the same underlying model, this assumption may no longer hold.  \nIn this paper, we investigate this fundamental assumption by examining whether the generation of faulty code by an LLM biases the generation of tests. When tasked with producing both artifacts for the same problem, an LLM is likely to rely on similar internal representations, assumptions, and reasoning trajectories. This shared generative process increases the likelihood that errors are not independent but systematica","cbCaiu7BPlYnpFtf","https://ap.wps.com/l/cbCaiu7BPlYnpFtf","pdf",610640,4,1,12,"English","en",105,"# Abstract\n# Introduction\n## Motivation and oracle independence\n## Error propagation mechanism\n## Implications for agentic workflows and validity threats","[{\"question\":\"What is the core assumption about LLM-generated tests that the paper challenges?\",\"answer\":\"The paper challenges the assumption that LLM-generated tests function as independent and reliable oracles when code and tests are produced through a coupled generative process.\"},{\"question\":\"What does the paper mean by error propagation?\",\"answer\":\"Error propagation refers to the systematic replication of faults from LLM-generated code into associated generated test artifacts, causing tests and implementations to agree on incorrect behavior.\"},{\"question\":\"How does generating tests after faulty code affect fault detection effectiveness?\",\"answer\":\"The results show that generating tests after faulty code significantly reduces fault detection effectiveness compared with generating tests independently (14% vs. 25%).\"}]",1784191366,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"on-the-risk-of-coding-before-testing-an-empirical-study-on-llm-based-test-generation-workflow","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/on-the-risk-of-coding-before-testing-an-empirical-study-on-llm-based-test-generation-workflow/83907/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the core assumption about LLM-generated tests that the paper challenges?","Question",{"text":75,"@type":76},"The paper challenges the assumption that LLM-generated tests function as independent and reliable oracles when code and tests are produced through a coupled generative process.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the paper mean by error propagation?",{"text":80,"@type":76},"Error propagation refers to the systematic replication of faults from LLM-generated code into associated generated test artifacts, causing tests and implementations to agree on incorrect behavior.",{"name":82,"@type":73,"acceptedAnswer":83},"How does generating tests after faulty code affect fault detection effectiveness?",{"text":84,"@type":76},"The results show that generating tests after faulty code significantly reduces fault detection effectiveness compared with generating tests independently (14% vs. 25%).","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]