[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84006-en":3,"doc-seo-84006-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84006,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Can Large Language Models Generate Observability-Aware Code?","Recent advances in coding agents enable generation of complex software, yet most evaluations emphasize functional correctness while production use requires failure evidence for observability. This study examines whether agents can restore source-level diagnostic semantics by reintroducing observability artifacts in 10 open-source and 8 industrial repositories. Runtime effectiveness is tested on 200 Kubernetes-deployed microservice systems with 13 injected faults.","Can Large Language Models Generate Observability-Aware Code?  \nYongliang Tao 1 , Hongyu Zhang 1 , Pengfei Gao2 , Minghua Ma2 , Zhiyu Fan2 , Yu Kang2 Jue Zhang2 , Si Qin2 , Liqun Li2 , Qingwei Lin2 , Saravan Rajmohan2  \n1 Chongqing University, Chongqing, China  \n2Microsoft  \narXiv :2607 .05785v 1 [ cs . SE] 7 Jul 2026  \nAbstract—Recent advances in coding agents have enabled the generation of increasingly complex software systems. While existing evaluations primarily focus on functional correctness, production systems must expose failure evidence to support observability. In this paper, we present a systematic study of observability in agent-generated systems. We examine whether agents can reconstruct source-level diagnostic semantics by restoring observability artifacts in 10 open-source and 8 industrial repositories. We also evaluate whether these artifacts translate into effective fault signals at runtime through 200 generated microservice systems deployed on Kubernetes with 13 injected faults. Our results reveal a consistent gap between diagnostic semantics at the source level and fault signals (i.e., explicit, fault-specific evidence) at runtime. At the source level, agents partially recover observability artifacts but struggle to capture key diagnostic semantics. At runtime, generated systems expose fault signals for only a small fraction of failures (up to 13.99%), despite the presence of logging, suggesting that the generated observability artifacts may lack the failure-specific semantics needed to effectively expose faults. We further introduce anobservability-oriented skill, which can serve as a guidance to improve both diagnostic semantics and fault-signal exposure, but the gains remain limited, indicating that the gap is not easily addressed. More broadly, our findings suggest that current evaluations focusing primarily on functional correctness may overlook observability as an important dimension of practical software quality.  \nIndex Terms—Coding Agent, Fault Signal, Observability  \nI. INTRODUCTION  \nRecent advances in coding agents, such as Copilot [1], Cursor [2], and Claude Code [3], have made it increasingly feasible to generate complete software systems. These systems may compile, deploy, and run successfully under normal workloads. However, software that is runnable is not necessarily software that is operable. When failures occur, developers need sufficient information to understand what happened. This shift from runnability to operability brings failure-time observability into focus. Beyond generating runnable code, agents should also generate observability artifacts (i.e., logs, traces, and metrics) that expose fault-indicating runtime evidence.  \nIn human-written software systems, even though developers may have a substantial understanding of the system, unexpected production failures still occur, making observability a fundamental need for debugging.  \nCoding agents further increase developers’ need for observability. While they can generate large amounts of code quickly, developers may not inspect or reason about all of the generated code in detail. As a result, developers may not fully understand the runtime behavior of the generated system. This creates knowledge debt: developers may need to maintain generated systems whose runtime behavior they do not fully understand.  \nKnowledge debt makes observability more difficult to improve after systems are deployed. When developers lack comprehensive understanding of the runtime behavior of agentgenerated systems, they may lack the semantic context needed to determine what diagnostic semantics (i.e., failure-relevant context captured by observability artifacts) should be encoded. This is because observability design requires reasoning about potential failure modes and identifying relevant runtime states, which depends on a deep understanding of system behavior. Consequently, post-hoc instrumentation may fail to capture a failure-relevant context and, therefo","cbCaid7ITozv48gD","https://ap.wps.com/l/cbCaid7ITozv48gD","pdf",845083,5,1,12,"English","en",105,"# Introduction\n## Motivation: from runnability to operability\n## Knowledge debt and observability limitations\n## Observability as a generation-time expectation\n# Evaluation Approach\n## Static restoration of observability artifacts\n## Runtime fault-signal effectiveness","[{\"question\":\"What gap does the study find between source-level observability and runtime fault signals?\",\"answer\":\"Generated systems show a consistent gap: agents partially recover observability artifacts at the source level but struggle to capture key diagnostic semantics, resulting in explicit fault signals for only a small fraction of failures at runtime.\"},{\"question\":\"How do the researchers test observability in agent-generated systems?\",\"answer\":\"They conduct a source-level restoration study by removing existing observability artifacts and asking agents to restore them using remaining repository context, then evaluate runtime impact using 200 generated microservice systems deployed on Kubernetes with 13 injected faults.\"},{\"question\":\"Why do the authors argue observability should be handled during code generation rather than after deployment?\",\"answer\":\"They suggest that improving observability post-hoc is difficult when developers lack semantic context about failure modes and relevant runtime states, so instrumentation may fail to expose the corresponding fault signals during failures.\"}]",1784191982,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"can-large-language-models-generate-observability-aware-code","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/can-large-language-models-generate-observability-aware-code/84006/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What gap does the study find between source-level observability and runtime fault signals?","Question",{"text":76,"@type":77},"Generated systems show a consistent gap: agents partially recover observability artifacts at the source level but struggle to capture key diagnostic semantics, resulting in explicit fault signals for only a small fraction of failures at runtime.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How do the researchers test observability in agent-generated systems?",{"text":81,"@type":77},"They conduct a source-level restoration study by removing existing observability artifacts and asking agents to restore them using remaining repository context, then evaluate runtime impact using 200 generated microservice systems deployed on Kubernetes with 13 injected faults.",{"name":83,"@type":74,"acceptedAnswer":84},"Why do the authors argue observability should be handled during code generation rather than after deployment?",{"text":85,"@type":77},"They suggest that improving observability post-hoc is difficult when developers lack semantic context about failure modes and relevant runtime states, so instrumentation may fail to expose the corresponding fault signals during failures.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]