[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85837-en":3,"doc-seo-85837-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85837,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Knowledge-Conditioned, Single-Pass LLM Synthesis of Executable Unity Game Scenes: A Compiler Error Census across 26 Goal Playable Concepts","Large language models can generate Unity C# scripts for game scenes, but common demonstrations rely on an iterative repair loop that recompiles until the output runs. This work removes the loop and evaluates single-pass generation, treating the first draft as final to isolate what the model has learned in its parameters. Goal Playable Concepts are instantiated across 10,400 generations spanning multiple models, IR conditioning levels, and goal patterns. None compile into runnable scenes.","Knowledge-Conditioned, Single-Pass LLM Synthesis of Executable Unity Game Scenes: A Compiler Error Census across 26 Goal  \nPlayable Concepts  \nHugh Xuechen Liu and Kıvanc¸ Tatar  \narXiv :2607 . 10 187v 1 [ cs .LG] 11 Jul 2026  \nAbstract—Large language models (LLMs) write Unity C\\# for game scenes. Yet nearly all demonstrations rest on an iterative repair loop that regenerates code until it compiles, conflating what the model writes with what the loop fixes. We remove the loop and evaluate a single pass, where the first draft is final. This isolates the model’s parametric knowledge, the most stringent test of unaided generation. Models instantiate Goal Playable Concepts, playable counterparts of goal patterns, across 10,400 generations (four open-weight models, 7B–30B; two generation modes; four intermediate-representation (IR) conditioning levels;  \n26 goal patterns; 20 seeds). None compiled into a runnable scene, leaving no survivorship bias. To understand how the generated C\\# scripts fail, we categorize the 99 error codes behind 90,673 compiler-error occurrences as Grounding (invented or misused Unity types and APIs) or Hygiene (structural defects needing no Unity knowledge). The split differs sharply by goal pattern (e.g., Stealth fails mostly on invented engine references; Capture on plain C\\# structure). Larger models, stricter IRs, and different generation modes move the errors but never yield a compiling scene. The bottleneck is missing engine-specific knowledge. The census orders goal patterns by that demand, showing designers where single-pass generation breaks.  \nIndex Terms—error taxonomy, code generation, large language models, Unity, gameplay design patterns, goal playable concepts  \nI. INTRODUCTION  \nLARGE language models (LLMs) have become a practical  \ntool for game content generation. Demos and tutorials routinely show LLMs writing Unity C\\# scripts, placing game objects, and implementing mechanics on demand 1. Yet almost every such demonstration rests on an iterative repair loop: the model generates a candidate, the compiler errors (if any) return to a human or to the model itself, and the loop repeats until the artifact compiles and runs. This workflow is productive in practice but conflates what the model writes for the scene with what the repair loop fixes. Also, the gameplay on show is likewise picked ad hoc rather than drawn from a gameplay design vocabulary [1]–[4] .  \nWe ask a simpler and harder question: what is the intrinsic capability ceiling of single-pass LLM generation of executable game artifacts with no human feedback and no iterative repair? Single-pass evaluation isolates the model’s parametric knowledge (what is stored in its weights) from the  \nHugh Xuechen Liu and Kıvanc¸ Tatar are with Chalmers University of Technology and University of Gothenburg, SE-412 96 Gteborg, Sweden (email: [xuechen@chalmers.se](xuechen@chalmers.se); [tatar@chalmers.se](tatar@chalmers.se)) .  \n1 Some examples: [https://www.youtube.com/watch?v=gSFHyso](https://www.youtube.com/watch?v=gSFHyso) uuI; [https:](https:)//[github.com/keijiro/DungeonMatchHeroes](github.com/keijiro/DungeonMatchHeroes)  \npractitioner’s domain expertise. It is the most stringent and most diagnostic condition. In an iterative workflow, each error reflects the model plus the feedback it received. Removing the loop makes every error attributable to the model alone. Removing the loop also levels the comparison. Every model faces the same condition on every task. A difference in the error profile can be hence read as a difference in what the model knows or in what the task demands.  \nExisting game generation studies typically condition on genre labels (e.g., First-person Shooter, platformer) or freeform natural language. However, genre is criticised as a culturally constructed, semantically unstable concept that cannot serve as a reproducible evaluation target for cumulative research [5]–[7] . We instead ground our evaluation in Goal Playable Con","cbCaisAhTSI05WXy","https://ap.wps.com/l/cbCaisAhTSI05WXy","pdf",613872,3,1,21,"English","en",105,"# Introduction\n## Goal Playable Concepts (GPC)\n## Evaluation Setup and Generation Grid\n## Results and Error Taxonomy","[{\"question\":\"What problem does the paper address with existing LLM-to-Unity scene demos?\",\"answer\":\"Most demonstrations use an iterative repair loop that regenerates code until it compiles. The loop mixes the model’s initial output with the fixes applied during repair, obscuring the model’s true capability under single-pass generation.\"},{\"question\":\"How is the evaluation made more diagnostic in the paper?\",\"answer\":\"The paper evaluates a single generation pass with no human feedback and no iterative repair. The first draft is treated as final, so any compilation failure can be attributed to the model and the task demands under the same conditions for all models.\"},{\"question\":\"How are compilation errors categorized?\",\"answer\":\"Errors are counted and categorized into Grounding (invented or misused Unity types and APIs) and Hygiene (structural defects that do not require Unity-specific knowledge). The distribution varies by goal pattern, revealing where engine-specific knowledge is missing.\"}]",1784206625,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"knowledge-conditioned-single-pass-llm-synthesis-of-executable-unity-game-scenes-a-compiler-error-census-across-26-goal-playable-concepts","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/knowledge-conditioned-single-pass-llm-synthesis-of-executable-unity-game-scenes-a-compiler-error-census-across-26-goal-playable-concepts/85837/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address with existing LLM-to-Unity scene demos?","Question",{"text":75,"@type":76},"Most demonstrations use an iterative repair loop that regenerates code until it compiles. The loop mixes the model’s initial output with the fixes applied during repair, obscuring the model’s true capability under single-pass generation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the evaluation made more diagnostic in the paper?",{"text":80,"@type":76},"The paper evaluates a single generation pass with no human feedback and no iterative repair. The first draft is treated as final, so any compilation failure can be attributed to the model and the task demands under the same conditions for all models.",{"name":82,"@type":73,"acceptedAnswer":83},"How are compilation errors categorized?",{"text":84,"@type":76},"Errors are counted and categorized into Grounding (invented or misused Unity types and APIs) and Hygiene (structural defects that do not require Unity-specific knowledge). The distribution varies by goal pattern, revealing where engine-specific knowledge is missing.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]