[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86414-en":3,"doc-seo-86414-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86414,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Can AI Reason Like an Urban Planner? Benchmarking Large Language Models Against Professional Judgment","Large language models confront urban planning with an urgent epistemological test: which dimensions of professional planning knowledge can be replicated by AI, and which remain irreducibly human. The paper proposes Urban Planning Bench (UPBench), a domain-specific evaluation framework using a 4×5 matrix across four knowledge pillars and five Bloom-adapted cognitive levels. Results from 25 LLMs reveal non-monotonic reasoning, where higher-order analytical tasks outperform lower-order factual recall and integrative judgment due to context-, value-, and institution-bound planning knowledge.","arXiv :2606 . 11678v1 [ cs .CL] 10 Jun 2026  \nCan AI Reason Like an Urban Planner?  \nBenchmarking Large Language Models Against Professional Judgment  \nYijie Deng 1,2,* , He Zhu 1,2,* , Wen Wang 1,2 , Junyou Su 1,2 , Minxin Chen 1,2 , and Wenjia Zhang 1,3,†  \n1 Behavioral and Spatial AI Lab, Tongji University; 2 Behavioral and Spatial AI Lab, Peking University; 3 College of Architecture and Urban Planning, Tongji University  \n*Yijie Deng and He Zhu contributed equally to this work. †Corresponding author: Wenjia Zhang.  \nABSTRACT  \nProblem, Research Strategy, and Findings  \nThe emergence of large language models (LLMs) confronts urban planning with an urgent epistemological question: what dimensions of professional planning knowledge can artificial intelligence replicate, and what remains irreducibly human? Despite growing deployment of AI tools in planning practice, we lack systematic frameworks for evaluating whether these systems can reason with the contextual sensitivity, value awareness, and institutional literacy that characterize professional planning judgment. This paper introduces Urban Planning Bench (UPBench), a domainspecific evaluative framework that assesses LLM reasoning across a 4 ×5 matrix encompassing four knowledge pillars (Principles of Urban Planning, Cross-Disciplinary Integration, Planning Governance, and Planning Practice) and five cognitive levels adapted from Bloom’s revised taxonomy. Evaluating 25 LLMs through a dual-track protocol combining automated scoring with expert panel assessment, we identify an non-monotonic cognitive curve: models perform more robustly on higher-order analytical tasks than on ostensibly lower-order factual recall and integrative judgment. This counterintuitive finding reveals that planning’s “lower-order” knowledge is infact deeply embedded in institutional, jurisdictional, and temporal context—making it resistant to pattern-generalization strategies. We codify these limitations into four epistemic diagnostics—regulatory hallucination, conceptual conflation, wickedness paralysis, and phronetic deficit—each illuminating a specific dimension of planning expertise that resists computational replication.  \nTakeaway for Practice  \nThese findings provide planning practitioners and educators with an evidencebased framework for differential delegation—determining which tasks can be responsibly augmented by AI and which require irreducibly human professional judgment. LLMs demonstrate competence in cross-disciplinary synthesis and broad analytical reasoning, suggesting productive augmentation potential for literature review, scenario generation, and preliminary policy analysis. However, they exhibit persistent incapacity in jurisdiction-specific regulatory interpretation, normative conflict resolution, and context-sensitive procedural application—tasks that constitute the core of planning’s phronetic expertise. Planning agencies should implement structured verification protocols for any AI-assisted regulatory analysis, while planning education should reorient from knowledge transmission toward cultivating the institutional literacy, normative judgment, and contextual sensitivity that constitute planning’s distinctive professional contribution.  \nKEYWORDS  \nartificial intelligence; planning knowledge; professional judgment; phronesis; benchmark; large language models; planning education  \n1. Introduction  \nPlanning has long grappled with a foundational disciplinary question: what constitutes distinctively professional planning knowledge? Friedmann (1987) framed this as the problem of linking knowledge to action—identifying what planners know that enables them to intervene meaningfully in the trajectory of cities and regions. Schön (1992a) recast professional knowledge not as applied science but as reflection-in-action, a form of knowing embedded in practice that resists codification into rules. More recently, Flyvbjerg (2001) argued that planning’s core intellectual contribution lies ","cbCaigatXewHH7tl","https://ap.wps.com/l/cbCaigatXewHH7tl","pdf",1516312,6,1,32,"English","en",105,"# Abstract\n# Key Takeaway for Practice\n# Keywords\n# Introduction","[{\"question\":\"What problem does the paper address about AI and urban planning?\",\"answer\":\"It examines which aspects of professional planning knowledge LLMs can replicate and which aspects remain irreducibly human, focusing on contextual sensitivity, value awareness, and institutional literacy.\"},{\"question\":\"How does Urban Planning Bench (UPBench) evaluate LLM reasoning?\",\"answer\":\"UPBench assesses LLMs across four planning knowledge pillars and five cognitive levels, using a dual-track protocol with automated scoring plus expert panel assessment.\"},{\"question\":\"What key finding challenges expectations about LLM performance?\",\"answer\":\"The study finds a non-monotonic cognitive curve: models perform more robustly on higher-order analytical tasks than on lower-order factual recall and integrative judgment, implying resistance to pattern-generalization for planning’s contextual knowledge.\"}]",1784211594,81,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"can-ai-reason-like-an-urban-planner-benchmarking-large-language-models-against-professional-judgment","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/can-ai-reason-like-an-urban-planner-benchmarking-large-language-models-against-professional-judgment/86414/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper address about AI and urban planning?","Question",{"text":76,"@type":77},"It examines which aspects of professional planning knowledge LLMs can replicate and which aspects remain irreducibly human, focusing on contextual sensitivity, value awareness, and institutional literacy.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does Urban Planning Bench (UPBench) evaluate LLM reasoning?",{"text":81,"@type":77},"UPBench assesses LLMs across four planning knowledge pillars and five cognitive levels, using a dual-track protocol with automated scoring plus expert panel assessment.",{"name":83,"@type":74,"acceptedAnswer":84},"What key finding challenges expectations about LLM performance?",{"text":85,"@type":77},"The study finds a non-monotonic cognitive curve: models perform more robustly on higher-order analytical tasks than on lower-order factual recall and integrative judgment, implying resistance to pattern-generalization for planning’s contextual knowledge.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]