[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83962-en":3,"doc-seo-83962-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83962,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Retrieval over Reasoning A Cost-Controlled Benchmark of Language Models for Energy Retrofit Recommendation","Recommending the correct set of energy conservation measures (ECMs) for a building is a structured multi-label prediction challenge: task-specific supervised models have limited training signal, while general language models lack grounding in the local building stock. This study analyzes 10,422 real New York City Local Law 87 (LL87) energy-audit records, using certified auditors’ recommended ECM categories as ground truth. The paper establishes tree ensembles for EUI prediction, studies task framing effects, and evaluates eight LLMs with retrieval and explicit reasoning under a cost-controlled benchmark.","Retrieval over Reasoning: A Cost-Controlled Benchmark of Language Models for Energy  \nRetrofit Recommendation  \nEliseo Curcio  \nAbstract  \nRecommending the correct set of energy conservation measures (ECMs) for a building is a structured, multi-label prediction problem in which a task-specific supervised model has weak training signal and a general language model has no grounding in the local building stock. We study this problem on 10,422 real New York City Local Law 87 (LL87) energy-audit records, taking as ground truth the set ofECM categories that certified auditors actually recommended. We make four contributions. First, we establish that energyuse-intensity (EUI) prediction the upstream task is effectively solved by tree ensembles: across fifteen trained models, a stacking ensemble reaches a coefficient of determination R² = 0.757, and every one of six neural architectures is outperformed by gradient-boosted trees. Second, we show that the framing of the recommendation task dominates model choice: recasting ECM recommendation as 19-way multi-label classification rather than single-label categorization lifts a gradient-boosted-tree baseline from a previously reported 25.9% accuracy to a micro-F1 of 0.571. Third, we benchmark eight large language models (LLMs) from four providers in a 2×2 design that independently toggles retrieval grounding and explicit reasoning, scoring each arm on per-label F1, U.S.-dollar cost per building, and latency; retrieval-augmented generation (RAG) improves micro-F1 by +0.11 to +0.20 on every model, while explicit reasoning yields no measurable accuracy change (−0.018 to +0.010) at up to 8.4× the cost. Fourth, we show LLMs systematically overrecommend high recall, low precision and that retrieval closes the gap chiefly by improving precision. A 70-billion-parameter open-weight model with a fifteen-line nearest-neighbor retrieval step reaches 0.511 micro-F1 at $0.00032 per building, comparable to a frontier model at roughly 10.1× lower cost.  \nKeywords: LLM evaluation; retrieval-augmented generation; structured prediction; multi-label classification; gradient boosting; building energy efficiency.  \n1. Introduction  \nBuildings account for roughly 40% of primary energy use in developed economies, and deciding which retrofit measures a given building requires is a precondition for any large-scale efficiency program [1] . In New York City this judgment is produced by certified energy auditors under Local Law 87 (LL87), which, together with the Local Law 84 (LL84) benchmarking mandate, requires large buildings to disclose energy use and undergo periodic audits [2]. Auditing is costly and slow, so a system that could reproduce auditor recommendations from already-available building characteristics would let program managers triage a portfolio before committing inspection budgets. Prior data-driven work on this corpus has focused on predicting energy use and assigning performance grades [3] [4], and broad reviews confirm that machine learning is now the dominant approach to building-energy estimation [5][6] .  \nECM recommendation is intrinsically a structured-prediction problem: a building does not receive a single fix but a set of measures. We make three observations. First, the upstream EUI-prediction task is, on this data, effectively solved by tree ensembles such as XGBoost [7], LightGBM [8], and CatBoost [9]; the open  \nproblem is recommending the measures themselves. Second, the framing of the recommendation task matters more than the model: a prior single-label, twelve-class formulation reported only 25.9% accuracy, whereas multi-label classification one binary decision per ECM category lifts the same supervised family to a micro-F1 of 0.571. Third, general LLMs are attractive because they require no task-specific training, but without grounding they cannot know what auditors in this stock recommend; retrieval-augmented generation (RAG) [10] supplies that grounding, while the value of explicit ","cbCaitfzkBVUJhmD","https://ap.wps.com/l/cbCaitfzkBVUJhmD","pdf",951896,2,1,22,"English","en",105,"# Introduction\n# Related Work\n## Machine learning for building energy\n## Tree ensembles and deep tabular models","[{\"question\":\"Why is ECM recommendation treated as a structured multi-label problem?\",\"answer\":\"A building receives a set of retrofit measures rather than a single fix, so the task predicts multiple ECM categories. The work frames recommendation as multi-label structured prediction.\"},{\"question\":\"What role does retrieval-augmented generation (RAG) play in the benchmarked LLMs?\",\"answer\":\"RAG provides grounding from the local audit context, improving accuracy across models. The paper reports a micro-F1 gain of about +0.11 to +0.20 on every evaluated LLM.\"},{\"question\":\"How does explicit reasoning affect performance and cost?\",\"answer\":\"Explicit reasoning yields no measurable accuracy change and mainly increases cost. The study reports cost rising up to 8.4× with accuracy shifts in the range of −0.018 to +0.010.\"}]",1784191679,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"retrieval-over-reasoning-a-cost-controlled-benchmark-of-language-models-for-energy-retrofit-recommendation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/retrieval-over-reasoning-a-cost-controlled-benchmark-of-language-models-for-energy-retrofit-recommendation/83962/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is ECM recommendation treated as a structured multi-label problem?","Question",{"text":75,"@type":76},"A building receives a set of retrofit measures rather than a single fix, so the task predicts multiple ECM categories. The work frames recommendation as multi-label structured prediction.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What role does retrieval-augmented generation (RAG) play in the benchmarked LLMs?",{"text":80,"@type":76},"RAG provides grounding from the local audit context, improving accuracy across models. The paper reports a micro-F1 gain of about +0.11 to +0.20 on every evaluated LLM.",{"name":82,"@type":73,"acceptedAnswer":83},"How does explicit reasoning affect performance and cost?",{"text":84,"@type":76},"Explicit reasoning yields no measurable accuracy change and mainly increases cost. The study reports cost rising up to 8.4× with accuracy shifts in the range of −0.018 to +0.010.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]