[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85441-en":3,"doc-seo-85441-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85441,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","Beyond Naïve Prompting: Strategies for Improved Context-aided Forecasting with LLMs","Real-world forecasting must combine historical signals with contextual information expressed in text, yet existing context-aided LLM approaches face three gaps: limited diagnostic understanding of why models fail, accuracy well below potential, and high compute that restricts deployment. This work proposes a unified framework of four strategies spanning model diagnostics, accuracy, and efficiency. Extensive evaluation across model families from small open-source models to frontier systems (Gemini, GPT, Claude) yields: diagnostic “Execution Gap” insights, 25–50% accuracy gains, and adaptive routing that approaches large-model accuracy while reducing inference costs, enabling flexible practitioner integration.","arXiv :2508 .09904v 3 [ cs .LG] 11 Jul 2026  \nBeyond Naïve Prompting: Strategies for Improved Contextaided Forecasting with LLMs  \nArjun Ashok  \nServiceNow Research, Mila - Québec AI Institute, Université de Montréal  \nAndrew R. Williams  \nServiceNow Research, Mila - Québec AI Institute, Université de Montréal  \nVincent Zhihao Zheng  \nServiceNow Research, McGill University  \nIrina Rish  \nMila - Québec AI Institute, Université de Montréal  \nNicolas Chapados  \nServiceNow Research, Mila - Québec AI Institute, Polytechnique Montréal  \nÉtienne Marcotte  \nServiceNow Research  \nValentina Zantedeschi  \nServiceNow Research, Université Laval  \nAlexandre Drouin  \n[arjun.ashok.psg@gmail. com](arjun.ashok.psg@gmail. com)  \n[andrew.williams@umontreal. ca](andrew.williams@umontreal. ca)  \n[z.vincent.zheng@gmail. com](z.vincent.zheng@gmail. com)  \n[irina.rish@umontreal. ca](irina.rish@umontreal. ca)  \n[nicolas.chapados@mila.quebec](nicolas.chapados@mila.quebec)  \n[etienne. marcotte@gmail. com](etienne. marcotte@gmail. com)  \n[valentina.zantedeschi@servicenow. com](valentina.zantedeschi@servicenow. com)  \n[alexandre. drouin@servicenow. com](alexandre. drouin@servicenow. com)  \nServiceNow Research, Mila - Québec AI Institute, Université Laval  \nReviewed on OpenReview: [https: // openreview. net/ forum? id= dkjHHFJkVI](https: // openreview. net/ forum? id= dkjHHFJkVI)  \nAbstract  \nReal-world forecasting requires models to integrate not only historical data but also relevant contextual information provided in textual form. While large language models (LLMs) show promise for context-aided forecasting, critical challenges remain: we lack diagnostic tools to understand failure modes, performance remains far below their potential, and high computational costs limit practical deployment. We introduce a unified framework of four strategies that address these limitations along three orthogonal dimensions: model diagnostics, accuracy, and efficiency. Through extensive evaluation across model families from small open-source models to frontier models including Gemini, GPT, and Claude, we uncover both fundamental insights and practical solutions. Our findings span three key dimensions: diagnostic strategies reveal the “Execution Gap” where models correctly explain how context affects forecasts but fail to apply this reasoning; accuracy-focused strategies achieve substantial performance improvements of 25-50%; and efficiency-oriented approaches show that adaptive routing between small and large models can approach large model accuracy on average while significantly reducing inference costs. These orthogonal strategies can be flexibly integrated based on deployment constraints, providing practitioners with a comprehensive toolkit for practical LLM-based context-aided forecasting. Code is made available at [https://github.com/ashok-arjun/beyond-naive-prompting](https://github.com/ashok-arjun/beyond-naive-prompting).  \n1 Introduction  \nProbabilistic time series forecasting is essential for optimal decision-making, involving predicting the evolution of various quantities over time, as well as estimating the likelihood of various scenarios (Hyndman & Athanasopoulos, 2021; Peterson, 2017) . This problem has been extensively studied by both the statistical and machine learning communities (Hyndman et al., 2008; Box et al., 2015; Hyndman & Athanasopoulos, 2021), culminating in different methods such as classical methods (Hyndman et al., 2008; Gardner Jr., 1985), deep learning methods (Salinas et al., 2020; Drouin et al. , 2022; Ashok et al., 2024), hybrid methods (Oreshkin et al., 2019), and more recently, foundation models (Rasul et al. , 2023; Ansari et al. , 2024; Woo et al. , 2024) . Research in forecasting has largely focused on building models that use nu-  \nFigure 1: Scope of our study. We propose four complementary strategies that extend naïve Direct Prompting (DP) (Williamset al., 2025) along different dimensions. FxDP (top) enables model diagnostics by eli","cbCaimDG8wguNhih","https://ap.wps.com/l/cbCaimDG8wguNhih","pdf",7286834,3,1,93,"English","en",105,"# Introduction\n## Context-aided forecasting with LLMs\n## Limitations of existing prompting approaches\n## Overview of four proposed strategies","[{\"question\":\"What problem does the paper address in context-aided forecasting with LLMs?\",\"answer\":\"It addresses how to incorporate textual context into probabilistic forecasting while overcoming three limitations: lack of diagnostic tools for failure modes, insufficient accuracy, and high computational cost for deployment.\"},{\"question\":\"What are the four strategies proposed in the unified framework?\",\"answer\":\"The paper introduces four complementary strategies: FxDP for forecast-effect explanations and diagnostics, RouteDP for adaptive model routing to reduce inference cost, and IC-DP and CorDP to substantially improve forecasting accuracy, especially for smaller models.\"},{\"question\":\"How do the proposed strategies improve accuracy and efficiency?\",\"answer\":\"Accuracy-focused approaches deliver substantial performance improvements of about 25–50%. Efficiency-oriented methods use adaptive routing between small and large models to approach large-model accuracy on average while significantly reducing inference costs.\"}]",1784203544,234,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"beyond-naive-prompting-strategies-for-improved-context-aided-forecasting-with-llms","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/beyond-naive-prompting-strategies-for-improved-context-aided-forecasting-with-llms/85441/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in context-aided forecasting with LLMs?","Question",{"text":75,"@type":76},"It addresses how to incorporate textual context into probabilistic forecasting while overcoming three limitations: lack of diagnostic tools for failure modes, insufficient accuracy, and high computational cost for deployment.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the four strategies proposed in the unified framework?",{"text":80,"@type":76},"The paper introduces four complementary strategies: FxDP for forecast-effect explanations and diagnostics, RouteDP for adaptive model routing to reduce inference cost, and IC-DP and CorDP to substantially improve forecasting accuracy, especially for smaller models.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the proposed strategies improve accuracy and efficiency?",{"text":84,"@type":76},"Accuracy-focused approaches deliver substantial performance improvements of about 25–50%. Efficiency-oriented methods use adaptive routing between small and large models to approach large-model accuracy on average while significantly reducing inference costs.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]