[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85574-en":3,"doc-seo-85574-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85574,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs","Bayesian Linguistic Forecaster (BLF) is an agentic system for binary forecasting that delivers state-of-the-art results on the ForecastBench benchmark. BLF combines a semi-structured linguistic belief state updated by an LLM at each tool-use step, multi-trial aggregation via logit-space averaging with shrinkage, and hierarchical calibration using Platt scaling with a hierarchical prior. Experiments on 400 backtesting questions show consistent superiority across question types, supported by ablations and mixed-effects analysis controlling for variability, plus a robust back-testing framework with low leakage.","arXiv :2604 . 18576v4 [ cs .AI] 12 Jul 2026  \nAgentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs  \nKevin Murphy  \nDepartment of Computer Science  \nUniversity of British Columbia  \nVancouver, BC, Canada  \n[murphyk@cs.ubc.ca](murphyk@cs.ubc.ca)  \nAbstract  \nWe present the Bayesian Linguistic Forecaster (BLF), an agentic system for binary forecasting that achieves state-of-the-art performance on the ForecastBench benchmark. The system is built on three ideas. (1) Linguistic belief state: a semi-structured representation combining numerical probability estimates with natural-language evidence summaries, updated by the LLM at each step of an iterative tool-use loop. This contrasts with the common approach of appending all retrieved evidence to an ever-growing, unstructured context. (2) Hierarchical multi-trial aggregation: running K independent trials and combining them using logit-space averaging shrinkage with a data-dependent prior. (3) Hierarchical calibration: Platt scaling with a hierarchical prior, which avoids over-shrinking extreme predictions for sources with skewed base rates. On 400 questions from the ForecastBench leaderboard, BLF outperforms all the top public methods, including Cassi, GPT-5, Grok 4.20, and Foresight-32B. Careful ablation studies, using mixed effects analysis to control for question variability (which accounts for 62% of the variance in performance), reveals that all 3 components contribute to the overall gains, but some components matter more than others, depending on the base LLM, and the setting (e.g. with or without a crowd prior) . All our experiments are based on a robust back-testing framework which we develop, which has a leakage rate below 1.5%, and may be of independent interest.  \n1 Introduction  \nForecasting the probability of future events is a fundamental challenge with applications in geopolitics, finance, and public health [Tetlock and Gardner, 2015, Spiegelhalter, 2025] . Recent work has shown that LLMs can approach human-level forecasting when given web search access [Halawi et al., 2024], and benchmarks such as ForecastBench [Karger et al., 2025] provide standardized evaluation with online leaderboards. We present BLF (Bayesian Linguistic Forecaster), an agentic system that achieves a new state-of-the-art (SOTA) performance on ForecastBench. Our approach is organized around three key ideas:  \n1. Linguistic belief states. Most forecasting agents either search in parallel then reason once, or sequentially accumulate raw search results in context. BLF instead maintains a semi-structured belief state—a probability estimate paired with natural-language evidence summaries—updated by the LLM at each step. We refer to this loosely as “Bayesian-style” updating: the slots are designed to mirror the form of a sequential Bayesian update (prior + evidence → posterior), but the actual update is an LLM forward pass, which may not satisfy the consistency criteria required to correspond to proper Bayesian inference Qiu et al. [2025], Falck et al. [2024] . (In Sec. J, we evaluated a more  \nPreprint.  \ntraditional Bayesian approach based on sequential updating with explicit LLM-estimated likelihoods, but it was much worse.)  \n2. Multi-trial aggregation. LLM forecasting exhibits high variance across runs. We run K=5 independent trials and aggregate by averaging in logit space. We also explore hierarchical shrinkage toward the empirical or uniform prior (inspired by James-Stein / empirical Bayes), which further helps on datasets with high trial variance.  \n3. Hierarchical calibration. To ensure the forecasts are calibrated, we use Platt scaling [Platt, 1999] . However, global Platt scaling can over-shrink well-calibrated extreme predictions. We use hierarchical Platt scaling with per-source intercept offsets, which is critical when empirical priors produce source-specific biases, especially in the zero-shot setting.  \nOn 400 backtesting questions from ForecastBench (FB), BLF ac","cbCaidwxxmLFOCg0","https://ap.wps.com/l/cbCaidwxxmLFOCg0","pdf",5159297,2,1,64,"English","en",105,"# Abstract\n# Introduction\n# Experimental Setup\n## Problem definition","[{\"question\":\"What is the Bayesian Linguistic Forecaster (BLF) designed to do?\",\"answer\":\"BLF is an agentic system for binary forecasting, estimating event probabilities in a sequential interaction loop and achieving state-of-the-art performance on ForecastBench.\"},{\"question\":\"How does BLF represent and update forecasting information over time?\",\"answer\":\"BLF maintains a semi-structured linguistic belief state that pairs probability estimates with natural-language evidence summaries, which the LLM updates at each step of an iterative tool-use loop.\"},{\"question\":\"What techniques does BLF use to improve forecast accuracy and calibration?\",\"answer\":\"BLF uses hierarchical multi-trial aggregation in logit space with shrinkage and performs hierarchical calibration via Platt scaling with a hierarchical prior to avoid over-shrinking extreme predictions.\"}]",1784204695,161,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"agentic-forecasting-using-sequential-bayesian-updating-of-linguistic-beliefs","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/agentic-forecasting-using-sequential-bayesian-updating-of-linguistic-beliefs/85574/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the Bayesian Linguistic Forecaster (BLF) designed to do?","Question",{"text":75,"@type":76},"BLF is an agentic system for binary forecasting, estimating event probabilities in a sequential interaction loop and achieving state-of-the-art performance on ForecastBench.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does BLF represent and update forecasting information over time?",{"text":80,"@type":76},"BLF maintains a semi-structured linguistic belief state that pairs probability estimates with natural-language evidence summaries, which the LLM updates at each step of an iterative tool-use loop.",{"name":82,"@type":73,"acceptedAnswer":83},"What techniques does BLF use to improve forecast accuracy and calibration?",{"text":84,"@type":76},"BLF uses hierarchical multi-trial aggregation in logit space with shrinkage and performs hierarchical calibration via Platt scaling with a hierarchical prior to avoid over-shrinking extreme predictions.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]