[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81711-en":3,"doc-seo-81711-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81711,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","The Limits of LLM Forecasting: Parametric Knowledge Gaps Across Conflict Zones","Media coverage of armed conflict is uneven, creating a 224× gap between the most and least covered conflict zones in English-language media across 22 countries (2020–2026). Zero-shot conflict escalation forecasting is evaluated on a 660-case held-out set comparing Llama-3.3-70B and GPT-4o against structured baselines. Results show qualitative failure: LLMs categorize rather than forecast; logistic regression with temporal-window features outperforms them. Coverage-stratified benchmarking and better under-covered conflict datasets are proposed.","The Limits of LLM Forecasting: Parametric Knowledge Gaps Across Conflict Zones  \nPoli Nemkova  \nUniversity of North Texas, Department of Computer Science and Engineering  \n[poli.nemkova@unt.edu](poli.nemkova@unt.edu)  \narXiv :2607 .00018v1 [ cs .CY] 29 May 2026  \nAbstract  \nMedia coverage of armed conflict is deeply asymmetric: we document a 224× gap between the most and least covered conflict zones in English-language media across 22 countries (2020–2026) . We evaluate zero-shot conflict escalation forecasting across all 22 countries on a 660-case held-out test set, comparing Llama- 3.3-70B and GPT-4o against three structured baselines.  \nThe central finding is not a performance gradient but a qualitative failure: LLMs do not forecast conflict—they categorize it. Llama predicts escalation on every under-covered case, matching the trivial Always-YES baseline to three decimals; GPT-4o predicts NO on every over-covered case, missing all five actual escalation events. A logistic regression using only eleven observation-window features with no country information achieves F1 = 0.402, outperforming both LLMs in every measurable tier. This failure cannot be resolved at inference time: adding structured ACLED evidence degrades performance on under-covered zones (GPT-4o F1: 0 .323 → 0. 168) and falls below LR by a factor of 2.4 . The bottleneck is not data availability but the LLM’s interpretation of temporal signal under a country-categorical prior.  \nUnder-covered populations receive not just less accurate AI, but qualitatively different AI that cannot distinguish stable from escalating periods. We call for coverage-stratified benchmarking, conflict NLP datasets for under-covered zones, and training data documentation standards for geographic conflict representation.  \n1 Introduction  \nAI systems for humanitarian decision support increasingly rely on LLMs for conflict early warning, displacement forecasting, and resource allocation (Karamolegkou et al., 2026 ; Rost and Ronco, 2026 ; UNHCR, 2022) . Yet these systems inherit world  \nknowledge from training corpora shaped by highly uneven English-language media coverage. This paper asks whether media attention asymmetries translate into asymmetries in how LLMs reason about conflict, and what this means for humanitarian AI.  \nWestern media bias in conflict and disaster coverage is well documented. Prior work shows that proximity, cultural consonance, elite-nation involvement, and framing shape which crises become salient (Galtung and Ruge, 1965 ; Entman, 1993) . Quantitative studies further show that news attention affects real-world aid allocation: Eisensee and Strömberg (2007) found that African disasters require far more casualties than comparable Eastern European disasters to receive similar US television coverage. Recent computational work confirms persistent under-coverage of conflicts in Sub-Saharan Africa and Southeast Asia relative to Europe and the Middle East, even after accounting for conflict intensity (Croicu and von der Maase, 2024) . What remains unknown is whether such asymmetries propagate into LLM-based conflict reasoning.  \nThis question matters because LLMs are increasingly considered as reasoning engines for forecasting, while prior conflict prediction work has mostly relied on structured datasets and hand-engineered features (Ward et al., 2013 ; Blair and Sambanis, 2020) . At the same time, NLP research shows that models inherit biases from web-scale corpora (Bender et al., 2021 ; Dodge et al., 2021 ; Gebru et al., 2021), including geographic bias in prediction tasks (Manvi et al., 2024), multilingual instability (Wu and Dredze, 2020 ; Nemkova et al., 2025a), and downstream failures in humanitarian contexts (Nemkova et al., 2025b) . However, geographic conflict coverage asymmetry has not been studied as a source of qualitative reasoning failure.  \nWe make five contributions. First, using ACLED event data (Raleigh et al., 2010) and GDELT media coverage (Leetaru and","cbCaijc18ACZlXX8","https://ap.wps.com/l/cbCaijc18ACZlXX8","pdf",379791,4,1,13,"English","en",105,"# Introduction\n# Measuring the Gap\n## Data Sources\n## Media Attention Ratio\n# Forecasting and Evaluation","[{\"question\":\"What is the main issue with LLMs in conflict escalation forecasting found in this work?\",\"answer\":\"LLMs do not truly forecast conflict dynamics; they categorize conflict contexts. This shows up as systematic Always-YES or Always-NO behavior depending on coverage levels.\"},{\"question\":\"How large is the media attention gap across conflict zones, and how is it measured?\",\"answer\":\"A 224× gap is observed between the most and least covered zones. It is measured as the ratio of English-language GDELT news articles per ACLED conflict event (2020–2026).\"},{\"question\":\"Why can structured evidence added at inference time fail to improve performance?\",\"answer\":\"Adding ACLED evidence can degrade performance on under-covered zones, because the bottleneck is not data availability but how LLMs interpret temporal signals under country-level categorical priors.\"}]",1784175564,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-limits-of-llm-forecasting-parametric-knowledge-gaps-across-conflict-zones","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/the-limits-of-llm-forecasting-parametric-knowledge-gaps-across-conflict-zones/81711/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main issue with LLMs in conflict escalation forecasting found in this work?","Question",{"text":75,"@type":76},"LLMs do not truly forecast conflict dynamics; they categorize conflict contexts. This shows up as systematic Always-YES or Always-NO behavior depending on coverage levels.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How large is the media attention gap across conflict zones, and how is it measured?",{"text":80,"@type":76},"A 224× gap is observed between the most and least covered zones. It is measured as the ratio of English-language GDELT news articles per ACLED conflict event (2020–2026).",{"name":82,"@type":73,"acceptedAnswer":83},"Why can structured evidence added at inference time fail to improve performance?",{"text":84,"@type":76},"Adding ACLED evidence can degrade performance on under-covered zones, because the bottleneck is not data availability but how LLMs interpret temporal signals under country-level categorical priors.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]