[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85492-en":3,"doc-seo-85492-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85492,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","BizFinBench.v2 Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation","Large language models are increasingly important for financial applications, yet common benchmarks rely on simulated or generic data, creating a large gap between reported scores and real-world usefulness. BizFinBench.v2 introduces an integrated offline and online benchmark built from authentic Chinese and U.S. equity user query-response data. It contains 28,860 questions over eight offline and two online tasks, enabling evaluation of core business capabilities alongside online performance, and supporting more reliable deployment decisions.","arXiv :2601 .0640 1v2 [ cs .AI] 13 Jul 2026  \nBizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation  \nXin Guo 1 2 * Rongjunchen Zhang 1 * † Guilong Lu 1 Xuntao Guo 1 Shuai Jia 1 Zhi Yang 2 Liwen Zhang 2 †  \nAbstract  \nLarge language models are becoming increasingly significant in financial applications. Nevertheless, prevailing benchmarks are largely dependent on simulated or generic data, which leads to a significant gap between reported performance and actual efficacy in real-world scenarios. To tackle this challenge, we present BizFinBench.v2, the first integrated offline and online benchmark built upon authentic user query-response data from both Chinese and U.S. equity markets. It comprises 28,860 questions across eight offline and two online tasks. Experimental results show that GPT-5 achieves a mere 61.5% accuracy, still failing to meet the practical business requirement (84.8%) . Among the evaluated commercial models, DeepSeek-R1 exhibits superior investment efficacy. Error analysis grounded in real financial practice reveals persistent limitations in existing models. By overcoming the constraints of prior benchmarks, BizFinBench.v2 provides a substantiated foundation for advancing LLM deployment in the financial sector. Our data and code are available at [https://github.com/](https://github.com/)[ ](https://github.com/)HiThink-Research/BizFinBench .v2 .  \n1. Introduction  \nLarge Language Models (LLMs) have developed rapidly in recent years, and their application boundaries in the financial sector continue to expand, making them an important technological direction for promoting the intelligent upgrade of financial services (Chen et al., 2024 ; Liu et al., 2025a ; Zhang et al., 2023b ; Lu et al., 2024) . However, as the  \n*Equal contribution 1HiThink Research 2 Shanghai University of Finance and Economics. Correspondence to: Rongjunchen Zhang \u003C[zhangrongjunchen@myhexin.com](zhangrongjunchen@myhexin.com) >, Liwen Zhang \u003C[zhang.liwen@shufe.edu.cn](zhang.liwen@shufe.edu.cn) >.  \nProceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \nFigure 1. BizFinBench.v2 comprises eight foundational tasks and two online tasks distributed across four major scenarios. The topright corner displays a real-time screenshot of the Portfolio Asset Allocation task.  \nscope of LLMs’ application in the financial sector becomes increasingly broad, the discrepancy between model performance under existing evaluation paradigms and their actual performance in real financial business scenarios has become increasingly prominent, directly leading to a conflict between real financial business needs and the limitations of various evaluation benchmarks.  \nThe core characteristics of financial scenarios lie in authenticity and online capability, yet the vast majority of current financial benchmarks are undermined by two fundamental limitations:  \n(1) Disconnected from Real-world Business Scenarios: Most existing studies rely on simulated samples or general datasets that are disconnected from actual business operations. These datasets generally exhibit low difficulty and are decoupled from the requirements of real-world financial business, as they lack the core logic of business workflows  \nFigure 2. We have ranked the performance of the LLMs participating in the evaluation under the zero-shot setting, and these results reflect their authentic practical business capabilities.  \nand fail to reflect the practical challenges inherent in the financial domain. (Zhu et al., 2024 ; Jiang et al., 2025) .  \n(2) Focusing on Purely Static Offline Tasks: Existing benchmarks universally focus on purely offline static tasks (Xie et al., 2024), lacking coverage of online tasks in financial scenarios. This prevents the support for evaluating model performance in online business scenarios such as real-time market data push and dynamic ","cbCaitNWkWpMDVPz","https://ap.wps.com/l/cbCaitNWkWpMDVPz","pdf",6564660,1,35,"English","en",105,"# Introduction\n## Motivation and limitations of existing benchmarks\n## BizFinBench.v2 design and scenarios\n## Benchmark structure and tasks\n## Contributions","[{\"question\":\"What problem does BizFinBench.v2 address in current LLM finance benchmarks?\",\"answer\":\"Existing benchmarks often use simulated or generic data and focus on static offline tasks, which fails to reflect authentic business workflows and online requirements. This leads to a mismatch between benchmark results and real operational performance.\"},{\"question\":\"What data and scope does BizFinBench.v2 use?\",\"answer\":\"BizFinBench.v2 is built from authentic user query-response data from both Chinese and U.S. equity markets. It includes 28,860 QA pairs across offline and online evaluation settings.\"},{\"question\":\"How is BizFinBench.v2 evaluated, and what types of tasks does it include?\",\"answer\":\"It uses a dual-track evaluation covering core business capabilities and online performance. The benchmark contains eight offline tasks and two online tasks, including Stock Price Prediction and Portfolio Asset Allocation.\"}]",1784203988,88,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"bizfinbenchv2-towards-reliable-llms-in-finance-via-real-user-data-and-offlineonline-bilingual-evaluation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/bizfinbenchv2-towards-reliable-llms-in-finance-via-real-user-data-and-offlineonline-bilingual-evaluation/85492/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does BizFinBench.v2 address in current LLM finance benchmarks?","Question",{"text":75,"@type":76},"Existing benchmarks often use simulated or generic data and focus on static offline tasks, which fails to reflect authentic business workflows and online requirements. This leads to a mismatch between benchmark results and real operational performance.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What data and scope does BizFinBench.v2 use?",{"text":80,"@type":76},"BizFinBench.v2 is built from authentic user query-response data from both Chinese and U.S. equity markets. It includes 28,860 QA pairs across offline and online evaluation settings.",{"name":82,"@type":73,"acceptedAnswer":83},"How is BizFinBench.v2 evaluated, and what types of tasks does it include?",{"text":84,"@type":76},"It uses a dual-track evaluation covering core business capabilities and online performance. The benchmark contains eight offline tasks and two online tasks, including Stock Price Prediction and Portfolio Asset Allocation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]