[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128729-en":3,"doc-seo-128729-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128729,1099523882182,"Eliana","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",8,"Research & Report","Data-driven simulations and policy gradients for limit order books","Central limit order books (LOBs) are a widely used execution method, yet classical parametric modelling can produce bias and a mismatch with observed market dynamics. This thesis develops a data-driven paradigm that fits a generative, non-parametric model to historical LOB data and uses reinforcement learning for optimization. It introduces K-nearest neighbor resampling to estimate policy performance from episodic counterfactual data, prove statistical consistency, and avoid explicit parametric assumptions. It further uses the method for simulation, calibration of trading strategies, and liquidation order sizing. Finally, it applies policy gradient approaches—enhanced via variance reduction and path-wise estimators—and analyzes linear convergence in continuous time.","Data-driven simulations and policy gradients for  \nlimit order books  \nMichael Giegrich St Anne’s College University of Oxford  \nThesis submitted for the degree of DPhil in Mathematics  \nMichaelmas 2024  \n1  \nAcknowledgements  \nFirst and foremost, I would like to express my sincere and deep gratitude to my supervisor Professor Christoph Reisinger for his continued support and his genuine interest in my research and my well-being. I would also like thank my industry partner and co-author Dr. Roel Oomen for his insightful comments and suggestions, which enhanced the applicability of my research. I am grateful to my co-author Dr. Yufei Zhang for the profound and motivating discussions on stochastic control and other topics. Furthermore, I would like to acknowledge the helpful feedback I received on my research from Professor Álvaro Cartea, Professor Sam Cohen, Professor Mihai Cucuringu, Professor Ben Hambly, Dr. Leandro Sanchez Betancourt, Professor Renyuan Xu, and other researchers I had the pleasure to interact with. I am also grateful to my colleagues and friends at the Math Department, who shared this journey with me, in particular, my (unoﬃcial) oﬃce mates Filippo de Angelis, Philipp Jettkant, and Aldaïr Petronilia. In my previous studies, I was fortunate enough to learn and receive mentorship from great mathematicians, in particular, Professor Antje Berndt, Professor Lukas Gonon, Tobias Leopold, Professor Enno Mammen, and Professor Josef Teichmann. I acknowledge the funding from Deutsche Bank and the Mathematics of Random Systems CDT.  \nI am deeply grateful to my family, especially to my parents, Gisela and Jürgen, for helping me ﬁnd my way. Furthermore, I thank my friends, in particular Carlo, Clemens, Martin, and Wasilios, for enriching my life. Finally, I could not have undertaken this journey without my wife, María, and her constant support, love, understanding, and encouragement.  \nAbstract  \nOver the last decades, central limit order books (LOBs) became a main method for asset execution in exchanges around the world for a wide variety of asset classes (e.g. [39, 85, 213]) . Their practical relevance and inherent complex strategic interactions made LOBs a primary objective of study for academia and industry. The classical approach to modelling limit order books is based on identifying the most important features of the problem, encoding them in a parametric model and calibrating this model to market data. However, parametric modelling can introduce biases leading to a mismatch between actual market data and model output (e.g. [48]) . Recent success in generative artiﬁcial intelligence (e.g. [75, 100]) and reinforcement learning (e.g. [93, 134, 194]) combined with the wide availability of LOB data motivates the usage of a new paradigm for decision making in LOBs, namely, ﬁtting a generative, non-parametric model to the historical data and using reinforcement learning for optimization. In this thesis, we explore diﬀerent aspects of this new paradigm.  \nFirstly, we propose a novel generative model, K-nearest neighbor resampling, for estimating the performance of a policy from historical data containing realized episodes of a decision process generated under a diﬀerent policy. We provide statistical consistency results under weak conditions by generalizing a well-known result in non-parametric statistics on local averaging to include episodic data and counterfactual estimation. Compared to similar methods, our algorithm does not require optimization, can be eﬃciently implemented via treebased nearest neighbor search and parallelization, and does not explicitly assume a parametric model for the environment’s dynamics. Numerical experiments demonstrate the eﬀectiveness of the algorithm compared to existing baselines in a variety of stochastic control settings.  \nSecondly, we show how K-nearest neighbor (K-NN) resampling can be applied to simulate LOB markets and how it can be used to evaluate and calibrate trading strategies","cbCaigpsemyqebKf","https://ap.wps.com/l/cbCaigpsemyqebKf","pdf",36484309,4,1,153,"English","en",105,"# Introduction\n## Limit order books\n## Stochastic control background","[{\"question\":\"What limitation of classical limit order book modelling motivates this thesis?\",\"answer\":\"Classical parametric modelling can introduce biases, causing discrepancies between model output and actual market data.\"},{\"question\":\"How does K-nearest neighbor resampling estimate policy performance?\",\"answer\":\"It uses historical data with realized episodes generated under a different policy, combining local averaging ideas with episodic and counterfactual estimation under weak conditions.\"},{\"question\":\"How are policy gradient methods applied in the thesis?\",\"answer\":\"Policy gradients are used for optimal execution, both in a parametric LOB model and in a generative adversarial (GAN) LOB model, including modifications such as variance reduction and a path-wise gradient estimator.\"}]","Data-driven simulations and policy gradients for limit order books | PDF",1786002909,386,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"data-driven-simulations-and-policy-gradients-for-limit-order-books","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/data-driven-simulations-and-policy-gradients-for-limit-order-books/128729/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What limitation of classical limit order book modelling motivates this thesis?","Question",{"text":76,"@type":77},"Classical parametric modelling can introduce biases, causing discrepancies between model output and actual market data.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does K-nearest neighbor resampling estimate policy performance?",{"text":81,"@type":77},"It uses historical data with realized episodes generated under a different policy, combining local averaging ideas with episodic and counterfactual estimation under weak conditions.",{"name":83,"@type":74,"acceptedAnswer":84},"How are policy gradient methods applied in the thesis?",{"text":85,"@type":77},"Policy gradients are used for optimal execution, both in a parametric LOB model and in a generative adversarial (GAN) LOB model, including modifications such as variance reduction and a path-wise gradient estimator.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]