[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85267-en":3,"doc-seo-85267-105":29,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85267,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Reinforcement Learning for Execution under Dynamic Fees in a Closed-Loop DEX Simulator","Trader-facing dynamic fees are increasingly proposed for automated market makers, yet historical data cannot reveal how order flow responds, because fee schedules may not vary, trader types are latent, and replayed tapes do not form sequential decision environments. This work builds a minimal closed-loop simulator where the missing signal exists: two constant-product pools with an equilibrium-inspired dynamic-fee rule, fee-sensitive noise flow, and closed-form CEX–AMM arbitrage. Model-free DQN control reduces implementation shortfall under dynamic fees across ordering regimes, while gains vanish under constant fees, yielding model-conditioned counterfactual evidence for execution control rather than historical-trader identification.","arXiv :2607 . 10960v 1 [ cs .LG] 12 Jul 2026  \nReinforcement Learning for Execution under  \nDynamic Fees in a Closed-Loop DEX Simulator  \nWen-Ting Wang  \nJuly 2026  \nEmail: [egpivo@gmail.com](egpivo@gmail.com)  \nAbstract  \nTrader-facing dynamic fees are increasingly proposed for automated market makers (AMMs), but historical data do not identify how order flow would respond: trader-facing fees do not vary, trader types are latent, and a replayed tape is not a sequential decision environment. We therefore construct a minimal closed-loop simulator in which the missing signal exists by construction: two constant-product pools repriced by an equilibrium-inspired dynamic-fee rule, fee-sensitive noise flow, and closed-form CEX–AMM arbitrage. Equilibrium is used as a closure principle, not as an object the trader learns. Against a tuned benchmark ladder of schedule, planning, lookahead, and tabular policies, a small DQN is the only evaluated valid policy whose paired improvement over tuned one-step routing excludes zero. On a reserved final block of 1,000 seeds with completion forced to 1.0 for every policy, it reduces implementation shortfall under every tested intra-step ordering, by 13 .3bps of order notional under the pre-specified agent-last ordering, and the edge is concentrated in, and learned from, dynamic-fee environments: under constant fees the paired difference is indistinguishable from zero. The result is model-conditioned counterfactual evidence about execution control in AMMs, not evidence about historical traders, equilibrium play, or deployable profit.  \nKeywords: automated market makers; decentralized exchanges; dynamic fees; optimal execution; reinforcement learning; market simulation.  \n1 Introduction  \nIn decentralized exchange research, it is of interest to know whether trader-facing dynamic fees protect liquidity providers, and, dually, whether execution agents can defend themselves against fee rules that reprice as they trade. A commonly used approach is empirical: estimate responses from historical swaps and liquidity events. However, historical identification stops at specific margins. In our companion event study of the Uniswap protocol-fee switch [25], the LP-side response to take-rate changes is identified by the design; the trader-facing question isnot, because trader-facing fees do not vary, trader types are latent, and the routing decision seta trader faced at each block is unobserved. Questions of the form “what would order flow do if fees moved against it” are counterfactuals the record cannot answer.  \nA second approach applies reinforcement learning to trade execution [21, 17 , 22], in our case by training an agent on the historical tape. In preliminary experiments we ran that program across four data regimes and closed it with a negative result: the historical tape does not identify an action-dependent transition kernel for counterfactual policies. A replayed agent can condition on history, but its actions cannot change subsequent states, and each apparent pocket of adaptive  \nheadroom was traced to a measurable artifact by pre-specified audits. The binding constraint was not sample size but the information content of the signal.  \nTo address this issue, we stop asking the tape for a signal it does not contain and instead construct a minimal market in which the signal exists by construction: a closed-loop simulator in which an execution trader’s actions move pool inventory, which moves quotes, which moves a dynamic fee rule, noise-flow routing, and an arbitrageur’s response. Two design commitments separate this from a generic market game. First, the fee rule is equilibrium-inspired closure: alinearization of the approximate Nash fee structure derived for competing constant-function market makers by Baggiani, Herdegen, and S´anchez-Betancourt [4, 5], whose role is to make the environment defend itself, not to certify equilibrium play. Second, evaluation discipline is inherited from that audit proto","cbCaisQNHYUn0Qoq","https://ap.wps.com/l/cbCaisQNHYUn0Qoq","pdf",2092233,1,19,"English","en",105,"# 1 Introduction\n## Motivation and limitations of historical identification\n## Reinforcement learning on historical tapes\n## Closed-loop simulator design and evaluation discipline\n## Research question and key findings\n## Related work and positioning\n# 2 Evaluation method\n# 3 Dynamic-fee DEX simulation study\n# 4 Scope and limitations","[{\"question\":\"Why can’t historical data identify how order flow responds to trader-facing dynamic fees?\",\"answer\":\"Because trader-facing fee schedules may not vary in the record, trader types are unobserved, and the routing and action-dependent transitions needed for counterfactuals are not observable in a replayed tape.\"},{\"question\":\"What makes the proposed simulator a closed-loop environment for execution?\",\"answer\":\"The trader’s actions change pool inventory, which updates quotes and then a dynamic-fee rule; the resulting fee-sensitive routing flow and a responding arbitrageur close the loop with sequential state changes.\"},{\"question\":\"What is the main experimental result about a DQN policy versus tuned execution heuristics?\",\"answer\":\"A small DQN is the only evaluated policy that shows statistically non-zero paired improvement over tuned one-step routing, reducing implementation shortfall under dynamic-fee environments, with the advantage concentrated in those settings.\"},{\"question\":\"Does the improvement persist under constant fees?\",\"answer\":\"No. Under constant fees, the paired difference between the learner and the heuristic is indistinguishable from zero.\"}]",1784202167,48,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":27},"reinforcement-learning-for-execution-under-dynamic-fees-in-a-closed-loop-dex-simulator","",{"@graph":35,"@context":89},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/reinforcement-learning-for-execution-under-dynamic-fees-in-a-closed-loop-dex-simulator/85267/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"Why can’t historical data identify how order flow responds to trader-facing dynamic fees?","Question",{"text":75,"@type":76},"Because trader-facing fee schedules may not vary in the record, trader types are unobserved, and the routing and action-dependent transitions needed for counterfactuals are not observable in a replayed tape.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What makes the proposed simulator a closed-loop environment for execution?",{"text":80,"@type":76},"The trader’s actions change pool inventory, which updates quotes and then a dynamic-fee rule; the resulting fee-sensitive routing flow and a responding arbitrageur close the loop with sequential state changes.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the main experimental result about a DQN policy versus tuned execution heuristics?",{"text":84,"@type":76},"A small DQN is the only evaluated policy that shows statistically non-zero paired improvement over tuned one-step routing, reducing implementation shortfall under dynamic-fee environments, with the advantage concentrated in those settings.",{"name":86,"@type":73,"acceptedAnswer":87},"Does the improvement persist under constant fees?",{"text":88,"@type":76},"No. Under constant fees, the paired difference between the learner and the heuristic is indistinguishable from zero.","https://schema.org",{"og:url":51,"og:type":91,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":93,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":96},[97,101,105,109,114,119,124,127,132,135,139],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},5,"Comic",60,"comic",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},6,"Technology",50,"technology",{"id":120,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":122,"slug":123},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":125,"slug":126},30,"research-report",{"id":128,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":130,"slug":131},9,"Religion & Spirituality",20,"religion-spirituality",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":130,"slug":134},"World Cup","world-cup",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":136,"slug":138},10,"Lifestyle","lifestyle",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":140,"show_sort_weight":110,"slug":141},"General","general"]