[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83416-en":3,"doc-seo-83416-105":30,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83416,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","MPFlow Learning Budgeted Max-Flow Optimization on the Lightning Network with Deep Graph Reinforcement Learning","MPFlow addresses liquidity placement in the Bitcoin Lightning Network under a fixed budget, aiming to maximize routing capacity by choosing which channels a node opens. The work formulates the problem as budget-constrained combinatorial optimization on graphs: selecting k edge additions to maximize the s–t max-flow on the observed capacity graph. A lightweight graph reinforcement learning agent combines a message-passing policy network with PPO and action masking, trained using a hub-exclusion curriculum to learn capacity-aware placement rather than hub attachment. Experiments on real Lightning Network snapshots show consistent improvements over strong heuristic baselines across seeds and unseen graphs, and the approach is deployed for peer recommendations with thousands of executed channel-open decisions totaling 267.3 BTC.","arXiv :2607 .08703v 1 [ cs .LG] 9 Jul 2026  \nMPFlow: Learning Budgeted Max-Flow Optimization on the Lightning Network with Deep Graph Reinforcement Learning  \nHarrison Rush  \nAmboss Technologies, Green Cove Springs, USA  \nVincent Davis  \nAmboss Technologies, Green Cove Springs, USA  \nSimone Antonelli  \nCISPA Helmholtz Center for Information Security  \nVikash Singh  \nStillmark, San Francisco, USA  \nJesse Shrader  \nAmboss Technologies, Green Cove Springs, USA  \nEmanuele Rossi  \nAmboss Technologies & Sapienza University of Rome, Barcelona, Spain  \n[harrison@amboss.tech](harrison@amboss.tech)  \n[vincentmdavis@protonmail. com](vincentmdavis@protonmail. com)  \n[simone. antonelli@cispa. de](simone. antonelli@cispa. de)  \n[vikash@stillmark. com](vikash@stillmark. com)  \n[j@amboss.tech](j@amboss.tech)  \n[emanuele. rossi1909@gmail. com](emanuele. rossi1909@gmail. com)  \nAbstract  \nB  \nWe address liquidity placement in the Bitcoin Lightning Network (LN): given a fixed budget, which channels should a node open to maximize its routing capacity? We cast this asa budget-constrained combinatorial optimization problem on graphs, selecting k edge additions that maximize s–t max-flow, a theory-grounded measure of routing capacity, and solve it with graph reinforcement learning. Our lightweight agent combines a message-passing policy network with proximal policy optimization (PPO) and action masking, and is trained under a hub-exclusion curriculum: the network’s top hubs are removed from training subgraphs, forcing the policy to learn capacity-aware placement rather than hub attachment. In extensive experiments on real Lightning Network snapshots, our method consistently outperforms strong heuristic baselines on the max-flow objective across multiple seeds and unseen graphs. The agent has been deployed in production for peer recommendations, executing 4,640 channel-open decisions that cumulatively allocate 267.3 BTC (over $16 million) across 30 managed nodes.  \nLightning Network allocation as a Markov decision process  \na new node opens k = 5 channels: which peers maximize its routing capacity? Inside the box, one step:  \nFigure 1: Liquidity placement as learned sequential graph construction. Left: MPFlow maps each state st to a peer at (budget k=5 × 0.20 BTC); reward: marginal max-flow rt = Ft − Ft−1 . Right: two max-aggregation MPNN layers, a masked-softmax actor, and a max-pooled critic, trained with PPO.  \n1 Introduction  \nPayment-channel networks such as the Bitcoin Lightning Network (LN) enable fast, low-cost payments by moving transactions off-chain. Their performance, however, hinges on where liquidity is placed: poor placement creates bottlenecks that throttle routing capacity, while targeted placement can unlock throughput. For LN operators the question is simple to state and hard to solve: which peers should I connect to, and how should I allocate scarce liquidity to participate effectively? Existing practice leans on static graph heuristics (e.g., degree or betweenness) that ignore directed balances, budgeted interventions, and interactions among multiple allocations (Newman, 2010; Freeman, 1977) .  \nThe LN offers little external observability and no supervised traces; realistic traffic simulation demands strong assumptions about demand matrices, retry logic, and hidden balances that can dominate conclusions (see, e.g., LN measurement/topology studies) (Seres et al., 2020; Rohrer et al., 2019) . We therefore frame liquidity placement as budgeted combinatorial optimization on a graph: select k edge additions that maximizes–t max-flow on the observed capacity graph (Ford & Fulkerson, 1956; Ahuja et al., 1993), a theory-grounded objective that directly measures routing capacity and upper-bounds deliverable payment volume. In this we follow the learning-to-optimize-on-graphs line initiated by Dai et al. (2018), which trains graph-embedding RL policies to construct combinatorial solutions incrementally; in our case, the task is budget-cons","cbCaipwwvk4kBPSo","https://ap.wps.com/l/cbCaipwwvk4kBPSo","pdf",1264136,2,1,24,"English","en",105,"# Abstract\n# Introduction\n# Background\n## Bitcoin & Lightning Network\n# Lightning Network allocation as a Markov decision process","[{\"question\":\"What problem does MPFlow solve in the Lightning Network?\",\"answer\":\"MPFlow determines which peers (channels) a node should open under a fixed liquidity budget to maximize routing capacity.\"},{\"question\":\"How is routing capacity measured in the paper’s optimization objective?\",\"answer\":\"The objective selects k edge additions to maximize the s–t max-flow on the observed capacity graph, treating it as a theory-grounded routing-capacity measure.\"},{\"question\":\"What learning approach does MPFlow use to construct the graph decisions?\",\"answer\":\"MPFlow uses a graph reinforcement learning agent with a message-passing policy network, PPO updates, and feasibility-aware action masking.\"},{\"question\":\"Why does the training use a hub-exclusion curriculum?\",\"answer\":\"Top hubs are removed from training subgraphs so the policy learns capacity-aware placement rather than relying on attaching to prominent hubs, improving generalization.\"}]",1784187438,60,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":28},"mpflow-learning-budgeted-max-flow-optimization-on-the-lightning-network-with-deep-graph-reinforcement-learning","",{"@graph":36,"@context":89},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/mpflow-learning-budgeted-max-flow-optimization-on-the-lightning-network-with-deep-graph-reinforcement-learning/83416/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does MPFlow solve in the Lightning Network?","Question",{"text":75,"@type":76},"MPFlow determines which peers (channels) a node should open under a fixed liquidity budget to maximize routing capacity.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is routing capacity measured in the paper’s optimization objective?",{"text":80,"@type":76},"The objective selects k edge additions to maximize the s–t max-flow on the observed capacity graph, treating it as a theory-grounded routing-capacity measure.",{"name":82,"@type":73,"acceptedAnswer":83},"What learning approach does MPFlow use to construct the graph decisions?",{"text":84,"@type":76},"MPFlow uses a graph reinforcement learning agent with a message-passing policy network, PPO updates, and feasibility-aware action masking.",{"name":86,"@type":73,"acceptedAnswer":87},"Why does the training use a hub-exclusion curriculum?",{"text":88,"@type":76},"Top hubs are removed from training subgraphs so the policy learns capacity-aware placement rather than relying on attaching to prominent hubs, improving generalization.","https://schema.org",{"og:url":51,"og:type":91,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":93,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":96},[97,101,105,109,113,118,123,126,131,134,138],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":29,"slug":112},5,"Comic","comic",{"id":114,"doc_module":4,"doc_module_name":46,"category_name":115,"show_sort_weight":116,"slug":117},6,"Technology",50,"technology",{"id":119,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":121,"slug":122},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":124,"slug":125},30,"research-report",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":129,"slug":130},9,"Religion & Spirituality",20,"religion-spirituality",{"id":129,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":129,"slug":133},"World Cup","world-cup",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":135,"slug":137},10,"Lifestyle","lifestyle",{"id":139,"doc_module":4,"doc_module_name":46,"category_name":140,"show_sort_weight":110,"slug":141},19,"General","general"]