[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85385-en":3,"doc-seo-85385-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85385,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Learning to Schedule in Parallel-Server Queues with Stochastic Bilinear Rewards","Scheduling in multiclass parallel-server queues is studied when assignment rewards are stochastic and unknown in expectation, while jobs accumulate holding costs before completion. Assignment rewards are observed but their means are modeled by a bilinear function of job and server features. The goal is to minimize regret by maximizing cumulative assignment reward across a horizon, subject to a bounded holding cost that preserves queue stability and throughput. A weighted proportional-fair scheduling rule with marginal reward costs is paired with a bilinear-reward bandit to control the regret–queue length tradeoff and provide uniform expected queue and cost bounds.","Learning to Schedule in Parallel-Server Queues with Stochastic Bilinear Rewards  \nJung-hun Kim and Milan Vojnovi  \narXiv :2112 .06362v 5 [ cs .LG] 12 Jul 2026  \nAbstract—We consider the problem of scheduling in multiclass, parallel-server queuing systems with uncertain rewards from job-server assignments. In this scenario, jobs incur holding costs while awaiting completion, and job-server assignments yield observable stochastic rewards with unknown mean values. The mean rewards for job-server assignments are assumed to follow abilinear model with respect to features that characterize jobs and servers. Our objective is to minimize regret by maximizing the cumulative reward of job-server assignments over a time horizon, while keeping the total job holding cost bounded to ensure the stability of the queueing system. This problem is motivated by applications requiring resource allocation in network systems.  \nA central challenge is to control the tradeoff between reward maximization and fair allocation for the stability of the underlying queuing system (i.e., maximizing network throughput). To address this challenge, we propose a scheduling algorithm based on a weighted proportional fair criteria augmented with marginal costs for reward maximization, incorporating a bandit algorithm tailored for bilinear rewards. Our algorithm admitsa regret–queue length tradeoff. For any fixed control parameter V > 0, it ensures a uniform expected queue length and timeaverage holding-cost bounds. For a target horizon T, choosing VT = Θ( √IT) at initialization yields eO(( √I + d2 )√T + 1/δ) regret. Under this regret-optimized tuning, the corresponding expected queue length and time-average holding-cost bounds remain uniform over the execution time and scales as O( √IT +1/δ)  \nand O ( √IT/δ), respectively.  \nIndex Terms—Resource Allocation, Scheduling Jobs, Queuing, Reward Maximization, Stability, Online Learning.  \nI. INTRODUCTION  \nIn this work, we address the problem of scheduling jobs in multi-class, parallel-server queuing systems—such as those found in data centers, edge computing infrastructures, and communication networks. In such systems, both flow types (jobs) and processing units (servers) can have heterogeneous characteristics, necessitating differentiated services. Assigning a job to a server or network function yields an observable stochastic reward—e.g., processing rate, or an applicationspecific quality of the job output dependent on the server assignment—with an unknown mean value that depends on the compatibility between job and server characteristics. Note that considering rewards of assignments accommodates assignment costs, treating them as negative rewards.  \nSpecifically, we consider the case of noisy rewards, where the rewards for job-server assignments follow a bilinear model based on the features characterizing jobs and servers. This  \nJung-hun Kim is with CREST, ENSAE Paris, France (email: [junghun.kim@ensae.fr](junghun.kim@ensae.fr))  \nMilan Vojnovi is with London School of Economics, United Kingdom ([email: m.vojnovic@lse.ac.uk](email: m.vojnovic@lse.ac.uk))  \nreward model can capture complex interactions between job and server characteristics and can be leveraged to make effective scheduling decisions in uncertain environments, as demonstrated in our work. The scheduler’s objective is to maximize the expected cumulative reward over a prescribed horizon while controlling the expected job holding cost.  \nThis problem arises in a wide range of networking systems where resource allocation decisions must be made under uncertainty. For example, in data centers and distributed cloud computing systems [1], [2], computational jobs composed of multiple tasks must be assigned to servers with heterogeneous processing capabilities and varying data locality preferences. A common goal is to maximize system throughput when processing dynamic workloads–a challenging task further exacerbated by uncertainty or lack of knowledge about cer","cbCaihDjAle67KfF","https://ap.wps.com/l/cbCaihDjAle67KfF","pdf",3451533,5,1,33,"English","en",105,"# Introduction\n## Problem Setting and Motivation\n## Related Work\n## Contributions","[{\"question\":\"What problem does the document address?\",\"answer\":\"It studies scheduling jobs in multiclass parallel-server queuing systems where rewards from job-server assignments are stochastic and have unknown mean values modeled from job and server features.\"},{\"question\":\"How are the rewards modeled?\",\"answer\":\"The reward means follow a bilinear model with respect to features characterizing jobs and servers, capturing interactions between job and server characteristics.\"},{\"question\":\"What is the main tradeoff and how is it handled?\",\"answer\":\"The key tradeoff is between maximizing reward and ensuring stability via bounded queue length/holding cost. The proposed weighted proportional-fair scheduling rule with marginal reward costs, combined with a bilinear-reward bandit, provides a regret–queue length tradeoff with uniform bounds under proper tuning.\"}]",1784203056,83,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"learning-to-schedule-in-parallel-server-queues-with-stochastic-bilinear-rewards","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/learning-to-schedule-in-parallel-server-queues-with-stochastic-bilinear-rewards/85385/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the document address?","Question",{"text":76,"@type":77},"It studies scheduling jobs in multiclass parallel-server queuing systems where rewards from job-server assignments are stochastic and have unknown mean values modeled from job and server features.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How are the rewards modeled?",{"text":81,"@type":77},"The reward means follow a bilinear model with respect to features characterizing jobs and servers, capturing interactions between job and server characteristics.",{"name":83,"@type":74,"acceptedAnswer":84},"What is the main tradeoff and how is it handled?",{"text":85,"@type":77},"The key tradeoff is between maximizing reward and ensuring stability via bounded queue length/holding cost. The proposed weighted proportional-fair scheduling rule with marginal reward costs, combined with a bilinear-reward bandit, provides a regret–queue length tradeoff with uniform bounds under proper tuning.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]