[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85543-en":3,"doc-seo-85543-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85543,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","RL+AHP: A Novel Reinforcement Learning driven AHP for Slice Aware mode selection in D2D enabled Heterogeneous Networks","The mode selection problem in device-to-device (D2D) enabled fifth generation (5G) heterogeneous networks (HetNet) prioritizes four KPIs—data rate, latency, reliability, and jitter—across eMBB, uRLLc, and mMTC slices. Existing methods often neglect slice-specific QoS requirements while balancing LTE-A, NR, and D2D options and limiting handover frequency. A two-level Analytic Hierarchy Process (AHP) combined with reinforcement learning (RL) is proposed to decide among modes by using RL-updated criteria weights from environment feedback. Simulations show KPI gains for all slices and reduced CPU usage versus DRL approaches when criteria count is below 6.","RL+AHP: A Novel Reinforcement Learning driven AHP for Slice Aware mode selection in D2D enabled Heterogeneous Networks  \nSouvik Deb, Sumita Majhi, Shankar K. Ghosh, Member, IEEE, Avirup Das, Member, IEEE, Rajib Mall, Sridevi S  \nand Jacob Augustine  \narXiv :2603 . 1455 1v2 [ cs .NI] 11 Jul 2026  \nAbstract—The mode selection problem in device-to-device communication (D2D) enabled Fifth generation (5G) heterogeneous networks (HetNet) aims prioritizing four key performance indicators (KPIs) namely data rate, latency, reliability and jitter across three slices: enhanced mobile broadband (eMBB), ultra reliable low latency (uRLLc) and massive machine type communications (mMTC). Such priority assignment must be traded off among three access technologies, i.e., Long Term Evolution advanced (LTE-A), New Radio (NR) and D2D, while minimizing handover frequency. In existing mode selection approaches for HetNet, slice specific quality of service (QoS) requirements are largely ignored. In this work, a novel mode selection algorithm is proposed by combining a two level Analytic Hierarchy Process (AHP) with a Reinforcement Learning (RL) method. While the two level AHP facilitates decision making based on multiple criteria (i.e., KPIs) and options (i.e., LTE-A, NR, D2D mode), the RL approach computes the weights of each criteria based on the feedback from the environment. Simulation results show that our proposed algorithm outperforms related works in terms of the major KPIs for all three slices. For eMBB applications, our approach increases throughput by 33%; for uRLLc applications, our approach significantly decreases latency and BER (27% and 10% respectively) and for mMTc applications, our approach significantly decreases latency (44%). Moreover, it has been shown that the proposed RL+AHP approach outperforms the existing DRL based approaches in terms of CPU usage when the number of criteria is reasonably low ( \u003C 6).  \nIndex Terms—5G, Device-to-device communication, mode selection, slice awareness, Reinforcement Learning, Analytic Hierarchy process.  \nI. INTRODUCTION  \nIn Non-standalone (NSA) deployment of fifth generation (5G) cellular network, Long Term Evolution advanced (LTEA) and New Radio (NR) systems co-exist [1] . Therein, LTE-A  \nSouvik Deb is affiliated with School of Computer Science, University of Petroleum and Energy Studies, Dehradun, India. Email: sou[vik.deb@upes.ac.in](vik.deb@upes.ac.in)  \nSumita Majhi is affiliated with Department of Computer Science and Engineering, Indian Institute of Technology Guwahati, Guwahati, India. Email:{[sumit176101013](sumit176101013}@iitg.ac.in)[}](sumit176101013}@iitg.ac.in)[@iitg.ac.in](sumit176101013}@iitg.ac.in)  \nShankar K. Ghosh and Rajib Mall are affiliated with Department of Computer Science and Engineering, Shiv Nadar Institution of Eminence, Delhi NCR, [India. Emails:](India. Emails: {shankar.it46)[ {](India. Emails: {shankar.it46)[shankar.it46](India. Emails: {shankar.it46), [rajib.mall](rajib.mall}@snu.edu.in)[}](rajib.mall}@snu.edu.in)[@snu.edu.in](rajib.mall}@snu.edu.in).  \nAvirup Das is affiliated with Singapore University of Technology and Design, Singapore. Emails: [avirup1987@gmail.com](avirup1987@gmail.com).  \nSridevi S and Jacob Augustine are affiliated with School of Computer Science and Engineering, Presidency University, Bengaluru, India. Emails:{sridevi.svs2809, [jacob.ku.augustine](jacob.ku.augustine}@gmail.com)[}](jacob.ku.augustine}@gmail.com)[@gmail.com](jacob.ku.augustine}@gmail.com).  \nShankar K. Ghosh and Avirup Das will be acting as corresponding authors.  \nA preliminary version of the manuscript has been archived on 15th March 2026 (DOI: [https://arxiv.org/pdf/2603.14551](https://arxiv.org/pdf/2603.14551)).  \nFig. 1: An example of D2D enabled HetNet: UE 3 is communicating in Mode-1; UE 4 is communicating in Mode-2; UE 5 is communicating in Mode-3; and UE 2 is communicating using Mode-4 .  \nmacro-cells provide ubiquitous coverage, whereas ultra dense deployment","cbCaigEJtKGgTods","https://ap.wps.com/l/cbCaigEJtKGgTods","pdf",2344650,2,1,15,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What KPIs and slices does the mode selection method target in D2D-enabled HetNets?\",\"answer\":\"It targets data rate, latency, reliability, and jitter across three slices: eMBB, uRLLc, and mMTC.\"},{\"question\":\"How does the proposed RL+AHP approach differ from existing mode selection methods?\",\"answer\":\"It combines a two-level AHP for multi-criteria decision-making with reinforcement learning to compute criteria weights from environment feedback, addressing slice-specific QoS requirements.\"},{\"question\":\"What performance improvements are reported for each slice?\",\"answer\":\"For eMBB, throughput increases by 33%. For uRLLc, latency and BER decrease by 27% and 10% respectively, and for mMTC latency decreases by 44%.\"}]",1784204327,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"rlahp-a-novel-reinforcement-learning-driven-ahp-for-slice-aware-mode-selection-in-d2d-enabled-heterogeneous-networks","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/rlahp-a-novel-reinforcement-learning-driven-ahp-for-slice-aware-mode-selection-in-d2d-enabled-heterogeneous-networks/85543/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What KPIs and slices does the mode selection method target in D2D-enabled HetNets?","Question",{"text":75,"@type":76},"It targets data rate, latency, reliability, and jitter across three slices: eMBB, uRLLc, and mMTC.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed RL+AHP approach differ from existing mode selection methods?",{"text":80,"@type":76},"It combines a two-level AHP for multi-criteria decision-making with reinforcement learning to compute criteria weights from environment feedback, addressing slice-specific QoS requirements.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance improvements are reported for each slice?",{"text":84,"@type":76},"For eMBB, throughput increases by 33%. For uRLLc, latency and BER decrease by 27% and 10% respectively, and for mMTC latency decreases by 44%.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]