<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Sig. Proc.</journal-id>
<journal-title>Frontiers in Signal Processing</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Sig. Proc.</abbrev-journal-title>
<issn pub-type="epub">2673-8198</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">864392</article-id>
<article-id pub-id-type="doi">10.3389/frsip.2022.864392</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Signal Processing</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A Tutorial on Bandit Learning and Its Applications in 5G Mobile Edge Computing (<italic>Invited Paper</italic>)</article-title>
<alt-title alt-title-type="left-running-head">Liu et al.</alt-title>
<alt-title alt-title-type="right-running-head">Bandit Learning Mobile Edge Computing</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Liu</surname>
<given-names>Sige</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1202754/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Cheng</surname>
<given-names>Peng</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1147618/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Chen</surname>
<given-names>Zhuo</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Vucetic</surname>
<given-names>Branka</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Li</surname>
<given-names>Yonghui</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/994059/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>School of Electrical and Information Engineering</institution>, <institution>The University of Sydney</institution>, <addr-line>Darlington</addr-line>, <addr-line>NSW</addr-line>, <country>Australia</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Department of Computer Science and Information Technology</institution>, <institution>La Trobe University</institution>, <addr-line>Melbourne</addr-line>, <addr-line>VIC</addr-line>, <country>Australia</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>CSIRO DATA61</institution>, <addr-line>Eveleigh</addr-line>, <addr-line>NSW</addr-line>, <country>Australia</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/990504/overview">Boya Di</ext-link>, Imperial College London, United Kingdom</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/836047/overview">Zehui Xiong</ext-link>, Nanyang Technological University, Singapore</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1168510/overview">Youjia Chen</ext-link>, Fuzhou University, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Peng Cheng, <email>p.cheng@latrobe.edu.au</email>, <email>peng.cheng@sydney.edu.au</email>; Yonghui Li, <email>yonghui.li@sydney.edu.au</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Signal Processing for Communications, a section of the journal Frontiers in Signal Processing</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>02</day>
<month>05</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>2</volume>
<elocation-id>864392</elocation-id>
<history>
<date date-type="received">
<day>28</day>
<month>01</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>28</day>
<month>03</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Liu, Cheng, Chen, Vucetic and Li.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Liu, Cheng, Chen, Vucetic and Li</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Due to the rapid development of 5G and Internet-of-Things (IoT), various emerging applications have been catalyzed, ranging from face recognition, virtual reality to autonomous driving, demanding ubiquitous computation services beyond the capacity of mobile users (MUs). Mobile cloud computing (MCC) enables MUs to offload their tasks to the remote central cloud with substantial computation and storage, at the expense of long propagation latency. To solve the latency issue, mobile edge computing (MEC) pushes its servers to the edge of the network much closer to the MUs. It jointly considers the communication and computation to optimize network performance by satisfying quality-of-service (QoS) and quality-of-experience (QoE) requirements. However, MEC usually faces a complex combinatorial optimization problem with the complexity of exponential scale. Moreover, many important parameters might be unknown <italic>a-priori</italic> due to the dynamic nature of the offloading environment and network topology. In this paper, to deal with the above issues, we introduce bandit learning (BL), which enables each agent (MU/server) to make a sequential selection from a set of arms (servers/MUs) and then receive some numerical rewards. BL brings extra benefits to the joint consideration of offloading decision and resource allocation in MEC, including the matched mechanism, situation awareness through learning, and adaptability. We present a brief tutorial on BL of different variations, covering the mathematical formulations and corresponding solutions. Furthermore, we provide several applications of BL in MEC, including system models, problem formulations, proposed algorithms and simulation results. At last, we introduce several challenges and directions in the future research of BL in 5G MEC.</p>
</abstract>
<kwd-group>
<kwd>5G</kwd>
<kwd>mobile edge computing</kwd>
<kwd>bandit learning</kwd>
<kwd>task offloading</kwd>
<kwd>resource allocation</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>The ever-increasing deployment of the fifth-generation (5G) communication and the Internet-of-Things (IoT) has created a large number of emerging applications ranging from face recognition, virtual reality to autonomous driving (<xref ref-type="bibr" rid="B31">Shi et al., 2016</xref>; <xref ref-type="bibr" rid="B36">Teng et al., 2019</xref>), and generated enormous volumes of data for transmission, storage, and execution. However, these tremendous computational requirements are usually beyond the capacity of mobile users (MUs), making it impossible to complete tasks in a prompt manner. Mobile cloud computing (MCC) (<xref ref-type="bibr" rid="B16">Khan et al., 2014</xref>) is a promising computational paradigm to relieve this situation by enabling MUs to offload their applications to the remote central cloud with strong computation and storage infrastructure. However, the inherent issue of MCC is a long communication distance between MUs and the remote cloud center, resulting in a long propagation and network delay.</p>
<p>Mobile edge computing (MEC) (<xref ref-type="bibr" rid="B37">Wang et al., 2017</xref>; <xref ref-type="bibr" rid="B43">Zhang et al., 2020</xref>; <xref ref-type="bibr" rid="B1">Asheralieva et al., 2021</xref>) has been envisioned as a key enabler to deal with the latency issue in MCC. MEC pushes the computation and storage resources to the network edge much closer to the local devices, benefiting from a low propagation latency and privacy/security enhancement (<xref ref-type="bibr" rid="B28">Mao et al., 2017</xref>). In MEC, joint consideration of communication and computation plays a pivotal role in network performance optimization to satisfy quality-of-service (QoS) and quality-of-experience (QoE) requirements (<xref ref-type="bibr" rid="B42">Yang et al., 2020</xref>; <xref ref-type="bibr" rid="B20">Lim et al., 2021</xref>).</p>
<p>Despite its potential benefits, MEC also suffers from several challenging issues. An MU needs to offload computation tasks to MEC servers in an opportunistic manner subject to the available edge computation and communication resources shared by a large number of MUs. Consequently the offloading decision and communication/computation resource allocation should be jointly optimized to maximize the network performance, typically leading to a complex combinatorial optimization problem with complexity of exponential scale. This is further exacerbated by the dynamic nature of the offloading environment and network topology, where many important parameters (e.g., channel state information and servers&#x2019; computation workload) in the formulated problem are either impossible or difficult to obtain <italic>a-priori</italic>.</p>
<p>Bandit learning (BL) (<xref ref-type="bibr" rid="B15">Gittins et al., 2011</xref>; <xref ref-type="bibr" rid="B35">Sutton and Barto, 2018</xref>), a typical online learning approach, offers a promising solution to deal with the aforementioned issues. It enables each agent (MU/server) to make a sequential selection from a set of arms (servers/MUs) in order to receive some numerical rewards available to the agent after pulling the arm. BL aims to strike a tradeoff between exploitation (exploit the learned knowledge and select the empirically optimal arm) and exploration (explore other arms than the optimal one to get more reward information). Consequently, the arms can be iteratively learned, and selection decisions will be improved progressively. The major advantages that BL brings to the joint consideration of offloading decision and network resource allocation can be summarized as follows.<list list-type="simple">
<list-item>
<p>&#x2022; Matched mechanism: The inherent idea behind MEC is to design policies to make a better selection for MUs or servers. This clearly coincides with the design purpose of BL. This match in mechanism provides selection policies to obtain better performance such as lower latency, lower energy consumption, and higher task completion ratio.</p>
</list-item>
<list-item>
<p>&#x2022; Situation awareness through learning: In the arm selection process, BL is able to learn the corresponding parameters of the offloading environment.</p>
</list-item>
<list-item>
<p>&#x2022; Adaptability: The structure of BL can be readily modified to accommodate a variety of characteristics, requirements, and constraints in an MEC system.</p>
</list-item>
</list>
</p>
<p>In this paper, we present a comprehensive tutorial on BL in the 5G MEC system. We first review the background of the BL, including its origin, concept of regret, objective, and workflow, then we introduce several basic mathematical parameters to formulate a general BL problem. To deal with the BL problem, we present several popular strategies, including <italic>&#x3f5;</italic>-greedy, upper confidence bound (UCB) algorithm, and weighted policy. Based on the number of agents, BL can be classified as single-agent BL (SA-BL), multi-agent BL (MA-BL), and other types. Specifically, SA-BL can be classified as stateful and stateless for arms with and without states, respectively. The stateful and stateless BL can be further classified into stochastic/non-stochastic and rested/restless, respectively. For each above BL type, we provide several widely used solutions to deal with the corresponding features and issues. Apparently, SA-BL is a special case of MA-BL, which, due to the participation of multiple agents, increases the learning efficiency and enhances the system capacity, yet has two more unique issues: collisions and communications. Collisions happen when one arm is selected by different agents simultaneously, and the topology of the network could be potentially complicated by communication modeling. We present two popular collision models and two communication models, in which we provide the corresponding distributed solutions to alleviate collisions and optimize system performance. Apart from the BL types in SA-BL and MA-BL, we also introduce several important BL variations (i.e., contextual, sleeping, and combinatorial BL) to cover more features, improving the BL model structure from different perspectives. Furthermore, to show how to deploy BL into the 5G MEC system, we introduce three BL applications: contextual sleeping BL, restless BL and contextual calibrated BL. Specifically, we introduce their system models and problem formulations, then provide simulation results to illustrate the excellent performances of the BL algorithms. At last, we envision the possible development avenues of BL In 5G MEC, and identified several research directions together with the associated challenges.</p>
<p>The remainder of this paper is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> introduces the fundamentals of the BL problem and presents a representative classification. SA-BL, MA-BL, and other types are elaborated on, and their features and solutions are given in <xref ref-type="sec" rid="s3">Sections 3</xref>&#x2013;<xref ref-type="sec" rid="s5">5</xref>, respectively. Several applications of BL in 5G MEC and the simulation results are introduced in <xref ref-type="sec" rid="s6">Section 6</xref>. Finally, future challenges and directions are drawn in <xref ref-type="sec" rid="s7">Section 7</xref>.</p>
</sec>
<sec id="s2">
<title>2 Bandit Learning Problem</title>
<sec id="s2-1">
<title>2.1 Background</title>
<p>The classical BL problem comes from a hypothetical experiment where an agent pulls a gambling machine (arm) from a set of such machines, each successive selection of which yields a reward and a state. The agent attempts to obtain a higher total reward through a set of selections. However, due to the lack of prior information about arms, the agent may pull an inferior arm in terms of reward at each round, yielding regret measuring the expected performance loss of the BL process. In other words, regret indicates the reward deviation of the pulled arm from the optimal one. Hence, the agent&#x2019;s objective is to find a policy to improve its arm selection decision by minimizing the expected cumulative regret in the long term.</p>
<p>The workflow of an agent in each learning round is presented in <xref ref-type="fig" rid="F1">Figure 1</xref>, where the agent pulls one arm based on the previously collected knowledge of each arm (i.e., historical reward and the number of selections). The idea is to strike a tradeoff between exploitation and exploration, where exploitation is defined as the investigation of the learned knowledge about the arms and selection of the empirically optimal one and the exploration is defined as the exploration of other arms than the optimal one to get more reward information. The selected arm yields a reward and state associated with itself and/or time. The agent then updates the corresponding parameters (i.e., empirical reward and pulled times) for each arm. Thus, the agent iteratively learns the reward performance of each arm and progressively improves the arm selection.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>The workflow of BL and a representative Classification of BL.</p>
</caption>
<graphic xlink:href="frsip-02-864392-g001.tif"/>
</fig>
</sec>
<sec id="s2-2">
<title>2.2 Mathematical Formulation</title>
<p>We consider a BL system with <italic>K</italic> arms indexed by <inline-formula id="inf1">
<mml:math id="m1">
<mml:mi mathvariant="script">K</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mi>K</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> and the timeline is divided into <italic>T</italic> slots indexed by <inline-formula id="inf2">
<mml:math id="m2">
<mml:mi mathvariant="script">T</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>. We denote by <italic>X</italic>
<sub>
<italic>k</italic>,<italic>t</italic>
</sub> and <inline-formula id="inf3">
<mml:math id="m3">
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bc;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> the reward and the expected reward of pulling arm <italic>k</italic> at slot <italic>t</italic>, respectively. We define <italic>a</italic>
<sub>
<italic>t</italic>
</sub> as the action selection of an arm at slot <italic>t</italic>. The indicator function <inline-formula id="inf4">
<mml:math id="m4">
<mml:mn mathvariant="double-struck">1</mml:mn>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:math>
</inline-formula> indicates that arm <italic>k</italic> is pulled at slot <italic>t</italic>. We then denote by <italic>r</italic>
<sub>
<italic>k</italic>,<italic>t</italic>
</sub> the reward yielded by pulling arm <italic>k</italic> at slot <italic>t</italic>. The objective of the agent is to find a policy <inline-formula id="inf5">
<mml:math id="m5">
<mml:mi mathvariant="script">G</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> to select the optimal arm with the highest expected reward and decrease the regret of the BL process accordingly. Therefore, we define the optimal policy <inline-formula id="inf6">
<mml:math id="m6">
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">G</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2a;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula>, which has the prior knowledge of all arms. It makes sure that the agent can always pull the optimal arm <inline-formula id="inf7">
<mml:math id="m7">
<mml:msubsup>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2a;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> at slot <italic>t</italic> with the highest reward. Hereafter, each symbol with superscript &#x201c;&#x2a;&#x201d; corresponds to that achieved by the optimal policy. Then the general form of the regret can be written as<disp-formula id="e1">
<mml:math id="m8">
<mml:mi>R</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">G</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2a;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msub>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msubsup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x002A;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2a;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="script">G</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(1)</label>
</disp-formula>
</p>
<p>On this basis, the agent decides which arms to pull in a sequence of trials to minimize its BL regret across slots by striking a tradeoff between exploitation and exploration.</p>
</sec>
<sec id="s2-3">
<title>2.3 Bandit Learning Approaches</title>
<p>In the sequel, we review several widely used strategies to tackle the BL problem. These strategies will be modified and extended to obtain a few state-of-the-art algorithms to resolve specific problems, to be detailed in <xref ref-type="sec" rid="s3">Sections 3</xref>&#x2013;<xref ref-type="sec" rid="s6">6</xref>.</p>
<sec id="s2-3-1">
<title>2.3.1 <italic>&#x3f5;</italic>-Greedy </title>
<p>As an intuitive strategy, <italic>&#x3f5;</italic>-greedy (<xref ref-type="bibr" rid="B18">Kuleshov and Precup, 2014</xref>) enables the agent to select an arm with the maximum observed value of the reward based on the current knowledge with a probability smaller than <italic>&#x3f5;</italic> (0 &#x3c; <italic>&#x3f5;</italic> &#x3c; 1). If the probability is larger than <italic>&#x3f5;</italic>, it randomly selects arms. For <italic>&#x3f5;</italic>-greedy, its regret grows linearly in time and its performance can be improved by adjusting the value of <italic>&#x3f5;</italic>.</p>
</sec>
<sec id="s2-3-2">
<title>2.3.2 Upper Confidence Bound Algorithm</title>
<p>UCB (<xref ref-type="bibr" rid="B2">Auer et al., 2002a</xref>) utilizes the confidence intervals on the empirical estimate of the reward of arms and calculates a UCB index for each arm. The UCB index consists of the average reward of each arm and a padding function. The padding function adjusts the exploration-exploitation according to the current slot and the pulled times of each arm. The arm with the highest UCB index will be pulled by the agent at each slot, and the index will be updated according to the received reward and selected times.</p>
</sec>
<sec id="s2-3-3">
<title>2.3.3 Weighted Policy</title>
<p>The weighted policy (<xref ref-type="bibr" rid="B4">Bubeck and Cesa-Bianchi, 2012</xref>) enables the agent to select an arm at each slot based on a mixed probability distribution. The distribution combines a uniform distribution and another one, which weights the arms according to their average regret performance in the past. This method is usually exploited to deal with the non-stochastic reward problems where the reward generation model of each arm cannot be classified into any specific probability distribution.</p>
</sec>
</sec>
<sec id="s2-4">
<title>2.4 Classification of Bandit Learning Models</title>
<p>We can classify BL into several different models as shown in <xref ref-type="fig" rid="F1">Figure 1</xref> based on its settings in terms of the number of agents and the reward generation models. Note that the classification cannot cover all types of BL, and we chose the most popular and representative ones in this paper.</p>
<sec id="s2-4-1">
<title>2.4.1 SA-BL</title>
<p>When the system has only one agent or multiple agents but with a centralized controller, it is referred to as single-agent BL (SA-BL) or centralized BL model. SA-BL can be further classified as stateful (Markov) and stateless for arms with and without states, respectively. The stateless BL can be classified into two forms. It is referred to as the stochastic BL, if the reward is stochastically drawn from a probability distribution, and as the non-stochastic (adversarial) BL otherwise. For the stateful BL, it could also be embodied in two forms: the rested BL if only the state of the pulled arm changes at each slot, and the restless BL if the states of all the arms change.</p>
</sec>
<sec id="s2-4-2">
<title>2.4.2 MA-BL</title>
<p>Furthermore, when the system has multiple agents without a centralized controller, it is referred to as multi-agent BL (MA-BL) or distributed BL model. Consequently, the issue of collision arises due to multiple agents simultaneously pulling the same arm, adding to the calculation complexity of the regret and the reward allocation. In addition, the communication model needs to be carefully designed to accommodate the information exchange among multiple agents.</p>
</sec>
<sec id="s2-4-3">
<title>2.4.3 Other BL Models</title>
<p>Apart from the SA/MA-BL models aforementioned, other significant BL variations, such as sleeping, contextual, and combinatorial BL, can cover more features (e.g., arms availability, contextual information, and multiple selections). These features can also be incorporated into SA-BL and MA-BL, resulting in more complex BL models such as contextual distributed BL (<xref ref-type="bibr" rid="B25">Lu et al., 2010</xref>; <xref ref-type="bibr" rid="B8">Chu et al., 2011</xref>), restless combinatorial BL (<xref ref-type="bibr" rid="B13">Gai et al., 2012</xref>).</p>
</sec>
</sec>
</sec>
<sec id="s3">
<title>3 Single-Agent (Centralized) Bandit Learning</title>
<p>In this section, we elaborate on the single-agent (centralized) BL following the classification in <xref ref-type="fig" rid="F1">Figure 1</xref>. Specifically, the stateless BL and stateful (Markov) BL will be covered.</p>
<sec id="s3-1">
<title>3.1 Stateless Bandit Learning</title>
<p>In stateless BL, the agent only receives the reward of the pulled arm, and independence holds for reward across slots for each arm. In other words, arms do not have states. We have two typical stateless BL types based on different reward generation models: stochastic BL and non-stochastic (adversarial) BL.</p>
<sec id="s3-1-1">
<title>3.1.1 Stochastic Bandit Learning</title>
<p>For stochastic BL, the reward of pulling each arm is stochastically drawn from a specific probability distribution which can be stationary or non-stationary. For the stationary case, the expected reward of each arm is time-independent, i.e., <italic>&#x3bc;</italic>
<sub>
<italic>k</italic>,<italic>t</italic>
</sub> &#x3d; <italic>&#x3bc;</italic>
<sub>
<italic>k</italic>
</sub>. On this basis, the regret function can be written as<disp-formula id="e2">
<mml:math id="m9">
<mml:msub>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>stationary</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>T</mml:mi>
<mml:mo>&#x22c5;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3bc;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2a;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="script">G</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(2)</label>
</disp-formula>
</p>
<p>For the non-stationary case, the expected reward of each arm might change with time. In this case, the previously learned knowledge may not truly reflect the real expected reward of current arms, potentially rendering the historical observations of the pulled arms less useful. Consequently, the BL problem becomes complicated, and the probability of pulling a suboptimal arm increases. Then we have its regret function as<disp-formula id="e3">
<mml:math id="m10">
<mml:msub>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>non</mml:mtext>
<mml:mo>-</mml:mo>
<mml:mtext>sta</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3bc;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x002A;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="script">G</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(3)</label>
</disp-formula>
</p>
<p>UCB family algorithms have been widely adopted to resolve the exploration-exploitation dilemma of the stochastic BL problems. For the stationary case, the basic UCB algorithms suffice. To deal with the non-stationary cases, the sliding-window UCB algorithm (<xref ref-type="bibr" rid="B10">Ding et al., 2019</xref>) could be adopted, only considering the previous observation of a fixed length. Furthermore, we can utilize the discount UCB algorithm (<xref ref-type="bibr" rid="B14">Garivier and Moulines, 2008</xref>), which emphasizes the recent actions by averaging the rewards of arms with a discount factor placing more weight on the recent observations. In addition, some statistical test methods [e.g., generalized likelihood ratio or Page-Hinkly test (<xref ref-type="bibr" rid="B26">Maghsudi and Hossain, 2016</xref>)] can be drawn upon to detect the expected reward changes to improve the arm selection process.</p>
</sec>
<sec id="s3-1-2">
<title>3.1.2 Non-Stochastic (Adversarial) Bandit Learning</title>
<p>For non-stochastic (adversarial) BL, the reward generation of each arm does not have any specific probability distribution. In other words, the reward of each arm is determined by an adversary at each slot rather than by a stochastic generation process. A special case of regret, referred to as weak regret, is usually used to measure the loss of adversarial BL. It considers the single globally optimal arm and can be written as<disp-formula id="e4">
<mml:math id="m11">
<mml:mi>W</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mi>max</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:munder>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="script">G</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(4)</label>
</disp-formula>
</p>
<p>Weighted policy algorithms are usually used to resolve the adversarial reward generation problem in the non-stochastic BL. For full information cases, where the agent observes the total rewards after each selection, we introduce the <italic>Hedge</italic> algorithm (<xref ref-type="bibr" rid="B33">Slivkins, 2019</xref>), whose main idea is to pull an arm with a probability proportional to the average performance of arms in the past. The arms with high rewards quickly gain a high probability of being pulled. For partial information cases, where the agent only observes the reward of the pulled arms, we can utilize the EXP3 algorithm (<xref ref-type="bibr" rid="B18">Kuleshov and Precup, 2014</xref>), which performs the Hedge algorithm as a subroutine with a mixed probability distribution. The distribution combines the uniform distribution and a distribution, which is determined by the weight of each arm. In addition, there are many other variations of the EXP3 algorithm, including EXP3.S, EXP3.P, and EXP4 algorithms. Interested readers are referred to (<xref ref-type="bibr" rid="B3">Auer et al., 2002b</xref>) for details.</p>
</sec>
</sec>
<sec id="s3-2">
<title>3.2 Stateful (Markov) Bandit Learning</title>
<p>In the stateful (Markov BL), each arm has some finite states and a Markov chain, where the probability of the following state only depends on the current state. At each slot, the pulled arm yields a reward drawn from a probability distribution, and the state of the arm changes to a new one based on the Markov state evolution probability. According to different state evolution models, stateful BL can be classified into two types: rested (frozen) BL and restless BL.<list list-type="simple">
<list-item>
<p>&#x2022; In the rested BL model, at each round, only the state of the selected arm evolves with time and the states of other arms are frozen.</p>
</list-item>
<list-item>
<p>&#x2022; In the restless BL model, at each round, all the states of arms (including the unselected arms) might evolve with time.</p>
</list-item>
</list>
</p>
<p>In order to resolve the stateful BL problems, we usually leverage index policies (<xref ref-type="bibr" rid="B39">Whittle, 1980</xref>; <xref ref-type="bibr" rid="B22">Liu and Zhao, 2010a</xref>), which, for each arm, calculate a defined index and provide a proxy to measure the expected reward in the current state. To deal with rested BL, the Gittins index policy (<xref ref-type="bibr" rid="B39">Whittle, 1980</xref>) is usually adopted. This policy works under the Bayesian framework and transforms an <italic>N</italic>-dimension rested BL problem into <italic>N</italic> independent 1-dimension ones, significantly reducing the computational complexity. Furthermore, in the restless BL model, we can leverage Whittle&#x2019;s index policy (<xref ref-type="bibr" rid="B40">Whittle, 1988</xref>). The policy first needs to prove that each arm is indexable, which guarantees the existence of Whittle&#x2019;s index in restless BL. Then it decouples the restless BL problem into multiple sub-problems by applying Lagrangian relaxation for computational simplification.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Multi-Agent (Distributed) Bandit Learning</title>
<p>In this section, we extend single-agent (centralized) BL to multi-agent (distributed) BL. The participation of multiple agents in the arm selection process brings the benefits of increased learning efficiency and enhanced system capacity at the expense of increasing network complication and computational complexity. Two critical issues, collisions and communications between agents, naturally arise due to the introduction of multiple agents. Collisions occur when different agents pull the same arm simultaneously, and communication modeling could potentially complicate the topology of the network.</p>
<p>Apparently, SA-BL is a particular case of MA-BL, and the aforementioned different features of the classifications in SA-BL, shown in <xref ref-type="fig" rid="F1">Figure 1</xref>, are also applicable to MA-BL. For simplicity, we do not repeat the classifications and focus on the issues of collisions and communications.</p>
<sec id="s4-1">
<title>4.1 Modelling of Collisions</title>
<p>We consider <italic>M</italic> agents in MA-BL indexed by <inline-formula id="inf8">
<mml:math id="m12">
<mml:mi mathvariant="script">M</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>m</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, and the policy <inline-formula id="inf9">
<mml:math id="m13">
<mml:mi mathvariant="script">G</mml:mi>
</mml:math>
</inline-formula> for single agent is extended to <inline-formula id="inf10">
<mml:math id="m14">
<mml:mi mathvariant="script">G</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, where the vector <bold>a</bold>
<sub>
<italic>t</italic>
</sub> &#x3d; [<italic>a</italic>
<sub>1,<italic>t</italic>
</sub>, &#x2026; , <italic>a</italic>
<sub>
<italic>m</italic>,<italic>t</italic>
</sub>, &#x2026; , <italic>a</italic>
<sub>
<italic>M</italic>,<italic>t</italic>
</sub>] indicates the actions of all the <italic>M</italic> agents at slot <italic>t</italic>. Other settings are in line with those of the mathematical formulation in <xref ref-type="sec" rid="s2-2">Section 2.2</xref>. Two popular collision models are summarized as follows.<list list-type="simple">
<list-item>
<p>&#x2022; Collision model I: When multiple agents pull the same arm, they share the reward of the arm in a specific (e.g., uniform and arbitrary) manner. In this model, the reward at slot <italic>t</italic> can be written as</p>
</list-item>
</list>
<disp-formula id="e5">
<mml:math id="m15">
<mml:msub>
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msubsup>
<mml:mrow>
<mml:mn mathvariant="double-struck">1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>I</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x22c5;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
</mml:math>
<label>(5)</label>
</disp-formula>where <inline-formula id="inf11">
<mml:math id="m16">
<mml:msubsup>
<mml:mrow>
<mml:mn mathvariant="double-struck">1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>I</mml:mtext>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> equals 1 if arm <italic>k</italic> is pulled at least once, and 0 otherwise.<list list-type="simple">
<list-item>
<p>&#x2022; Collision model II: When multiple agents pull the same arm, none of them obtains a reward. In this model, the reward at slot <italic>t</italic> can be written as</p>
</list-item>
</list>
<disp-formula id="e6">
<mml:math id="m17">
<mml:msub>
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msubsup>
<mml:mrow>
<mml:mn mathvariant="double-struck">1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>II</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x22c5;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
</mml:math>
<label>(6)</label>
</disp-formula>where <inline-formula id="inf12">
<mml:math id="m18">
<mml:msubsup>
<mml:mrow>
<mml:mn mathvariant="double-struck">1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>II</mml:mtext>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> equals 1 if arm <italic>k</italic> is pulled exactly once, and 0 otherwise.</p>
<p>Note that if a collision happens and all the collided agents have the full reward of the arm, it is a trivial problem as it carries no difference from SA-BL.</p>
<p>Following the above models, the regret function of MA-BL can be written as<disp-formula id="e7">
<mml:math id="m19">
<mml:msub>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>MA</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">G</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2a;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msub>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msubsup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2a;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2a;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="script">G</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(7)</label>
</disp-formula>
</p>
<p>To minimize the regret function, a distributed policy needs to be carefully designed to alleviate collisions, although such a policy might not be optimal from the perspective of a single agent.</p>
</sec>
<sec id="s4-2">
<title>4.2 Modelling of Communications</title>
<p>The level of information exchange among agents could vary significantly. Here we introduce two communication models and give the corresponding solutions in the following.</p>
<sec id="s4-2-1">
<title>4.2.1 No Communication</title>
<p>One extreme case is that no information exchange among agents takes place. In such a case, each agent pulls an arm only based on its local observation of the previous selections and rewards. To alleviate collisions, we usually utilize order-optimal policies (<xref ref-type="bibr" rid="B23">Liu and Zhao, 2010b</xref>). These policies design a pre-determined slot allocation pattern so that each agent can use a different slot to select the optimal arm independently. We exemplify such a concept for the case with two agents. At even slots, agent one selects the arm with the highest BL index (e.g., UCB index) and agent two selects the arm with the second-highest index, and vice versa for the odd slots.</p>
</sec>
<sec id="s4-2-2">
<title>4.2.2 Partial Communications</title>
<p>In partial communications, different agents partially communicate with each other. For instance, either an agent can observe other agents&#x2019; selections only when they select the same arm, or an agent can only communicate with agents being its neighbor or within a given number of hops. With this information exchange, agents involved can make a collaborative policy to facilitate the collision reduction, usually with the aid of a graph-based method. We define an undirected coordinate graph <inline-formula id="inf13">
<mml:math id="m20">
<mml:mi>G</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="script">V</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="script">E</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, where <inline-formula id="inf14">
<mml:math id="m21">
<mml:mi mathvariant="script">V</mml:mi>
</mml:math>
</inline-formula> is a set of nodes and <inline-formula id="inf15">
<mml:math id="m22">
<mml:mi mathvariant="script">E</mml:mi>
</mml:math>
</inline-formula> is a set of edges between nodes. We model the MA-BL network as a coordinate graph, where each node <inline-formula id="inf16">
<mml:math id="m23">
<mml:mi>v</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">V</mml:mi>
</mml:math>
</inline-formula> can be viewed as an agent and each edge between a pair of nodes <inline-formula id="inf17">
<mml:math id="m24">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">E</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>&#x2200;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>q</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">V</mml:mi>
</mml:math>
</inline-formula> indicates that the agent <italic>v</italic>
<sub>
<italic>p</italic>
</sub> and <italic>v</italic>
<sub>
<italic>q</italic>
</sub> are neighbors and the two agents may make arm selection collaboratively. Based on the graph, coop-UCB and coop-UCL algorithms (<xref ref-type="bibr" rid="B19">Landgren, 2019</xref>) are developed to deal with the issue of collisions in the BL with partial communications. As another solution, calibrated-based BL (<xref ref-type="bibr" rid="B12">Foster and Vohra, 1997</xref>) enables each agent to simultaneously learn the reward performance of arms and predict the selections of other agents. The equilibrium point will be progressively reached in the BL process, reducing the collision frequency.</p>
</sec>
</sec>
</sec>
<sec id="s5">
<title>5 Other Types of Bandit Learning</title>
<p>Apart from the classifications in SA-BL and MA-BL, shown in <xref ref-type="fig" rid="F1">Figure 1</xref>, there are many other BL models and we will introduce some of them in this section. Here many other BL models exist, which are independent of those in SA-BL and MA-BL, potentially covering more features. They can be incorporated into SA-BL and MA-BL, generating more comprehensive BL models. For simplicity, we only introduce these models with a single agent.</p>
<sec id="s5-1">
<title>5.1 Contextual BL</title>
<p>Contextual BL assumes that the expected reward of each arm is a function of the contextual information. It enables the agent to make a sequential selection from a set of arms based on the observed contextual information and previous knowledge, followed by receiving numerical rewards. This extra information can accelerate the convergence process by learning the underlying connection between arms and contextual information.</p>
<p>LinUCB algorithm (<xref ref-type="bibr" rid="B8">Chu et al., 2011</xref>), developed based on the classical UCB algorithm, is usually leveraged in contextual BL. It takes advantage of the contextual information by maintaining a context-related vector, assuming that the reward of pulling an arm is a function of the vector and the corresponding contextual information. In addition, there are many other solutions, e.g., LinREL and KernelUCB, and interested readers are referred to (<xref ref-type="bibr" rid="B45">Zhou, 2015</xref>) for details.</p>
</sec>
<sec id="s5-2">
<title>5.2 Sleeping BL</title>
<p>In sleeping BL, the set of available arms is time-varying and each arm&#x2019;s state can be exchanged between &#x201c;awakening&#x201d; (can be pulled) and &#x201c;sleeping&#x201d; (cannot be pulled).</p>
<p>The optimal arm may be time-varying and the arm selection becomes more complex, decreasing the learning speed. Moreover, apart from awakening and sleeping, the states of arms can also be mortal (<xref ref-type="bibr" rid="B6">Chakrabarti et al., 2008</xref>). It means that arms are available only for a finite time period, which can be known/unknown and deterministic/stochastic.</p>
<p>The sleeping BL is often resolved by using FTAL and AUER algorithms (<xref ref-type="bibr" rid="B17">Kleinberg et al., 2010</xref>). Since the optimal arm may be sleeping in some slot, these algorithms aim to order in advance all the arms in terms of the expected rewards, and view the optimal ordering of the selections as the optimal policy. As for the mortal case, DetOpt algorithm (<xref ref-type="bibr" rid="B6">Chakrabarti et al., 2008</xref>) solves the selection by pulling each arm several times and abandoning the arm unless it seems promising.</p>
</sec>
<sec id="s5-3">
<title>5.3 Combinatorial BL</title>
<p>In combinatorial BL, we relax the setting that the agent selects one arm for each slot, and assume that a set of arms (a super arm) can be pulled at each slot.</p>
<p>The reward of the pulled super arm is a sum function of weighted value of all the pulled arms&#x2019; rewards, due to the dependencies of arms. Here, the selection space exponentially increases due to the explosion of combinations of arms.</p>
<p>To deal with the issues of selection space and dependencies of arms, we can utilize LLR algorithm (<xref ref-type="bibr" rid="B13">Gai et al., 2012</xref>), which selects a super arm and records the observations both for the arms and the super arm. As an arm might belong to different super arms, LLR could exploit this dependency information to accelerate the learning process. Other algorithms for solving combinatorial BL, such as CUCB and ComBand, can be found in (<xref ref-type="bibr" rid="B5">Cesa-Bianchi and Lugosi, 2012</xref>) for interested readers.</p>
</sec>
</sec>
<sec id="s6">
<title>6 Applications of BL in 5G Mobile Edge Computing</title>
<p>In this section, we introduce several applications of BL in the 5G MEC system. We first present their system models and problem formulations, then provide their simulation results to show the excellent performances of the proposed different BL algorithms. We consider SA-BL in the first two applications and MA-BL in the last one.</p>
<sec id="s6-1">
<title>6.1 Contextual User-Centric Task Offloading for Mobile Edge Computing</title>
<p>In <xref ref-type="bibr" rid="B24">Liu et al. (2022)</xref>, we consider a general user-centric task offloading scheme for MEC in ultra-dense networks (UDN), where a mobile user (MU) randomly moves around the whole network without any predictable tracks. The MU may remain static or move in any direction and computational tasks can be generated sequentially at any location. These tasks will be offloaded to its nearby small base stations (SBSs), which can be referred to as arms. As shown in <xref ref-type="fig" rid="F2">Figure 2</xref>, the dotted lines represent the uncertain future tracks of the MU, and <italic>t</italic>
<sub>
<italic>k</italic>
</sub> and <italic>t</italic>
<sub>
<italic>k</italic>&#x2b;1</sub> indicate the time when <italic>k</italic>th and (<italic>k</italic> &#x2b; 1)-th tasks are generated, respectively.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Illustration of user-centric task offloading for MEC in UDN. The two dotted lines represent the uncertain future tracks of the MU.</p>
</caption>
<graphic xlink:href="frsip-02-864392-g002.tif"/>
</fig>
<p>We formulate the mobile task offloading problem, aiming to minimize the long-term total delay in finishing the tasks of the MU. However, due to the unpredictability of the MU&#x2019;s tracks, many information (e.g., moving tracks, SBS computation capacity, and channel fading gain) cannot be obtained in advance. Besides, the channel conditions (between SBSs and MU) and the delays (of executing different tasks) are changing over time and exhibit randomness. Therefore, the conventional methods cannot tackle the problem.</p>
<p>To address the challenges, we propose the contextual sleeping bandit learning (CSBL) algorithm. The idea is to incorporate the contextual information (e.g., SBS location, service provider, and task type) into the bandit learning to accelerate the convergence process by exploring the underlying relationship between contextual information and arms. At each selection round, the MU first collects the current contextual information of the SBSs and then makes a selection based on the contextual information and previous knowledge of the delay performance of the SBSs. From the perspective of the MU, an available SBS may disappear and then appear again due to the random movement. Hence, we leverage the sleeping bandit learning to enable the MU to identify the status of the arms (sleeping or awakening), thereby accelerating the learning process. We compare the proposed CSBL algorithm with the existing BL algorithms in terms of the cumulative delay, i.e., the cumulated delay over slots, in <xref ref-type="fig" rid="F3">Figure 3</xref>. It can be seen that CSBL outperforms other BL algorithms [i.e., UCB1 (<xref ref-type="bibr" rid="B2">Auer et al., 2002a</xref>), AUER (<xref ref-type="bibr" rid="B17">Kleinberg et al., 2010</xref>), S-LinUCB (<xref ref-type="bibr" rid="B29">Mohamed et al., 2021</xref>), and SEEN (<xref ref-type="bibr" rid="B7">Chen et al., 2018</xref>)] and close to Oracle and Delay-Myopic, which are two algorithms with prior knowledge of arms.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Performance comparison in terms of the cumulative delay.</p>
</caption>
<graphic xlink:href="frsip-02-864392-g003.tif"/>
</fig>
</sec>
<sec id="s6-2">
<title>6.2 Task Offloading for Large-Scale Asynchronous Mobile Edge Computing: An Index Policy Approach</title>
<p>In <xref ref-type="bibr" rid="B41">Xu et al. (2021)</xref>, we consider a MEC system, shown in <xref ref-type="fig" rid="F4">Figure 4</xref>, which has one MEC server and multiple users. Each user runs mission-critical tasks generated randomly, and each task has a deadline. The local computation resource of each user is not sufficient to meet the deadline demand of the tasks. Therefore, a user seeks assistance from resourceful MEC servers by task offloading at the expense of transmission delay and energy.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>A typical asynchronous MEC system with 4 users.</p>
</caption>
<graphic xlink:href="frsip-02-864392-g004.tif"/>
</fig>
<p>Our objective is to design a task offloading policy to meet the deadline requirement of each task while keeping the total transmission energy costs at a low level. However, because the task arrival pattern is stochastic and unpredictable, it is impossible to solve a static combinatorial optimization problem before task arrival. Besides, the computational resources at the MEC server are limited. Thus, it is pretty challenging to balance the transmission energy costs and the deadline requirements.</p>
<p>To deal with the challenges, we develop a new reward function, considering both deadline requirements and transmission energy costs. Then we formulate the task offloading problem as a restless BL one with the objective to maximize the total discounted reward over the time horizon. To solve this problem, we propose a Whittle&#x2019;s index (WI) based policy and rigorously prove the indexability, which guarantees the existence of WI in restless BL. Then we focus on the task completion ratio and propose a shorter slack time less remaining workload (STLW) rule, which identifies the criticality of the tasks by giving priority to the tasks with shorter slack time and less remaining workload. By applying STLW into WI, we develop the STLW-WI algorithm, which could improve the performance of the policy by selecting the users with the highest WI value without violating the STLW. Then we compare our proposed algorithm STLW-WI with the existing BL algorithms in terms of task completion ratio, i.e., the proportion of tasks that can be completed before their deadlines. It is clearly shown in <xref ref-type="fig" rid="F5">Figure 5</xref> that the WI algorithm outperforms other existing algorithms (i.e., Greedy (<xref ref-type="bibr" rid="B41">Xu et al., 2021</xref>), LST (<xref ref-type="bibr" rid="B9">Davis et al., 1993</xref>), EDF (<xref ref-type="bibr" rid="B21">Liu and Layland, 1973</xref>)), and applying STLW into WI can further improve the task completion ratio compared with WI.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Performance comparison in terms of the task completion ratio.</p>
</caption>
<graphic xlink:href="frsip-02-864392-g005.tif"/>
</fig>
</sec>
<sec id="s6-3">
<title>6.3 Multi-Agent Calibrated Bandit Learning With Contextual Information in Mobile Edge Computing</title>
<p>In <xref ref-type="bibr" rid="B44">Zhang et al. (2022)</xref>, we consider an MA-BL MEC system composed of one macro base station (MaBS), multiple microcell base stations (MiBSs), and several randomly located users shown in <xref ref-type="fig" rid="F6">Figure 6</xref>. At each round, the users select MiBSs independently to offload their tasks, and the MaBS monitors the task offloading decisions of all users in this round and broadcasts such information. We incorporate contextual information into our model, such as task size and type. The competition of computational resources among users requires an advanced strategy to reduce the collision.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>An example of the system model with 5 MiBS and 8 users.</p>
</caption>
<graphic xlink:href="frsip-02-864392-g006.tif"/>
</fig>
<p>We formulate the task offloading as an optimization problem, aiming to minimize the long-term average task delay of all users with a restricted number of MiBSs. Different from the conventional static optimization, our problem is quite challenging due to the lack of information about MiBSs and other users. Thus, it is impossible to design a task offloading policy by solving a static non-convex nonlinear integer programming problem. We can develop a centralized task offloading policy by utilizing a collaborative BL method. However, centralized decision making is required, collecting all users&#x2019; task features and then returning the task offloading decisions. A significant transmission delay may be introduced, which is highly undesirable in practice.</p>
<p>To deal with the problem, we develop a decentralized task offloading policy without peer-to-peer information exchange, where all users can make decisions locally. We decouple the formulated problem into several independent contextual BL problems; each user minimizes its own long-term average task delay. We also leverage the calibrated method by developing a calibrated forecaster for each user to predict others&#x2019; actions. On this basis, we propose a contextual calibrated bandit learning (CCBL) algorithm by fusing the BL and calibrated learning. We compare CCBL with the existing BL algorithms in terms of the average delay in <xref ref-type="fig" rid="F7">Figure 7</xref>. It is clearly shown that the proposed CCBL algorithm outperforms other algorithms [i.e., LTS (<xref ref-type="bibr" rid="B44">Zhang et al., 2022</xref>), Myopic (<xref ref-type="bibr" rid="B38">Wang et al., 2020</xref>), <italic>&#x3f5;</italic>-Greedy (<xref ref-type="bibr" rid="B18">Kuleshov and Precup, 2014</xref>), and Calibration (<xref ref-type="bibr" rid="B12">Foster and Vohra, 1997</xref>)] and is only slightly inferior to CGA (<xref ref-type="bibr" rid="B11">Feng et al., 2019</xref>), which is the benchmark with a centralized algorithm.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>The average delay of all users versus task round <italic>r</italic>.</p>
</caption>
<graphic xlink:href="frsip-02-864392-g007.tif"/>
</fig>
</sec>
</sec>
<sec id="s7">
<title>7 Future Research Challenges and Directions</title>
<p>Although dozens of BL algorithms and several applications have been introduced and elaborated on, there are still many issues and solutions which are not involved. Here we present a few possible future directions and challenges on BL in 5G and beyond MEC.</p>
<sec id="s7-1">
<title>7.1 UDN</title>
<p>UDN is a promising paradigm to enable future wireless communications supporting efficient and flexible massive connectivity. This is achieved by deploying a large number of MEC servers, each with a small transmission range at the network edge. Naturally, combining MEC with UDN will significantly increase the coverage of edge computing and provide ubiquitous task offloading services consistently across the whole network (<xref ref-type="bibr" rid="B34">Sun et al., 2017</xref>). From the perspective of users, they have more alternative selections of different MEC servers, increasing the selection space. However, the first issue in the UDN-MEC system is the energy constraint because the servers are usually powered by batteries. Another issue is the migration cost caused by a user selecting different servers to perform task offloading. This issue becomes even severe when users are constantly moving across the network. In addition, other issues, such as the Doppler effect, resource competition, and mobility management, should also be considered. These issues must be carefully addressed when we design BL algorithms. For instance, sleeping/mortal characteristics of servers can be leveraged to model the system, and the contextual information of the offloading environment can be utilized to facilitate the arm selection process.</p>
</sec>
<sec id="s7-2">
<title>7.2 IIoT</title>
<p>The booming of the IIoT brings an exponential increase in industrial devices, calling for more flexible and low-cost communications. The 5G-MEC technologies bring a cyber revolution and considerable benefits to the industry. This is achieved by supporting flexible communication services with low hardware complexity and providing massive and robust connectivity. Numerous industrial devices generate a massive volume of data, which can be utilized for the deployment of BL methods.</p>
<p>Note that industrial data has unique features. For example, some machine-type communication data is usually stable and has a relatively long transmission horizon. Meanwhile, some mission-critical applications generate ultra-reliable and low-latency data, which has short-length packet and require instance execution (<xref ref-type="bibr" rid="B32">Sisinni et al., 2018</xref>). Therefore, different application characteristics and requirements should be considered in the designing and deployment process of BL.</p>
</sec>
<sec id="s7-3">
<title>7.3 FGFA</title>
<p>In most existing communication networks, users trying to access wireless channels have to obtain access permission via a contention-based random access (RA) process with multiple handshakes (<xref ref-type="bibr" rid="B30">Shahab et al., 2020</xref>). However, the excessive delay and signaling overhead involved are unacceptable for many emerging mission-critical applications and those with small-size tasks. Therefore, to improve the access success ratio and the throughput of the MEC system, FGFA is developed to pre-allocate dedicated channels to specific users so that extra handshakes can be spared (<xref ref-type="bibr" rid="B27">Mahmood et al., 2019</xref>). Meanwhile, if the pre-allocated channels are not utilized by users, it will cost unnecessary network resource costs. As a solution, BL could be incorporated to learn the performance of each user and utilize the contextual information to streamline the design of pre-allocation selection policies achieving more efficient network resource utilization.</p>
</sec>
</sec>
<sec id="s8">
<title>8 Conclusion</title>
<p>In this paper, we provided a comprehensive tutorial on BL in 5G MEC, where BL is incorporated for joint consideration of the offloading decision and communication/computation resource allocation. Specifically, we reviewed the fundamental of BL, including background, mathematical formulation, and several popular solutions. We classified BL into three forms, ranging from SA-BL, MA-BL to other types, then we elaborated on each of them and presented several corresponding widely used algorithms. Furthermore, to show how to deploy BL in 5G MEC, we introduced several applications, where the system models and problem formulations were presented, followed by the simulation results to show the excellent performances of the BL algorithms. In addition, we introduced several future directions and challenges on BL in 5G MEC.</p>
</sec>
</body>
<back>
<sec id="s9">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/Supplementary Material, further inquiries can be directed to the corresponding authors.</p>
</sec>
<sec id="s10">
<title>Author Contributions</title>
<p>In this paper, BV and YL developed the whole framework. SL, PC, and ZC proposed the specific methodologies and conducted experimental results.</p>
</sec>
<sec sec-type="COI-statement" id="s11">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s12">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Asheralieva</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Niyato</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Xiong</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Auction-and-learning Based lagrange Coded Computing Model for Privacy-Preserving, Secure, and Resilient mobile Edge Computing</article-title>. <source>IEEE Trans. Mobile Comput.</source>, <fpage>1</fpage>. <pub-id pub-id-type="doi">10.1109/tmc.2021.3097380</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Auer</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Cesa-Bianchi</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Fischer</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Finite-time Analysis of the Multiarmed Bandit Problem</article-title>. <source>Mach. Learn.</source> <volume>47</volume> (<issue>2</issue>), <fpage>235</fpage>&#x2013;<lpage>256</lpage>. <pub-id pub-id-type="doi">10.1023/a:1013689704352</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Auer</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Cesa-Bianchi</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Freund</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Schapire</surname>
<given-names>R. E.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>The Nonstochastic Multiarmed Bandit Problem</article-title>. <source>SIAM J. Comput.</source> <volume>32</volume> (<issue>1</issue>), <fpage>48</fpage>&#x2013;<lpage>77</lpage>. <pub-id pub-id-type="doi">10.1137/s0097539701398375</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bubeck</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Cesa-Bianchi</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Regret Analysis of Stochastic and Nonstochastic Multi-Armed Bandit Problems</article-title>. <source>FNT Machine Learn.</source> <volume>5</volume> (<issue>1</issue>), <fpage>1</fpage>&#x2013;<lpage>122</lpage>. <pub-id pub-id-type="doi">10.1561/2200000024</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cesa-Bianchi</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Lugosi</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Combinatorial Bandits</article-title>. <source>J. Comput. Syst. Sci.</source> <volume>78</volume> (<issue>5</issue>), <fpage>1404</fpage>&#x2013;<lpage>1422</lpage>. <pub-id pub-id-type="doi">10.1016/j.jcss.2012.01.001</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Chakrabarti</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Kumar</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Radlinski</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Upfal</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2008</year>). &#x201c;<article-title>Mortal Multi-Armed Bandits</article-title>,&#x201d; in <conf-name>Proceedings of the 21st International Conference on Neural Information Processing Systems</conf-name>, <conf-loc>Vancouver, British Columbia, Canada</conf-loc>. (<publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates Inc.</publisher-name>), <fpage>273</fpage>&#x2013;<lpage>280</lpage>. </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Ren</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Spatio-Temporal Edge Service Placement: A Bandit Learning Approach</article-title>. <source>IEEE Trans. Wireless Commun.</source> <volume>17</volume> (<issue>12</issue>), <fpage>8388</fpage>&#x2013;<lpage>8401</lpage>. <pub-id pub-id-type="doi">10.1109/twc.2018.2876823</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chu</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Reyzin</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Schapire</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2011</year>). &#x201c;<article-title>Contextual Bandits with Linear Payoff Functions</article-title>,&#x201d; in <conf-name>Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics</conf-name>, <conf-loc>Fort Lauderdale, FL, USA</conf-loc>, <conf-date>11&#x2013;13 Apr</conf-date>, <fpage>208</fpage>&#x2013;<lpage>214</lpage>. </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Davis</surname>
<given-names>R. I.</given-names>
</name>
<name>
<surname>Tindell</surname>
<given-names>K. W.</given-names>
</name>
<name>
<surname>Burns</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>1993</year>). &#x201c;<article-title>Scheduling Slack Time in Fixed Priority Pre-emptive Systems</article-title>,&#x201d; in <conf-name>1993 Proceedings Real-Time Systems Symposium</conf-name>, <conf-loc>Raleigh-Durham, NC, USA</conf-loc>, <conf-date>December</conf-date>, <fpage>222</fpage>&#x2013;<lpage>231</lpage>. </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ding</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Yuan</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Liew</surname>
<given-names>S. C.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Sparsity Learning-Based Multiuser Detection in grant-free Massive-Device Multiple Access</article-title>. <source>IEEE Trans. Wireless Commun.</source> <volume>18</volume> (<issue>7</issue>), <fpage>3569</fpage>&#x2013;<lpage>3582</lpage>. <pub-id pub-id-type="doi">10.1109/twc.2019.2915955</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Feng</surname>
<given-names>W.-J.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>C.-H.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>X.-S.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Multi-user and Multi-Task Offloading Decision Algorithms Based on Imbalanced Edge Cloud</article-title>. <source>IEEE Access</source> <volume>7</volume>, <fpage>95 970</fpage>&#x2013;<lpage>995 977</lpage>. <pub-id pub-id-type="doi">10.1109/access.2019.2928377</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Foster</surname>
<given-names>D. P.</given-names>
</name>
<name>
<surname>Vohra</surname>
<given-names>R. V.</given-names>
</name>
</person-group> (<year>1997</year>). <article-title>Calibrated Learning and Correlated Equilibrium</article-title>. <source>Games Econ. Behav.</source> <volume>21</volume> (<issue>1-2</issue>), <fpage>40</fpage>. <pub-id pub-id-type="doi">10.1006/game.1997.0595</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gai</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Krishnamachari</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Jain</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Combinatorial Network Optimization with Unknown Variables: Multi-Armed Bandits with Linear Rewards and Individual Observations</article-title>. <source>IEEE/ACM Trans. Network.</source> <volume>20</volume> (<issue>5</issue>), <fpage>1466</fpage>&#x2013;<lpage>1478</lpage>. <pub-id pub-id-type="doi">10.1109/tnet.2011.2181864</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Garivier</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Moulines</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>On Upper-Confidence Bound Policies for Non-stationary Bandit Problems</article-title>. <source>arXiv</source>. <comment>arXiv preprint arXiv:0805.3415</comment>. </citation>
</ref>
<ref id="B15">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gittins</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Glazebrook</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Weber</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2011</year>). <source>Multi-armed Bandit Allocation Indices</source>. <publisher-loc>Hoboken, NJ, USA</publisher-loc>: <publisher-name>Wiley</publisher-name>. </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Khan</surname>
<given-names>A. u. R.</given-names>
</name>
<name>
<surname>Othman</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Madani</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Khan</surname>
<given-names>S. U.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>A Survey of mobile Cloud Computing Application Models</article-title>. <source>IEEE Commun. Surv. Tutor.</source> <volume>16</volume> (<issue>1</issue>), <fpage>393</fpage>&#x2013;<lpage>413</lpage>. <pub-id pub-id-type="doi">10.1109/surv.2013.062613.00160</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kleinberg</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Niculescu-Mizil</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sharma</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Regret Bounds for Sleeping Experts and Bandits</article-title>. <source>Mach. Learn.</source> <volume>80</volume> (<issue>2</issue>), <fpage>245</fpage>&#x2013;<lpage>272</lpage>. <pub-id pub-id-type="doi">10.1007/s10994-010-5178-7</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kuleshov</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Precup</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Algorithms for Multi-Armed Bandit Problems</article-title>. <source>CoRR</source>. <comment>arXiv: 1402.6028</comment> <volume>abs/1402.6028</volume>. </citation>
</ref>
<ref id="B19">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Landgren</surname>
<given-names>P. C.</given-names>
</name>
</person-group> (<year>2019</year>). <source>Distributed Multi-Agent Multi-Armed Bandits</source>. <comment>Ph.d. Dissertation</comment>. <publisher-loc>Princeton, NJ, USA</publisher-loc>: <publisher-name>Princeton University</publisher-name>. </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lim</surname>
<given-names>W. Y. B.</given-names>
</name>
<name>
<surname>Ng</surname>
<given-names>J. S.</given-names>
</name>
<name>
<surname>Xiong</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Niyato</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Miao</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>D. I.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Dynamic Edge Association and Resource Allocation in Self-Organizing Hierarchical Federated Learning Networks</article-title>. <source>IEEE J. Select. Areas Commun.</source> <volume>39</volume> (<issue>12</issue>), <fpage>3640</fpage>&#x2013;<lpage>3653</lpage>. <pub-id pub-id-type="doi">10.1109/jsac.2021.3118401</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>C. L.</given-names>
</name>
<name>
<surname>Layland</surname>
<given-names>J. W.</given-names>
</name>
</person-group> (<year>1973</year>). <article-title>Scheduling Algorithms for Multiprogramming in a Hard-Real-Time Environment</article-title>. <source>J. ACM</source> <volume>20</volume> (<issue>1</issue>), <fpage>46</fpage>&#x2013;<lpage>61</lpage>. <pub-id pub-id-type="doi">10.1145/321738.321743</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Indexability of Restless Bandit Problems and Optimality of Whittle index for Dynamic Multichannel Access</article-title>. <source>IEEE Trans. Inform. Theor.</source> <volume>56</volume> (<issue>11</issue>), <fpage>5547</fpage>&#x2013;<lpage>5567</lpage>. <pub-id pub-id-type="doi">10.1109/tit.2010.2068950</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Distributed Learning in Multi-Armed Bandit with Multiple Players</article-title>. <source>IEEE Trans. Signal. Process.</source> <volume>58</volume> (<issue>11</issue>), <fpage>5667</fpage>&#x2013;<lpage>5681</lpage>. <pub-id pub-id-type="doi">10.1109/tsp.2010.2062509</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Xiang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Vucetic</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Contextual User-Centric Task Offloading for mobile Edge Computing in Ultra-dense Networks</article-title>. <source>IEEE Trans. Mobile Comput.</source> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lu</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>P&#xe1;l</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>P&#xe1;l</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2010</year>). &#x201c;<article-title>Contextual Multi-Armed Bandits</article-title>,&#x201d; in <conf-name>Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics</conf-name>, <conf-loc>Chia Laguna Resort, Sardinia, Italy</conf-loc>, <conf-date>13&#x2013;15 May</conf-date>, <fpage>485</fpage>&#x2013;<lpage>492</lpage>. </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Maghsudi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Hossain</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Multi-armed Bandits with Application to 5g Small Cells</article-title>. <source>IEEE Wireless Commun.</source> <volume>23</volume> (<issue>3</issue>), <fpage>64</fpage>&#x2013;<lpage>73</lpage>. <pub-id pub-id-type="doi">10.1109/mwc.2016.7498076</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mahmood</surname>
<given-names>N. H.</given-names>
</name>
<name>
<surname>Abreu</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>B&#xf6;hnke</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Schubert</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Berardinelli</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Jacobsen</surname>
<given-names>T. H.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Uplink grant-free Access Solutions for Urllc Services in 5g New Radio</article-title>,&#x201d; in <conf-name>2019 16th International Symposium on Wireless Communication Systems (ISWCS)</conf-name>, <conf-loc>Oulu, Finland</conf-loc>, <conf-date>August</conf-date>, <fpage>607</fpage>&#x2013;<lpage>612</lpage>. <pub-id pub-id-type="doi">10.1109/iswcs.2019.8877253</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>You</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Letaief</surname>
<given-names>K. B.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>A Survey on mobile Edge Computing: The Communication Perspective</article-title>. <source>IEEE Commun. Surv. Tutor.</source> <volume>19</volume> (<issue>4</issue>), <fpage>2322</fpage>&#x2013;<lpage>2358</lpage>. <pub-id pub-id-type="doi">10.1109/comst.2017.2745201</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mohamed</surname>
<given-names>E. M.</given-names>
</name>
<name>
<surname>Hashima</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Hatano</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Aldossari</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Zareei</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Rihan</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Two-hop Relay Probing in Wigig Device-To-Device Networks Using Sleeping Contextual Bandits</article-title>. <source>IEEE Wireless Commun. Lett.</source> <volume>10</volume> (<issue>7</issue>), <fpage>1581</fpage>&#x2013;<lpage>1585</lpage>. <pub-id pub-id-type="doi">10.1109/lwc.2021.3074972</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shahab</surname>
<given-names>M. B.</given-names>
</name>
<name>
<surname>Abbas</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Shirvanimoghaddam</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Johnson</surname>
<given-names>S. J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Grant-free Non-orthogonal Multiple Access for Iot: A Survey</article-title>. <source>IEEE Commun. Surv. Tutor.</source> <volume>22</volume> (<issue>3</issue>), <fpage>1805</fpage>&#x2013;<lpage>1838</lpage>. <pub-id pub-id-type="doi">10.1109/comst.2020.2996032</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shi</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Edge Computing: Vision and Challenges</article-title>. <source>IEEE Internet Things J.</source> <volume>3</volume> (<issue>5</issue>), <fpage>637</fpage>&#x2013;<lpage>646</lpage>. <pub-id pub-id-type="doi">10.1109/jiot.2016.2579198</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sisinni</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Saifullah</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Jennehag</surname>
<given-names>U.</given-names>
</name>
<name>
<surname>Gidlund</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Industrial Internet of Things: Challenges, Opportunities, and Directions</article-title>. <source>IEEE Trans. Ind. Inf.</source> <volume>14</volume> (<issue>11</issue>), <fpage>4724</fpage>&#x2013;<lpage>4734</lpage>. <pub-id pub-id-type="doi">10.1109/tii.2018.2852491</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Slivkins</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Introduction to Multi-Armed Bandits</article-title>. <source>arXiv</source>. <comment>arXiv preprint arXiv:1904.07272</comment>. <pub-id pub-id-type="doi">10.1561/9781680836219</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Emm: Energy-Aware Mobility Management for mobile Edge Computing in Ultra Dense Networks</article-title>. <source>IEEE J. Select. Areas Commun.</source> <volume>35</volume> (<issue>11</issue>), <fpage>2637</fpage>&#x2013;<lpage>2646</lpage>. <pub-id pub-id-type="doi">10.1109/jsac.2017.2760160</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Sutton</surname>
<given-names>R. S.</given-names>
</name>
<name>
<surname>Barto</surname>
<given-names>A. G.</given-names>
</name>
</person-group> (<year>2018</year>). <source>Reinforcement Learning: An Introduction</source>. <publisher-loc>Cambridge, MA, USA</publisher-loc>: <publisher-name>MIT press</publisher-name>. </citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Teng</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>F. R.</given-names>
</name>
<name>
<surname>Leung</surname>
<given-names>V. C. M.</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Resource Allocation for Ultra-dense Networks: A Survey, Some Research Issues and Challenges</article-title>. <source>IEEE Commun. Surv. Tutor.</source> <volume>21</volume> (<issue>3</issue>), <fpage>2134</fpage>&#x2013;<lpage>2168</lpage>. <pub-id pub-id-type="doi">10.1109/comst.2018.2867268</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>A Survey on mobile Edge Networks: Convergence of Computing, Caching and Communications</article-title>. <source>IEEE Access</source> <volume>5</volume>, <fpage>6757</fpage>&#x2013;<lpage>6779</lpage>. <pub-id pub-id-type="doi">10.1109/access.2017.2685434</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Cui</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Optimal Energy Allocation and Task Offloading Policy for Wireless Powered mobile Edge Computing Systems</article-title>. <source>IEEE Trans. Wireless Commun.</source> <volume>19</volume> (<issue>4</issue>), <fpage>2443</fpage>&#x2013;<lpage>2459</lpage>. <pub-id pub-id-type="doi">10.1109/twc.2020.2964765</pub-id> </citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Whittle</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>1980</year>). <article-title>Multi-armed Bandits and the Gittins index</article-title>. <source>J. R. Stat. Soc. Ser. B Methodol.</source> <volume>42</volume> (<issue>2</issue>), <fpage>143</fpage>&#x2013;<lpage>149</lpage>. <pub-id pub-id-type="doi">10.1111/j.2517-6161.1980.tb01111.x</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Whittle</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>1988</year>). <article-title>Restless Bandits: Activity Allocation in a Changing World</article-title>. <source>J. Appl. Probab.</source> <volume>25</volume>, <fpage>287</fpage>&#x2013;<lpage>298</lpage>. <pub-id pub-id-type="doi">10.1017/s0021900200040420</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Ding</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Vucetic</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Task Offloading for Large-Scale Asynchronous mobile Edge Computing: An index Policy Approach</article-title>. <source>IEEE Trans. Signal. Process.</source> <volume>69</volume>, <fpage>401</fpage>&#x2013;<lpage>416</lpage>. <pub-id pub-id-type="doi">10.1109/tsp.2020.3046311</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Alphones</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Xiong</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Niyato</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Artificial-intelligence-enabled Intelligent 6g Networks</article-title>. <source>IEEE Netw.</source> <volume>34</volume> (<issue>6</issue>), <fpage>272</fpage>&#x2013;<lpage>280</lpage>. <pub-id pub-id-type="doi">10.1109/mnet.011.2000195</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Vucetic</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Online Learning Enabled Task Offloading for Vehicular Edge Computing</article-title>. <source>IEEE Wireless Commun. Lett.</source> <volume>9</volume> (<issue>7</issue>), <fpage>928</fpage>&#x2013;<lpage>932</lpage>. <pub-id pub-id-type="doi">10.1109/lwc.2020.2973985</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Vucetic</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Calibrated Bandit Learning for Decentralized Task Offloading in Ultra-dense Networks</article-title>. <source>IEEE Trans. Commun.</source> <pub-id pub-id-type="doi">10.1109/tcomm.2022.3152262</pub-id> </citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>A Survey on Contextual Multi-Armed Bandits</article-title>. <source>arXiv</source>. <comment>arXiv preprint arXiv:1508.03326</comment>. </citation>
</ref>
</ref-list>
</back>
</article>