<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comms. Net</journal-id>
<journal-title>Frontiers in Communications and Networks</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comms. Net</abbrev-journal-title>
<issn pub-type="epub">2673-530X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">739414</article-id>
<article-id pub-id-type="doi">10.3389/frcmn.2021.739414</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Communications and Networks</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Entropy-Driven Stochastic Federated Learning in Non-IID 6G Edge-RAN</article-title>
<alt-title alt-title-type="left-running-head">Aamer et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Entropy-Driven Stochastic Federated Learning</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Aamer</surname>
<given-names>Brahim</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1403191/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Chergui</surname>
<given-names>Hatim</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Benjillali</surname>
<given-names>Mustapha</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/167643/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Verikoukis</surname>
<given-names>Christos</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/572410/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>National Institute of Posts and Telecommunications (INPT), <addr-line>Rabat</addr-line>, <country>Morocco</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>Centre Tecnologic De Telecomunicacions De Catalunya, <addr-line>Barcelona</addr-line>, <country>Spain</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/970210/overview">Mingzhe Chen</ext-link>, Princeton University, United&#x20;States</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/980863/overview">Zhaohui Yang</ext-link>, King&#x2019;s College London, United&#x20;Kingdom</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1422989/overview">Sihua Wang</ext-link>, Beijing University of Posts and Telecommunications (BUPT), China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Brahim Aamer, <email>aamer.brahim@gmail.com</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Data Science for Communications, a section of the journal Frontiers in Communications and Networks</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>07</day>
<month>10</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>2</volume>
<elocation-id>739414</elocation-id>
<history>
<date date-type="received">
<day>10</day>
<month>07</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>25</day>
<month>08</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Aamer, Chergui, Benjillali and Verikoukis.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Aamer, Chergui, Benjillali and Verikoukis</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Scalable and sustainable AI-driven analytics are necessary to enable large-scale and heterogeneous service deployment in sixth-generation (6G) ultra-dense networks. This implies that the exchange of raw monitoring data should be minimized across the network by bringing the analysis functions closer to the data collection points. While federated learning (FL) is an efficient tool to implement such a decentralized strategy, real networks are generally characterized by time- and space-varying traffic patterns and channel conditions, making thereby the data collected in different points non independent and identically distributed (non-IID), which is challenging for FL. To sidestep this issue, we first introduce a new <italic>a priori</italic> metric that we call <italic>dataset entropy</italic>, whose role is to capture the distribution, the quantity of information, the unbalanced structure and the &#x201c;non-IIDness&#x201d; of a dataset independently of the models. This <italic>a priori</italic> entropy is calculated using a multi-dimensional spectral clustering scheme over both the features and the supervised output spaces, and is suitable for classification as well as regression tasks. The FL aggregation operations support system (OSS) server then uses the reported dataset entropies to devise 1) an entropy-based federated averaging scheme, and 2) a stochastic participant selection policy to significantly stabilize the training, minimize the convergence time, and reduce the corresponding computation cost. Numerical results are provided to show the superiority of these novel approaches.</p>
</abstract>
<kwd-group>
<kwd>dataset entropy</kwd>
<kwd>fast federated learning</kwd>
<kwd>non-iid</kwd>
<kwd>spectral clustering</kwd>
<kwd>stochastic policy</kwd>
<kwd>6G</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>6G wireless networks announces the era of massive heterogeneous digital services, that extend the vertical use cases to the final consumer, which is challenging from a network management point of view. Indeed, in this new context, classical centralized monitoring, analysis, and control would become impractical, as they usually represent a single point of failure and suffer from large overhead. Alternatively, decentralized service processing would bring scalability, low raw data exchange and therefore more system sustainability. In this regard, distributed artificial intelligence (AI) approaches, and in particular FL schemes, can play a pivotal role in leveraging the potential of scattered monitoring data across the network as well as the computing power of edge cloud, while reducing the computational costs and enabling fast local analysis and decision. Nevertheless, FL performance is often limited by the convergence delay due to several conceptual and operational issues that are reviewed in the sequel.</p>
<sec id="s1-1">
<title>1.1 Related Work</title>
<p>In (<xref ref-type="bibr" rid="B2">Brendan McMahan et&#x20;al., 2017</xref>), the authors have proposed the federated averaging (FedAvg) algorithm that synchronously aggregates the parameters, and is thus susceptible to the so-called straggler effect, i.e.,&#x20;each training round only progresses as fast as the slowest edge device since the FL server waits for all devices to complete local training before the global aggregation can be performed. Alternatively, the asynchronous model in (<xref ref-type="bibr" rid="B14">Sprague et&#x20;al., 2018</xref>) has been introduced to improve the scalability and efficiency of FL. For asynchronous FL, the server updates the global model whenever it receives a local update which grants more robustness against participants joining halfway during a training round, as well as when the federation involves participating devices with heterogeneous processing capabilities. However, the model convergence is found to be significantly delayed when data is non independent and identically distributed (non-IID) and unbalanced (<xref ref-type="bibr" rid="B25">Zhao et&#x20;al., 2018</xref>). To solve this issue, it has been proposed to distribute a public dataset to the FL clients at the beginning. However, such a dataset may not always exist, or the participants may refuse to download them for security reasons. Therefore, an alternative solution was to construct an approximately IID dataset using inputs from a limited number of privacy insensitive participants (<xref ref-type="bibr" rid="B23">Yoshida et&#x20;al., 2019</xref>). In the Hybrid-FL protocol, the server asks random participants if they allow their data to be uploaded. During the participant selection phase, apart from selecting participants based on computing capabilities, participants are selected such that their uploaded data can form an approximately IID dataset in the server, i.e.,&#x20;the amount of collected data in each class has close values. Thereafter, the server trains a model on the collected IID dataset, and merges this model with the global model trained by the participants. Nevertheless, requests for data sharing are not in line with the original intent of FL. As an improvement, the authors in (<xref ref-type="bibr" rid="B21">Xie et&#x20;al., 2019</xref>) have proposed the FedAsync algorithm in which newly received local updates are adaptively weighted according to staleness, that is defined as the difference between the current epoch and the iteration to which the received update belongs to. For example, a stale update from a straggler is outdated since it should have been received in previous training rounds. As such, it is given a smaller weight. In addition, the authors prove the convergence guarantee for a restricted family of non-convex problems. However, the current hyperparameters of the FedAsync algorithm still have to be tuned to ensure convergence in different settings. Hence, the algorithm is still unable to generalize to suit the dynamic computation constraints of heterogeneous devices. Given this uncertainty surrounding the reliability of asynchronous FL, synchronous FL remains the most commonly used approach (<xref ref-type="bibr" rid="B8">Keith et&#x20;al., 2019</xref>). In this context, it has been confirmed that the correlation between the model parameters of different clients is increasing as the training progresses, which implies that aggregating parameters directly by averaging may not be a reasonable approach in general (<xref ref-type="bibr" rid="B20">Xiao et&#x20;al., 2020</xref>). Besides, a fair resource federated learning approach has been studied recently in (<xref ref-type="bibr" rid="B15">Tian et&#x20;al., 2020</xref>), which introduces a weighted averaging that gives higher weights to devices with the worst performance (i.e.,&#x20;the largest loss) to let them dominate the objective, and thereby impose more uniformity to the training accuracy. Finally, authors in (<xref ref-type="bibr" rid="B11">Niknam et&#x20;al., 2019</xref>) and (<xref ref-type="bibr" rid="B22">Yang et&#x20;al., 2021</xref>) have listed the different FL motivations, challenges and applications on 6G and wireless communications, where FL has been presented as a solution to address energy, bandwidth, delay and privacy questions in wireless communications. As energy consumption is one of the important aspects to consider in FL, in (<xref ref-type="bibr" rid="B17">Tran et&#x20;al., 2019</xref>) the trade-off between learning time, learning accuracy and terminals power consumption has been investigated.</p>
</sec>
<sec id="s1-2">
<title>1.2 Contributions</title>
<p>In this paper, our contribution is two-fold.<list list-type="simple">
<list-item>
<p>&#x2022; We first introduce the concept of the <italic>entropy</italic> of a dataset in both classification and regression tasks, where we jointly consider the features and supervised outputs to characterize the distribution of its samples and the underlying quantity of information based on a custom spectral clustering strategy. This generalized entropy captures the diversity of a dataset as well as its unbalanced structure and non-IIDness.</p>
</list-item>
<list-item>
<p>&#x2022; By leveraging the proposed entropy as an <italic>a priori</italic> information, we develop two novel FL strategies to make central units (CUs) at 6G Edge-RAN collaborate in learning a certain resource usage, namely, 1) Entropy-weighted federated aggregation which involves all the CUs in the FL training task while prioritizing the most balanced and uncorrelated datasets (i.e.,&#x20;those maximizing the entropy) and 2) Entropy-driven stochastic policy for selecting only a subset of CUs to take part in the FL task. This consists on sampling, at each FL round, the active CUs according to an entropy-based probability distribution, which dramatically reduces the convergence stability and time, as well as the underlying resource consumption by avoiding concurrent training by all CUs at each&#x20;round.</p>
</list-item>
</list>
</p>
</sec>
</sec>
<sec id="s2">
<title>2 Network Description and Data Collection</title>
<sec id="s2-1">
<title>2.1&#x20;Edge-RAN</title>
<p>As depicted in <xref ref-type="fig" rid="F1">Figure&#x20;1</xref>, the considered network corresponds to a 6G edge-RAN under the central unit (CU)/distributed unit (DU) functional split, where each transmission/reception point (TRP) is co-located with its DU, while all CUs are hosted in an edge cloud where they run as virtual network functions (VNFs). Each CU <italic>k</italic> (<italic>k</italic>&#x20;&#x3d; 1, &#x2026;, <italic>K</italic>) performs RAN key performance indicators (KPIs) data collection to build its local dataset <inline-formula id="inf1">
<mml:math id="m1">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> of size <italic>D</italic>
<sub>
<italic>k</italic>
</sub>, where <inline-formula id="inf2">
<mml:math id="m2">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> stands for the input features vector while <inline-formula id="inf3">
<mml:math id="m3">
<mml:msubsup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> represents the corresponding output. Given that this dataset is generally non-exhaustive to train accurate analytical models, the CU takes part in a federated learning task wherein an OSS server&#x2014;located at the core cloud&#x2014;plays the role of a model aggregator. In this work, the CUs and the OSS are connected via fiber transport links, which present a very stable behavior (compared to the wireless channel), and have no effect on the accuracy of the&#x20;FL.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Network architecture.</p>
</caption>
<graphic xlink:href="frcmn-02-739414-g001.tif"/>
</fig>
</sec>
<sec id="s2-2">
<title>2.2 Data Collection</title>
<p>
<xref ref-type="table" rid="T1">Table&#x20;1</xref> shows the features and the supervised output of the local datasets, which have been collected from a live LTE-Advanced (LTE-A) RAN with a granularity of 1&#xa0;h. The considered TRPs cover areas with different traffic profiles&#x2014;both in space and time&#x2014;that tightly depend on the heterogeneous users distribution and behavior in each context (e.g., residential zones, business zones, entertainment events, &#x2026;). On the other hand, the radio KPIs are correlated with the time-varying channel conditions. These realistic datasets are therefore non-IID, which is more challenging for FL algorithms as studied in (<xref ref-type="bibr" rid="B10">Li et&#x20;al., 2021</xref>).</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Dataset features and output.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Feature</th>
<th align="center">Description</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Cell Throughput</td>
<td align="left">Cell Downlink Average Throughput</td>
</tr>
<tr>
<td align="left">User Throughput</td>
<td align="left">User Downlink Average Throughput</td>
</tr>
<tr>
<td align="left">BLER</td>
<td align="left">Average Block Error Rate</td>
</tr>
<tr>
<td align="left">&#x23; Users</td>
<td align="left">Downlink Average Active Users</td>
</tr>
<tr>
<td align="left">MIMO Rank</td>
<td align="left">Average MIMO Rank</td>
</tr>
<tr>
<td align="left">DL PRB</td>
<td align="left">Downlink PRB Usage Percentage</td>
</tr>
<tr>
<td align="left">TA</td>
<td align="left">Average Timing Advance</td>
</tr>
<tr>
<td align="left">CQI</td>
<td align="left">Average Channel Quality Index</td>
</tr>
<tr>
<td align="left">QPSK</td>
<td align="left">Percentage of QPSK modulation usage</td>
</tr>
<tr>
<td align="left">
<bold>Output</bold>
</td>
<td align="center">
<bold>Description</bold>
</td>
</tr>
<tr>
<td align="left">Traffic</td>
<td align="left">Traffic Volume (Output)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s3">
<title>3 Proposed Entropy-Based Federated Learning</title>
<p>To tackle the FL convergence in practical non-IID setups, we seek an objective and compressed metric capturing both the distribution of a dataset and its quantity of information, while not depending on the local models. In this regard, we introduce the notion of <italic>dataset entropy</italic> that is a sufficient statistic to characterize the unbalanced structure of a dataset, as well as its independence from other datasets. Specifically, the entropy is maximized under a uniform distribution with low probability mass function (PMF). By relying on the <italic>a priori</italic> entropies of all CUs, the aggregation server can implement novel CUs selection and models combining schemes to accelerate and stabilize the FL convergence.</p>
<sec id="s3-1">
<title>3.1 Dataset Entropy</title>
<p>Since we are targeting a generalized definition of the entropy, the labels of a classification dataset are not reliable to reflect the distribution of data since it does not apply to regression tasks where the supervised output is continuous, and it omits the effect of the input features. In particular, samples with different feature values but presenting approximately similar outputs are not providing the same information and might not necessarily fit in the same group of data. Therefore, in order to accurately discern the samples, we consider a joint approach where both the features and the supervised output are used. To that end, each CU uses a clustering algorithm that operates on the so-called <italic>similarity matrix</italic> <bold>S</bold>
<sub>
<italic>k</italic>
</sub> whose entries measure the logical correlations between the dataset samples vectors including both the features and the supervised output, i.e.,<disp-formula id="e1">
<mml:math id="m4">
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mspace width="0.17em"/>
<mml:msubsup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(1)</label>
</disp-formula>
</p>
<p>This matrix is built using a radial basis function (RBF) kernel with parameter <italic>&#x3c3;</italic>. As such, the (<italic>i</italic>, <italic>j</italic>)-th matrix element is given by<disp-formula id="e2">
<mml:math id="m5">
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(2)</label>
</disp-formula>where <italic>d</italic> stands for the pairwise logical distance between samples&#x2019; vectors <italic>i</italic> and <italic>j</italic>. Let <italic>F</italic> denote the number of features in the datasets. A general definition of this distance that involves both the features and the output can be written as<disp-formula id="e3">
<mml:math id="m6">
<mml:mi>d</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold">x</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>,</mml:mo>
</mml:math>
<label>(3)</label>
</disp-formula>where {<italic>&#x3b1;</italic>
<sub>
<italic>f</italic>
</sub>} stand for the weights of each feature/output in the distance and verify <inline-formula id="inf4">
<mml:math id="m7">
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:math>
</inline-formula>. They can be fine-tuned to orient the clustering towards the direction of the most relevant features or prioritize the output. For the sake of simplicity, and to avoid generating a high number of scenarios, we settle in this work to the typical setting where the weights are uniform, i.e.,&#x20;<italic>&#x3b1;</italic>
<sub>
<italic>f</italic>
</sub> &#x3d; 1/(<italic>F</italic>&#x20;&#x2b; 1). Further investigation on the effect of the weights on the entropy is left for future&#x20;works.</p>
<p>Since basic clustering algorithms usually require the target number of clusters as an input, we resort to the well-established self-tuning spectral clustering (STSC) technique that presents a time complexity of <inline-formula id="inf5">
<mml:math id="m8">
<mml:mi mathvariant="script">O</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, but is still practical as long as the dataset size <italic>D</italic>
<sub>
<italic>k</italic>
</sub> &#x3c; 10<sup>3</sup> (<xref ref-type="bibr" rid="B18">Tsironis et&#x20;al., 2013</xref>). Since in our case we have only small datasets for each CU (e.g., of size 100)&#x2014;which is by the way one of the reasons to resort to federating learning, the clustering scheme is viable in our&#x20;case.</p>
<p>The STSC relies on the eigenvalues and eigenvectors of the similarity matrix. To that end, we first define <bold>&#x39b;</bold> to be a diagonal matrix with<disp-formula id="e4">
<mml:math id="m9">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">&#x39b;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
</mml:math>
<label>(4)</label>
</disp-formula>and construct the normalized affinity matrix<disp-formula id="e5">
<mml:math id="m10">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">&#x39b;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">&#x39b;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>.</mml:mo>
</mml:math>
<label>(5)</label>
</disp-formula>
</p>
<p>When <bold>&#x39b;</bold> is strictly block diagonal, its eigenvalues and eigenvectors are the union of the eigenvalues and eigenvectors of its blocks padded appropriately with zeros. Let <bold>X</bold>
<sub>
<italic>k</italic>
</sub> denotes the block diagonal matrix gathering the eigenvectors. In this case, we can automatically cluster a dataset into an appropriate number of clusters that minimizes a custom cost function defined in terms of the coefficients of a rotated and normalized version of matrix <bold>X</bold>
<sub>
<italic>k</italic>
</sub> (<xref ref-type="bibr" rid="B24">Zelnik and Pietro, 2004</xref>). Let us assume that for CU <italic>k</italic>, the clustering yields <italic>n</italic>
<sub>
<italic>k</italic>
</sub> clusters <inline-formula id="inf6">
<mml:math id="m11">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> with probabilities <inline-formula id="inf7">
<mml:math id="m12">
<mml:mi>Pr</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>Pr</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula> over dataset <inline-formula id="inf8">
<mml:math id="m13">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, which are calculated via the number of samples per cluster &#x394;<sub>
<italic>k</italic>,<italic>p</italic>
</sub> as<disp-formula id="e6">
<mml:math id="m14">
<mml:mi>Pr</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">&#x394;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo>.</mml:mo>
</mml:math>
<label>(6)</label>
</disp-formula>
</p>
<p>The Corresponding Entropy Is Then Defined as <disp-formula id="e7">
<mml:math id="m15">
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:mi>Pr</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mi>log</mml:mi>
<mml:mspace width="-0.17em"/>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:mi>Pr</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(7)</label>
</disp-formula>
</p>
<p>By letting the CUs report their dataset entropies <inline-formula id="inf9">
<mml:math id="m16">
<mml:msubsup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> to the aggregation server before starting the training, it becomes possible to devise advanced entropy-driven FL strategies that prioritize the CUs with high entropy datasets.</p>
</sec>
<sec id="s3-2">
<title>3.2&#x20;Entropy-Driven FL Combining</title>
<p>In this strategy, the aggregation server directly uses the entropies to perform a weighted averaging of all CUs local models at each round <italic>t</italic>, i.e.,<disp-formula id="e8">
<mml:math id="m17">
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>&#x3b5;</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mfrac>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
</mml:math>
<label>(8)</label>
</disp-formula>where<disp-formula id="e9">
<mml:math id="m18">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>&#x3b5;</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
<label>(9)</label>
</disp-formula>is the cumulative sum of the different CUs entropies that serves as a factor. This allows the CUs with high entropies to dominate and orient the FL training, although this requires the participation of all&#x20;CUs.</p>
<p>
<statement content-type="algorithm" id="alg1">
<p>
<inline-graphic xlink:href="frcmn-02-739414-fx1.tif"/>
</p>
</statement>
</p>
</sec>
<sec id="s3-3">
<title>3.3&#x20;Entropy-Driven Stochastic FL Policy</title>
<p>To optimize the federated learning computation time as well as the underlying resource consumption, we aim at selecting only a number of active CUs in each FL round. In this respect, we introduce an entropy-driven stochastic CU selection policy wherein the aggregation server first generates a probability distribution over all the CUs using their received entropies. In fact, CUs with high entropies hold datasets that are rich in terms of quantity of information and can lead to more generalized models in the training. A direct strategy would consist on selecting the m CUs with highest entropies during all the training. But since the datasets of CUs with low entropy can also hold samples that are non-existing in the other high entropy datasets and yet can help in further generalizing the FL model, the idea we have proposed is to give them a chance by implementing a softmax stochastic policy, where each CU can participate in the training with a probability proportional to its entropy. Hence, in the long-term, even CUs with low entropy are given a chance in some rounds to train the model. This leads to a fast convergence (since it orients the training to the CUs with high potential) while ensuring a more general model at the&#x20;end.</p>
<p>This is achieved by a direct softmax activation layer, i.e.,<disp-formula id="e10">
<mml:math id="m19">
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c0;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>exp</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:mi>exp</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
<mml:mo>.</mml:mo>
</mml:math>
<label>(10)</label>
</disp-formula>
</p>
<p>Next, at each FL round <italic>t</italic>, as illustrated in <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>, the server selects&#x20;a subset of <italic>m</italic>&#x20;&#x3c; <italic>K</italic> CUs to participate in the training by sampling the non-uniform CUs set with probabilities {<italic>&#x3c0;</italic>
<sub>1</sub>, &#x2026;, <italic>&#x3c0;</italic>
<sub>
<italic>K</italic>
</sub>}, i.e.,<disp-formula id="e11">
<mml:math id="m20">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">C</mml:mi>
<mml:mi mathvariant="normal">U</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">C</mml:mi>
<mml:mi mathvariant="normal">U</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x223c;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c0;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c0;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x2223;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">C</mml:mi>
<mml:mi mathvariant="normal">U</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">C</mml:mi>
<mml:mi mathvariant="normal">U</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:math>
<label>(11)</label>
</disp-formula>which ensures that, by the convergence round, the CUs would have stochastically taken part in the FL task according to the initial probability distribution, while avoiding the concurrent training by all CUs at each round. In this case, the model averaging at round <italic>t</italic> is performed as<disp-formula id="e12">
<mml:math id="m21">
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:munder>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>.</mml:mo>
</mml:math>
<label>(12)</label>
</disp-formula>Where <italic>D</italic> is the total samples over all CUs datasets. This entropy-driven stochastic policy is summarized in <xref ref-type="other" rid="alg1">Algorithm 1</xref>, where <inline-formula id="inf10">
<mml:math id="m22">
<mml:mi mathvariant="script">L</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mo>&#x22c5;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mo>&#x22c5;</mml:mo>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> stands for the mean square error (MSE) loss function, and <bold>b</bold> is the bias, while the rest of FL setting parameters is provided in <xref ref-type="table" rid="T2">Table&#x20;2</xref>.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Entropy-driven stochastic federated learning policy.</p>
</caption>
<graphic xlink:href="frcmn-02-739414-g002.tif"/>
</fig>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>FL settings.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Parameter</th>
<th align="center">Description</th>
<th align="center">Value</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">
<italic>T</italic>
</td>
<td align="center">Number of rounds</td>
<td align="center">20</td>
</tr>
<tr>
<td align="center">
<italic>L</italic>
</td>
<td align="center">Number of epochs</td>
<td align="center">50</td>
</tr>
<tr>
<td align="center">
<italic>K</italic>
</td>
<td align="center">Number of CUs</td>
<td align="center">6</td>
</tr>
<tr>
<td align="center">
<italic>M</italic>
</td>
<td align="center">Number of selected CUs</td>
<td align="center">3</td>
</tr>
<tr>
<td align="center">
<italic>D</italic>
<sub>
<italic>k</italic>
</sub>
</td>
<td align="center">Local dataset size</td>
<td align="center">100</td>
</tr>
<tr>
<td align="center">
<italic>&#x03B7;</italic>
</td>
<td align="center">Learning rate</td>
<td align="center">0.001</td>
</tr>
<tr>
<td align="center">
<italic>&#x3a3;</italic>
</td>
<td align="center">Kernel parameter</td>
<td align="center">1.0</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s4">
<title>4 Stochastic Federated Learning Convergence Analysis</title>
<p>In this section, we analyze the convergence probability of the stochastic federated learning. In this intent, a closed-form expression for the lower bound of the convergence probability is derived, reflecting the effects of the CUs selection probability and the datasets&#x20;sizes.</p>
<p>
<statement content-type="theorem" id="Theorem_1">
<label>
<bold>Theorem 1 </bold>
</label>
<p>(Convergence Analysis of the Stochastic Federated Learning). <italic>Consider that the CUs selection in the stochastic federated learning follows a policy</italic> {<italic>&#x3c0;</italic>
<sub>1</sub>, &#x2026;, <italic>&#x3c0;</italic>
<sub>
<italic>K</italic>
</sub>}<italic>, and let</italic> &#x3a9; <italic>and</italic> <italic>B</italic>
<sub>
<italic>k</italic>
</sub> <italic>stand for the upper bounds on the weights and the</italic> norm <italic>of subgradient</italic> <inline-formula id="inf11">
<mml:math id="m23">
<mml:mo>&#x2207;</mml:mo>
<mml:mi mathvariant="script">L</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula>
<italic>, respectively. Let <inline-formula id="inf12">
<mml:math id="m33">
<mml:msub>
<mml:mi>&#x03B1;</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>&#x223C;</mml:mo>
<mml:mi>&#x212C;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>&#x03C0;</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> denote the CU activation bit. Then, the federated learning convergence probability satisfies</italic>
<disp-formula id="e13">
<mml:math id="m24">
<mml:mi mathvariant="bold">P</mml:mi>
<mml:mi mathvariant="bold">r</mml:mi>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:mi mathvariant="bold">E</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi mathvariant="script">L</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="script">L</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x002A;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2265;</mml:mo>
<mml:mi>&#x3d5;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(13)</label>
</disp-formula>where<disp-formula id="e14">
<mml:math id="m25">
<mml:mi>&#x3d5;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mo>{</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3b5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c0;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:msub>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="normal">&#x3a9;</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo>&#x002A;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
<mml:mo>}</mml:mo>
</mml:math>
<label>(14)</label>
</disp-formula>
<italic>Proof</italic>. First, by means of the subgradient inequality we have at round <italic>t</italic>:<disp-formula id="e15">
<mml:math id="m26">
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="script">L</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="script">L</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x002A;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2264;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">&#x27e8;</mml:mo>
<mml:mrow>
<mml:mo>&#x2207;</mml:mo>
<mml:mi mathvariant="script">L</mml:mi>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x002A;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="false">&#x27e9;</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:math>
<label>(15)</label>
</disp-formula>Using Cauchy-Schwarz inequality, we get<disp-formula id="e16">
<mml:math id="m27">
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2264;</mml:mo>
<mml:mfenced open="&#x2016;" close="&#x2016;">
<mml:mrow>
<mml:mo>&#x2207;</mml:mo>
<mml:mi mathvariant="script">L</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mfenced open="&#x2016;" close="&#x2016;">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x22c6;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(16)</label>
</disp-formula>
</p>
<p>By recalling the federated learning aggregation <xref ref-type="disp-formula" rid="e12">Eq. 12</xref>, we can write<disp-formula id="e17">
<mml:math id="m28">
<mml:mo>&#x2207;</mml:mo>
<mml:mi mathvariant="script">L</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2207;</mml:mo>
<mml:mi mathvariant="script">L</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(17)</label>
</disp-formula>
</p>
<p>Therefore, from <xref ref-type="disp-formula" rid="e16">Eqs 16</xref>, <xref ref-type="disp-formula" rid="e17">17</xref> and by invoking the triangle inequality we have<disp-formula id="e18">
<mml:math id="m29">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x2264;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mfenced open="&#x2016;" close="&#x2016;">
<mml:mrow>
<mml:mo>&#x2207;</mml:mo>
<mml:mi mathvariant="script">L</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mfenced open="&#x2016;" close="&#x2016;">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x22c6;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right"/>
<mml:mtd columnalign="left">
<mml:mo>&#x2264;</mml:mo>
<mml:mn>2</mml:mn>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:msub>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="normal">&#x3a9;</mml:mi>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(18)</label>
</disp-formula>
</p>
<p>By the monotonicity of the expectation, we have<disp-formula id="e19">
<mml:math id="m30">
<mml:mi mathvariant="bold">E</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2264;</mml:mo>
<mml:mn>2</mml:mn>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c0;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:msub>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="normal">&#x3a9;</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>C</mml:mi>
<mml:mo>.</mml:mo>
</mml:math>
<label>(19)</label>
</disp-formula>
</p>
<p>By means of Hoeffding-Azuma&#x2019;s inequality (<xref ref-type="bibr" rid="B6">Hoeffding, 1963</xref>), we have<disp-formula id="e20">
<mml:math id="m31">
<mml:mi>Pr</mml:mi>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3c4;</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3c4;</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:mi mathvariant="bold">E</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">E</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>&#x3f5;</mml:mi>
<mml:mo stretchy="false">&#x2223;</mml:mo>
<mml:mi>&#x3c4;</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2265;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mo>{</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3b5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:msup>
<mml:mrow>
<mml:mi>C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
<mml:mo>}</mml:mo>
</mml:math>
<label>(20)</label>
</disp-formula>
</p>
</statement>
</p>
</sec>
<sec id="s5">
<title>5 Numerical Results</title>
<sec id="s5-1">
<title>5.1 Settings and Baselines</title>
<sec id="s5-1-1">
<title>5.1.1 DNN Setting</title>
<p>The structure of the global model weights matrix <bold>W</bold> has been defined by the server to satisfy the findings of (<xref ref-type="bibr" rid="B7">Ke and Liu, 2008</xref>), where the authors have estimated the required number <italic>Q</italic> of neurons per layer based on the number <italic>H</italic> of hidden layers, the dataset sizes <italic>D</italic>
<sub>
<italic>k</italic>
</sub>, and the number of features <italic>F</italic> as<disp-formula id="e21">
<mml:math id="m32">
<mml:mi>Q</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:msqrt>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>max</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>D</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
<mml:mrow>
<mml:mi>H</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:math>
<label>(21)</label>
</disp-formula>which is confirmed via <xref ref-type="fig" rid="F3">Figures 3</xref>, <xref ref-type="fig" rid="F4">4</xref>, where the best setting of the DNN model neurons turns out to be <italic>Q</italic>&#x20;&#x3d; 4 for <italic>H</italic>&#x20;&#x3d; 3. As a benckmark, the performance of our proposed approaches is compared with LossFedAvg (<xref ref-type="bibr" rid="B10">Li et&#x20;al., 2021</xref>) and FedAvg (<xref ref-type="bibr" rid="B2">Brendan McMahan et&#x20;al., 2017</xref>). FL settings are listed on <xref ref-type="table" rid="T2">Table&#x20;2</xref>, where FL system consists of <italic>K</italic>&#x20;&#x3d; 6 DUs running local DNN with a learning rate <italic>&#x3b7;</italic> &#x3d; 0.001 for <italic>T</italic>&#x20;&#x3d; 20 rounds.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Different number of layers configurations comparison for <italic>&#x3b7;</italic> &#x3d; 0.001.</p>
</caption>
<graphic xlink:href="frcmn-02-739414-g003.tif"/>
</fig>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Different number of neurons per layer configurations comparison for <italic>&#x3b7;</italic> &#x3d; 0.001.</p>
</caption>
<graphic xlink:href="frcmn-02-739414-g004.tif"/>
</fig>
</sec>
<sec id="s5-1-2">
<title>5.1.2 Learning Rate</title>
<p>The learning rate is a key parameter in ML models, therefore we have to select carefully its right value. In this perspective, we have simulated different learning rate values to illustrate Entropy-Weighted model convergence behaviour. In this respect, <xref ref-type="fig" rid="F5">Figure&#x20;5</xref> shows fast convergence of Entropy-Weighted model with learning rate <italic>&#x3b7;</italic> &#x3d; 0.01, while for <italic>&#x3b7;</italic> &#x3d; 0.001 it is showing a stable yet more slow convergence to the same loss as the case of <italic>&#x3b7;</italic> &#x3d; 0.01. Note that the adopted DNN optimizer is <italic>Adam optimizer</italic> (<xref ref-type="bibr" rid="B9">Kingma and Ba, 2015</xref>).</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Convergence of entropy-weighted vs. learning&#x20;rate.</p>
</caption>
<graphic xlink:href="frcmn-02-739414-g005.tif"/>
</fig>
</sec>
</sec>
<sec id="s5-2">
<title>5.2 Numerical Results Analysis</title>
<sec id="s5-2-1">
<title>5.2.1 Convergence</title>
<p>
<xref ref-type="fig" rid="F6">Figures 6A,B</xref> illustrate the gains achieved by the entropy-weighted approach compared to the baseline FedAvg and LossFedAvg. The comparison is done for both balanced and unbalanced non IID datasets. As showcased in <xref ref-type="table" rid="T3">Table&#x20;3</xref>, the entropy metric varies in balanced datasets, since the clustering technique takes into account the correlation between features as well as the supervised output. In the unbalanced scenario, the entropy difference between CUs is even clearer and demonstrates also that datasets with smaller size can sometimes yield more clusters compared to larger datasets, which further corroborates the role of the introduced entropy metric in characterizing a dataset efficiently.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>FL training loss vs. number of rounds for <italic>&#x3b7;</italic> &#x3d; 0.001.</p>
</caption>
<graphic xlink:href="frcmn-02-739414-g006.tif"/>
</fig>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Results: Datasets clustering.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="center">CU number</th>
<th colspan="3" align="center">Balanced</th>
<th colspan="3" align="center">Unbalanced</th>
</tr>
<tr>
<th align="center">Nb samples</th>
<th align="center">Nb clusters</th>
<th align="center">Entropy</th>
<th align="center">Nb samples</th>
<th align="center">Nb clusters</th>
<th align="center">Entropy</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">1</td>
<td align="center">100</td>
<td align="center">2</td>
<td align="char" char=".">0.692</td>
<td align="center">100</td>
<td align="center">2</td>
<td align="char" char=".">0.692</td>
</tr>
<tr>
<td align="center">2</td>
<td align="center">100</td>
<td align="center">2</td>
<td align="char" char=".">0.592</td>
<td align="center">70</td>
<td align="center">3</td>
<td align="char" char=".">1.026</td>
</tr>
<tr>
<td align="center">3</td>
<td align="center">100</td>
<td align="center">2</td>
<td align="char" char=".">0.676</td>
<td align="center">90</td>
<td align="center">2</td>
<td align="char" char=".">0.515</td>
</tr>
<tr>
<td align="center">4</td>
<td align="center">100</td>
<td align="center">3</td>
<td align="char" char=".">0.998</td>
<td align="center">80</td>
<td align="center">4</td>
<td align="char" char=".">1.238</td>
</tr>
<tr>
<td align="center">5</td>
<td align="center">100</td>
<td align="center">3</td>
<td align="char" char=".">1.051</td>
<td align="center">50</td>
<td align="center">3</td>
<td align="char" char=".">1.068</td>
</tr>
<tr>
<td align="center">6</td>
<td align="center">100</td>
<td align="center">2</td>
<td align="char" char=".">0.676</td>
<td align="center">60</td>
<td align="center">2</td>
<td align="char" char=".">0.690</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>A slightly lower losses are met with the entropy-weighted approach rather than the entropy stochastic policy, but both methods have the same convergence trend. In <xref ref-type="fig" rid="F6">Figure&#x20;6A,B</xref> both entropy-based FL converge faster than FedAvg and LossFedAvg. Knowing how critical is the bandwidth occupation for FL exchanges, and how the CUs local model training is power consuming, especially in 6G mobile systems, our introduced entropy stochastic policy shows good results. This aspect becomes more critical if the FL result is an input for fast decision-making algorithms such as network slicing orchestration or resources scheduling.</p>
<p>Better than FedAvg and LossFedAvg, the entropy stochastic policy convergence trend is oscillating around entropy-weighted as in <xref ref-type="fig" rid="F6">Figure&#x20;6A,B</xref>.</p>
</sec>
<sec id="s5-2-2">
<title>5.2.2 Time Complexity and Scalability</title>
<p>Another important achievement with the entropy stochastic policy is the reduction of the required time for a given number of rounds and exchanges between the OSS server and the CUs towards convergence, as shown in <xref ref-type="fig" rid="F7">Figure&#x20;7</xref>, wherein the convergence time difference between the entropy-weighted approach and the entropy stochastic policy is exponentially growing with the number of FL rounds. Note that the corresponding wall-clock time performance is tightly dependent on the computation capabilities of both the OSS server and the CUs, but it shows that the stochastic policy FL minimizes the computation burden by selecting only a subset of CUs to take part in the training according to their <italic>prior</italic> entropy measure, no matter how the number of CUs grows in the network. This proves the scalability of the proposed stochastic FL in large-scale deployments scenarios. More results can be generated for different values of <italic>K</italic> and&#x20;<italic>m</italic>.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Models time convergence for <italic>&#x3b7;</italic> &#x3d; 0.001.</p>
</caption>
<graphic xlink:href="frcmn-02-739414-g007.tif"/>
</fig>
</sec>
<sec id="s5-2-3">
<title>5.2.3 Learning Rate Sets</title>
<p>We have trained both entropy-weighted and LossFedAvg models using specific learning rate per each FL CU. As illustrated in <xref ref-type="fig" rid="F8">Figure&#x20;8</xref>, better convergence is achieved with both used sets of learning rates compared to fixed <italic>&#x3b7;</italic> &#x3d; 0.001. Where <italic>set1</italic> is a random selection of CUs learning rates, while in <italic>set2</italic>, <italic>&#x3b7;</italic> has been chosen according to each CU&#x2019;s entropy value, i.e.,&#x20;CUs with high entropy are assigned small <italic>&#x3b7;</italic> values and vice-versa. Note that the random learning rate strategy exhibits unstable convergence since it allows CUs with low entropy to learn faster and therefore dominate in some&#x20;cases.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>Models convergence for heterogeneous learning rate&#x20;sets.</p>
</caption>
<graphic xlink:href="frcmn-02-739414-g008.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec id="s6">
<title>6 Conclusion</title>
<p>In this paper, we have introduced a novel <italic>a priori</italic> metric termed <italic>dataset entropy</italic> to characterize the distribution, the quantity of information, the unbalanced structure and the &#x201c;non-IIDness&#x201d; of a dataset independently of the models. This entropy is calculated via a generalized clustering strategy that relies on a custom similarity matrix defined over both the features and the supervised output spaces, and supporting both classification and regression tasks. The entropy metric has been then adopted to develop 1) an entropy-based federated averaging scheme, and 2) a stochastic CU selection policy to significantly stabilize the training, minimize the convergence time, and reduce the corresponding computation cost. Numerical results have been provided to corroborate these findings. In particular, the convergence time difference between Entropy-Weighted and Entropy Stochastic Policy schemes is exponentially growing with the number of FL rounds. Another important result is Entropy Stochastic Policy model convergence, which is better than FedAvg and LossFedAvg and oscillating near Entropy-Weighted&#x20;model.</p>
</sec>
</body>
<back>
<sec id="s7">
<title>Data Availability Statement</title>
<p>The datasets presented in this article are not readily available because the dataset is protected by IPR of the operator. Requests to access the datasets should be directed to <ext-link ext-link-type="uri" xlink:href="http://aamer.brahim@gmail.com">aamer.brahim@gmail.com</ext-link>.</p>
</sec>
<sec id="s8">
<title>Author Contributions</title>
<p>BA: Algorithms proposal and implementation. HC: Stochastic&#x20;policy proposal. MB: Results analysis and recommendations proposal. CV: Results analysis and recommendations proposal.</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s10" sec-type="disclaimer">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Brendan McMahan</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Moore</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Ramage</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Hampson</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ag&#xfc;era y Arcas</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Communication-efficient Learning of Deep Networks from Decentralized Data</article-title>,&#x201d; in <source>Proceedings of the 20 th International Conference on Artificial Intelligence and Statistics (AISTATS)</source>. <publisher-name>JMLR: W&#x0026;CP</publisher-name> <volume>54</volume>. </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hoeffding</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>1963</year>). <article-title>Probability Inequalities for Sums of Bounded Random Variables</article-title>. <source>J.&#x20;Am. Stat. Assoc.</source> <volume>58</volume> (<issue>301</issue>), <fpage>13</fpage>&#x2013;<lpage>30</lpage>. </citation>
</ref>
<ref id="B7">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ke</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2008</year>). &#x201c;<article-title>Empirical Analysis of Optimal Hidden Neurons in Neural Network Modeling for Stock Prediction</article-title>,&#x201d; in <conf-name>IEEE Pacific-Asia Workshop on Computational Intelligence and Industrial Application</conf-name>, <conf-loc>Wuhan, China</conf-loc>, <conf-date>December 19-20, 2008</conf-date>. <pub-id pub-id-type="doi">10.1109/paciia.2008.363</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Keith</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Eichner</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Grieskamp</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Huba</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Ingerman</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Ivanov</surname>
<given-names>V.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Towards Federated Learning at Scale: System Design</article-title>. <comment>[Online]</comment>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1902.01046">arxiv.org/abs/1902.01046</ext-link>
</comment>. </citation>
</ref>
<ref id="B9">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Kingma</surname>
<given-names>D. P.</given-names>
</name>
<name>
<surname>Ba</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Adam: A Method for Stochastic Optimization</article-title>,&#x201d; in <conf-name>3rd International Conference for Learning Representations</conf-name>, <conf-loc>San Diego</conf-loc>, <conf-date>Jul. 2015</conf-date>. </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Diao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Federated Learning on Non-IID Data Silos: An Experimental Study</article-title>. <source>Comput. Sci.</source> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Niknam</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Dhillon</surname>
<given-names>H. S.</given-names>
</name>
<name>
<surname>Reed</surname>
<given-names>J.&#x20;H.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Federated Learning for Wireless Communications: Motivation, Opportunities and Challenges</article-title>. <source>Electr. Eng. Syst. Sci.</source> <volume>58</volume> (<issue>6</issue>), <comment>June 2020</comment> <fpage>46</fpage>&#x2013;<lpage>51</lpage>. <pub-id pub-id-type="doi">10.1109/MCOM.001.1900461</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Sprague</surname>
<given-names>M. R.</given-names>
</name>
<name>
<surname>Jalalirad</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Scavuzzo</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Capota</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Neun</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Do</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Kopp</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Asynchronous Federated Learning for Geospatial Applications</article-title>,&#x201d; in <conf-name>Joint European Conference on Machine Learning and Knowledge Discovery in Databases</conf-name>, <conf-date>March, 2019</conf-date> (<publisher-name>Springer</publisher-name>) <volume>967</volume>, <fpage>21</fpage>&#x2013;<lpage>28</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-14880-5_2</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Tian</surname>
<given-names>Li.</given-names>
</name>
<name>
<surname>Sanjabi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ahmad</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Smith</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2020</year>). <source>Fair Resource Allocation in Federated Learning</source>. <publisher-loc>Addis Ababa, Ethiopia</publisher-loc>: <publisher-name>ICLR</publisher-name>. </citation>
</ref>
<ref id="B17">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Tran</surname>
<given-names>N. H.</given-names>
</name>
<name>
<surname>Bao</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Albert</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Minh</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Hong</surname>
<given-names>C. S.</given-names>
</name>
<name>
<surname>Nguyen</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Federated Learning over Wireless Networks: Optimization Model Design and Analysis</article-title>,&#x201d; in <conf-name>IEEE INFOCOM 2019-IEEE Conference on Computer Communications</conf-name>, <conf-loc>Paris, France</conf-loc>, <conf-date>April 29&#x2013;May 2, 2019</conf-date>. </citation>
</ref>
<ref id="B18">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Tsironis</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sozio</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Vazirgiannis</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Poltechnique</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2013</year>). &#x201c;<article-title>Accurate Spectral Clustering for Community Detection in Mapreduce</article-title>,&#x201d; in <conf-name>&#xc2;&#x17d;&#xc2;&#x17d; NIPS Workshops</conf-name>, <conf-loc>Serbia</conf-loc>, <conf-date>September 2018</conf-date>. <pub-id pub-id-type="doi">10.1109/SISY.2018.8524662</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xiao</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Stankovic</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Vukobratovic</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Averaging Is Probably Not the Optimum Way of Aggregating Parameters in Federated Learning</article-title>. <source>Entropy (Basel)</source> <volume>22</volume> (<issue>3</issue>). <pub-id pub-id-type="doi">10.3390/e22030314</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Xie</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Koyejo</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Gupta</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Asynchronous Federated Optimization</article-title>. <comment>[Online]</comment>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1903.03934">arxiv.org/abs/1903.03934</ext-link>
</comment>. </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kai-Kit Wong</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Poor</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Cui</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Federated Learning for 6G: Applications, Challenges, and Opportunities</article-title>. <source>Comput. Sci.</source> </citation>
</ref>
<ref id="B23">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Yoshida</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Nishio</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Morikura</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Yamamoto</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Yonetani</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Hybrid-FL: Cooperative Learning Mechanism Using Non-IID Data in Wireless Networks</article-title>. <comment>[Online]</comment>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1905.07210">arxiv.org/abs/1905.07210</ext-link>
</comment>. </citation>
</ref>
<ref id="B24">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zelnik</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Pietro</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2004</year>). &#x201c;<article-title>Self-Tuning Spectral Clustering</article-title>,&#x201d; in <conf-name>the 17th International Conference on Neural Information Processing Systems (NIPS&#x2019;04)</conf-name>, <conf-loc>Vancouver</conf-loc>, <conf-date>January 2004</conf-date>, <fpage>1601</fpage>&#x2013;<lpage>1608</lpage>. </citation>
</ref>
<ref id="B25">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Meng</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Lai</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Suda</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Civin</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Chandra</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Federated Learning with Non-IID Data</article-title>. <comment>[Online]</comment>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1806.00582">arxiv.org/abs/1806.00582</ext-link>
</comment>. </citation>
</ref>
</ref-list>
</back>
</article>