<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Appl. Math. Stat.</journal-id>
<journal-title>Frontiers in Applied Mathematics and Statistics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Appl. Math. Stat.</abbrev-journal-title>
<issn pub-type="epub">2297-4687</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fams.2023.1267034</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Applied Mathematics and Statistics</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Federated statistical analysis: non-parametric testing and quantile estimation</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Becher</surname> <given-names>Ori</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/2428295/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/>
<role content-type="https://credit.niso.org/contributor-roles/investigation/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Marcus-Kalish</surname> <given-names>Mira</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1663324/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/funding-acquisition/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Steinberg</surname> <given-names>David M.</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1366059/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/funding-acquisition/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/supervision/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff><institution>Department of Statistics and Operations Research, Tel Aviv University</institution>, <addr-line>Tel Aviv</addr-line>, <country>Israel</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Jun Fan, Hong Kong Baptist University, Hong Kong SAR, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Giovanni Cicceri, University of Messina, Italy; Zhan Yu, Hong Kong Baptist University, Hong Kong SAR, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: David M. Steinberg <email>dms&#x00040;tauex.tau.ac.il</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>13</day>
<month>11</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>9</volume>
<elocation-id>1267034</elocation-id>
<history>
<date date-type="received">
<day>25</day>
<month>07</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>03</day>
<month>10</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2023 Becher, Marcus-Kalish and Steinberg.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Becher, Marcus-Kalish and Steinberg</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>The age of big data has fueled expectations for accelerating learning. The availability of large data sets enables researchers to achieve more powerful statistical analyses and enhances the reliability of conclusions, which can be based on a broad collection of subjects. Often such data sets can be assembled only with access to diverse sources; for example, medical research that combines data from multiple centers in a federated analysis. However these hopes must be balanced against data privacy concerns, which hinder sharing raw data among centers. Consequently, federated analyses typically resort to sharing data summaries from each center. The limitation to summaries carries the risk that it will impair the efficiency of statistical analysis procedures. In this work, we take a close look at the effects of federated analysis on two very basic problems, non-parametric comparison of two groups and quantile estimation to describe the corresponding distributions. We also propose a specific privacy-preserving data release policy for federated analysis with the <italic>K</italic>-anonymity criterion, which has been adopted by the Medical Informatics Platform of the European Human Brain Project. Our results show that, for our tasks, there is only a modest loss of statistical efficiency.</p></abstract>
<kwd-group>
<kwd>federated analysis</kwd>
<kwd>Mann-Whitney test</kwd>
<kwd>medical informatics</kwd>
<kwd>privacy preservation</kwd>
<kwd>information loss</kwd>
</kwd-group>
<contract-num rid="cn001">785907</contract-num>
<contract-sponsor id="cn001">European Research Council<named-content content-type="fundref-id">10.13039/501100000781</named-content></contract-sponsor>
<counts>
<fig-count count="5"/>
<table-count count="3"/>
<equation-count count="20"/>
<ref-count count="26"/>
<page-count count="13"/>
<word-count count="7973"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Statistics and Probability</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>The ability to analyze large sets of medical data has clear potential for improving health care. Often, though, a large patient base is available only by combining data from multiple silos. Combining data faces immediate challenges: data quality is often not uniform, nor is granularity; sites may code data differently, requiring adjustment before analysis is possible. Additionally, given the personal and sensitive nature of medical information, sharing data across centers poses ethical and legal concerns. Many countries have enacted laws protecting privacy. For example, data sharing in Europe must be consistent with the General Data Protection Regulation (&#x0201C;GDPR&#x0201D;) and in the United States with the Health Insurance Portability and Accountability Act (&#x0201C;HIPAA&#x0201D;).</p>
<p>Federated data analysis addresses privacy concerns by limiting data release from a center to summary statistics, without revealing the raw data. The analysis must then rely on the summary statistics. Federated analyses have been used to study a variety of medical problems, including the clinical impact of atrial fibrillation for dementia [<xref ref-type="bibr" rid="B1">1</xref>], attenuation and scatter correction for PET images [<xref ref-type="bibr" rid="B2">2</xref>], mortality following transcatheter aortic valve replacement [<xref ref-type="bibr" rid="B3">3</xref>], detection of cancer boundaries [<xref ref-type="bibr" rid="B4">4</xref>] and histological response to chemotherapy in a rare form of breast cancer [<xref ref-type="bibr" rid="B5">5</xref>].</p>
<p>The applications above all exploit methods for federated analysis that have been proposed in the machine learning literature; however, little research has been done to examine the consequences of the methods for statistical inference. Our goal in this paper is to fill some of the gap, assessing the loss in statistical efficiency when using federated data for some basic statistical analyses within a particular privacy protection protocol.</p>
<p>The roots of our work are in the European Human Brain Project (&#x0201C;HBP&#x0201D;). Data sharing is a major priority for the HBP, but must be fully consistent with the GDPR. Salles et al. [<xref ref-type="bibr" rid="B6">6</xref>] spelled out a detailed Opinion and Action Plan on &#x0201C;Data Protection and Privacy&#x0201D; for the HBP. The plan gives important guidelines and a sound administrative framework for data protection, but does not present technical solutions. Several measures of privacy have been proposed. One of the measures is the degree of anonymization&#x02014;the extent to which one is able to identify an individual from the records in the data and link the sensitive information to her. A well-known criterion for anonymization is <italic>K</italic>-anonymity [<xref ref-type="bibr" rid="B7">7</xref>]. A dataset is <italic>K</italic>-anonymous if each data item cannot be distinguished from at least <italic>K</italic>&#x02212;1 other data items. Fulfilling this criterion introduces fuzziness into the data that makes it less likely to expose a certain individual. One of the techniques to achieve <italic>K</italic>-anonymity is generalization. For example, one could release that <italic>K</italic> patients were between age 10 and 30 instead of releasing the exact ages of each of these patients. Another popular criterion is differential privacy, in which querying a database must not reveal too much information about a specific individual&#x00027;s record in it [<xref ref-type="bibr" rid="B8">8</xref>].</p>
<p>The Medical Informatics Platform (&#x0201C;MIP&#x0201D;), the HBP vehicle for federated, multi-institutional, data analysis, adopted the <italic>K</italic>-anonymity criterion for privacy protection. Specifically, any data table exported from a member institution for use in federated analysis on the MIP must have at least 10 subjects in any cell of the table. Consequently, we chose to study the effect of federated analysis on statistical efficiency when using the MIP implementation of the <italic>K</italic>-anonymity criterion.</p>
<p>In Section 3, we propose a method for data summary that supports the <italic>K</italic>-anonymity criterion used in the MIP. We then address two common statistical problems: (i) use of the nonparametric Mann-Whitney U statistic (henceforth &#x0201C;MWU&#x0201D;) [<xref ref-type="bibr" rid="B9">9</xref>] to test the hypothesis that there is no difference between two groups (in Section 4); and (ii) quantile estimation to describe the corresponding distributions (in Section 5). We find that federated procedures are almost as sensitive as the full-data methods for these problems. For quantile estimation, they can be even more sensitive due to the need for a more sophisticated estimation strategy. Discussion and conclusions are in Section 6.</p></sec>
<sec id="s2">
<title>2. Related work</title>
<p>Most of the research on methods for federated data analysis has focused on the predictive models commonly used in machine learning, under the general header of Federated Learning. These works emphasize the adjustment of machine learning algorithms to federated settings, addressing algorithmic problems, security, and communication efficiency. Several recent surveys provide good summaries [<xref ref-type="bibr" rid="B10">10</xref>&#x02013;<xref ref-type="bibr" rid="B12">12</xref>]. A notable example is [<xref ref-type="bibr" rid="B13">13</xref>], who presented the &#x0201C;FederatedAveraging&#x0201D; algorithm, which combines local stochastic gradient descent on each client with a server that averages results across clients. Two related methods that also deal with inter-site heterogeneity are &#x0201C;FedProx&#x0201D; [<xref ref-type="bibr" rid="B14">14</xref>], and &#x0201C;FedBN&#x0201D; [<xref ref-type="bibr" rid="B15">15</xref>]. Hwang et al. [<xref ref-type="bibr" rid="B16">16</xref>] proposed the &#x0201C;FedPxN&#x0201D; algorithm, which modifies the way in which local site models are aggregated, and used publicly available medical data to compare the accuracy of these algorithms on several classification tasks.</p>
<p>Research with an emphasis on statistical inference has been less prominent. Nasirigerdeh et al. [<xref ref-type="bibr" rid="B17">17</xref>] created sPLINK, a system used to conduct Genome-Wide Association studies in a federated manner while respecting privacy. Algorithms such as linear and logistic regression were adjusted to the federated setting using data summaries from different data centers. Duan et al. [<xref ref-type="bibr" rid="B18">18</xref>, <xref ref-type="bibr" rid="B19">19</xref>] presented privacy-preserving distributed algorithms (&#x0201C;ODAL&#x0201D; and &#x0201C;ODAL2&#x0201D;) to perform logistic regression. With a focus on efficient communication, they made these <italic>one-shot algorithms</italic>, i.e., using only one information transfer from each center; by contrast, most algorithms are iterative and require multiple transfers. Liu and Ihler [<xref ref-type="bibr" rid="B20">20</xref>] considered federated maximum likelihood estimation for parameters in exponential family distribution models. Their idea was to combine local maximum likelihood estimates by minimizing the Kullback-Leibler divergence. Their method yields a federated estimator that outperforms any other linear combination in various scenarios and is equivalent to the global MLE when the underlying distribution belongs to the full exponential family. Spath et al. [<xref ref-type="bibr" rid="B21">21</xref>] developed an open-source platform for federated analysis of time-to-event data that includes common methods like survival curves, the log-rank test and the Cox proportional hazards model. They found that their analyses lost little efficiency by comparison with fully aggregated analysis. Their methods and results are not directly relevant to the MIP, as they adopted differential privacy and additive secret sharing to protect local data rather than <italic>K</italic>-anonymity.</p>
<p>Related statistical literature is concerned with distributed computing, in which the data is centralized but so large that calculations are split over multiple servers in parallel to accelerate calculations. For example, Rosenblatt and Nadler [<xref ref-type="bibr" rid="B22">22</xref>] showed that the estimator from averaging estimates from <italic>m</italic> servers is as accurate as the centralized solution when the number of parameters <italic>p</italic> is fixed and the amount of data <italic>n</italic> &#x02192; &#x0221E;.</p></sec>
<sec id="s3">
<title>3. The binning algorithm</title>
<p>This section describes a procedure for constructing a <italic>K</italic>-anonymous federated summary table when two groups are compared with respect to a numerical variable. We denote the groups by <italic>x</italic> and <italic>y</italic> and use the terms control and treatment for them. The summary table will have <italic>B</italic> bins, with the <italic>b</italic>th bin given by (<italic>c</italic><sub><italic>b</italic>&#x02212;1</sub>, <italic>c</italic><sub><italic>b</italic></sub>], and observation frequencies <italic>f</italic><sub><italic>bx</italic></sub> for the control group and <italic>f</italic><sub><italic>by</italic></sub> for the treatment group. The table preserves <italic>K</italic>-anonymity in that it is constructed from frequency tables released from the centers in which all cell counts are either 0 or are &#x02265;<italic>K</italic>.</p>
<p>Here, is an outline of our table construction process. We proceed sequentially to add information from each center, beginning with the largest center and proceeding in decreasing order of sample size. The initial summary table meets the cell count constraint while attempting to minimize the width of the cells. Data from the other centers are then added, generating new bins if it is possible to do so without violating the privacy constraint. Existing bins are never removed. When cell counts from a new center are between 0 and <italic>K</italic>, neighboring bins are combined and their total count is redistributed among the bins that were combined (See <xref ref-type="table" rid="T4">Algorithm 1</xref> for details).</p>
<table-wrap position="float" id="T4">
<label>Algorithm 1</label>
<caption><p>Binning algorithm.</p></caption>
<table frame="hsides" rules="groups">
<tbody>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0;Input: x1, x2</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0;bins, frequencies1, frequencies2 = empty list</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0;</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0;while true <bold>do</bold></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0; next_point = next_2d_point(x1, x2)</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0; if next_point is None <bold>then</bold></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0; frequencies1[length(frequencies1)] &#x0002B;= length(x1)</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0; frequencies2[length(frequencies2)] &#x0002B;= length(x2)</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0; bins[length(bins)] = &#x0221E;</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0; return bins, frequencies1, frequencies2</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0; end <bold>if</bold></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0; f1 = length(x1[x1 &#x0003C; next_point])</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0; x1 = x1[x1 &#x02265; next_point]</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0; f2 = length(x2[x2 &#x0003C; next_point])</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0; x2 = x2[x2 &#x02265; next_point]</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0; bins.append(anonymize_boundary(next_point, x1, x2))</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0; frequencies1.append(f1)</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0; frequencies2.append(f2)</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0;end <bold>while</bold></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;&#x000A0;Output{bins, f1, f2}</td></tr>
</tbody>
</table>
</table-wrap>
<sec>
<title>3.1. Binning the largest center</title>
<p>The process proceeds (arbitrarily) from small to large values. The first bin is, initially, from <italic>a</italic><sub>0</sub> to <italic>a</italic><sub>1</sub>, where <italic>a</italic><sub>0</sub> is the minimal value in the data and <italic>a</italic><sub>1</sub>, is the smallest data value for which [<italic>a</italic><sub><italic>o</italic></sub>, <italic>a</italic><sub>1</sub>] has at least <italic>K</italic> observations from one group and either 0 or at least <italic>K</italic> observations from the other group. The next tentative bin limit, <italic>a</italic><sub>2</sub>, is found in the same way, looking at the interval (<italic>a</italic><sub>1</sub>, <italic>a</italic><sub>2</sub>]. This continues so long as a new bin limit can be found. When a limit cannot be found, the number of unbinned data in at least one group is between 0 and <italic>K</italic>. Tentatively extend the upper limit of the previously formed bin to the maximal value of this group as the next limit. The unbinned data from the other group might permit continuation of the process, blocking off new bins in which that group has counts of at least <italic>K</italic>, vs. counts of 0 for the first group. When that group has fewer than <italic>K</italic> unbinned data, replace the last bin limit by the maximal value in the second group (See Algorithm S1 in <xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref> for details.).</p>
<p>The initial bin boundaries <italic>a</italic><sub>0</sub>, &#x02026;, <italic>a</italic><sub><italic>B</italic></sub> produced by the algorithm above are actual data values and, unless many subjects share the same value, violate the privacy condition. There is a simple fix for <italic>a</italic><sub>1</sub>, &#x02026;, <italic>a</italic><sub><italic>B</italic>&#x02212;1</sub>. All values in the <italic>j</italic>th bin are &#x02264; <italic>a</italic><sub><italic>j</italic></sub> and all values in the <italic>j</italic>&#x0002B;1st bin are &#x0003E;<italic>a</italic><sub><italic>j</italic></sub>. So we can replace <italic>a</italic><sub><italic>j</italic></sub> by <italic>c</italic><sub><italic>j</italic></sub> &#x0003D; <italic>wa</italic><sub><italic>j</italic></sub>&#x0002B;(1&#x02212;<italic>w</italic>)<italic>v</italic><sub><italic>j</italic>&#x0002B;1</sub> where <italic>v</italic><sub><italic>j</italic>&#x0002B;1</sub> is the smallest value in the (<italic>j</italic>&#x0002B;1)st bin and <italic>w</italic> is a uniform random variable on (0, 1). The extreme boundaries <italic>a</italic><sub>0</sub> and <italic>a</italic><sub><italic>B</italic></sub> are the minimum and maximum in the data, so a different approach is needed. One option is to take <italic>c</italic><sub>0</sub> &#x0003D; &#x02212;&#x0221E; and <italic>c</italic><sub><italic>B</italic></sub> &#x0003D; &#x0221E;. Another option is to impose natural limits; for example, if by definition a variable cannot assume negative values, we could choose <italic>c</italic><sub>0</sub> &#x0003D; 0. A final option is to extend the bin limits by &#x0201C;privacy buffers&#x0201D;. To make these reasonably close to the data, we base them on the observed gaps between successive observations in the extreme bin. For example, compute <italic>c</italic><sub><italic>B</italic></sub> as <inline-formula><mml:math id="M1"><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, where <inline-formula><mml:math id="M2"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the mean difference between consecutive data points in the last bin. (If <inline-formula><mml:math id="M3"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>, <italic>c</italic><sub><italic>B</italic></sub> &#x0003D; <italic>a</italic><sub><italic>B</italic></sub>, but this is now privacy preserving, as all observations in the last bin are equal to one another, with more than <italic>K</italic> in each group that has data.) Similarly, compute <italic>c</italic><sub>0</sub> as <inline-formula><mml:math id="M4"><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
</sec>
<sec>
<title>3.2. Joining additional centers</title>
<p>A new algorithm is needed to add the data from a new center, preserving all bin boundaries from the first center. The simple option of increasing the frequency counts in each current bin is not an option, as the incremental table from the new center will typically not be <italic>K</italic>-anonymous. Further, the incremental counts for some existing bin might be so large that data from the new center could actually be used to split it into two or more bins.</p>
<p>Algorithm S3 in <xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref> is used to add the information from a new center to an existing summary table. We first iterate over the current bins, creating finer bins if possible. Then we remove any counts that are not <italic>K</italic>-anonymous by combining and redistributing data from adjacent cells. Pseudo-code for Algorithm S3 and for two algorithms called by it are given in the <xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref>.</p>
<p>Splitting an existing bin into two bins forces us to reallocate the previous frequencies. We do so proportionally to the relative frequencies from the new center. For example, suppose a bin with a current count of 27 for one group is split into two new bins, which have equal counts at the new center. Then we split the 27 equally to the two new groups, adding 13.5 to each. Note that this procedure can result in counts that are not integers.</p>
<p>After creating new bins wherever possible, we iterate again and fix bins where the new center has frequencies between 0 and <italic>K</italic>. Proceeding from bin 1 to bin <italic>B</italic>, these non-private bins are combined with the next bin to the right until all counts from the new center are either 0 or at least <italic>K</italic>. Then the total counts are distributed among the original bins proportionally to the relative frequencies of the bins in the current table. <xref ref-type="supplementary-material" rid="SM1">Supplementary Table 1</xref> shows an example that illustrates how the algorithm works.</p>
<p>The extreme bin limits <italic>c</italic><sub>0</sub> and <italic>c</italic><sub><italic>B</italic></sub> must be compared with the minimum and maximum values, respectively, in the new center. If the new center has a more extreme data value, we need to revise these bin limits. We do so by applying the buffer method that was used to find <italic>c</italic><sub>0</sub> and <italic>c</italic><sub><italic>B</italic></sub> in the largest center, but now adding buffers that depend only on the data in the extreme bin from the new center.</p></sec></sec>
<sec id="s4">
<title>4. Testing</title>
<p>This section considers the problem of hypothesis testing with federated data, studying the common problem of determining whether numerical outcomes from two groups come from the same distribution (the null hypothesis, <italic>H</italic><sub>0</sub>); or whether one group has larger values than the other. The standard choice is the independent samples <italic>t</italic>-test, which requires the mean, the standard deviation and the number of observations in each group. All of these are privacy-preserving summary statistics, so the <italic>t</italic>-test can still be used with federated data. However, the <italic>t</italic>-test relies on the assumption, often invalid, that the data are normally distributed. We consider here the standard non-parametric alternative, the Mann-Whitney U (&#x0201C;MWU&#x0201D;) test [<xref ref-type="bibr" rid="B9">9</xref>] (or, equivalently, the Wilcoxon rank sum test).</p>
<sec>
<title>4.1. The Mann-Whitney <italic>U</italic>-test</title>
<p>The MWU statistic can be defined as follows. Denoting the observations in the two groups by <italic>X</italic><sub>1</sub>, &#x02026;, <italic>X</italic><sub><italic>n</italic></sub> and <italic>Y</italic><sub>1</sub>, &#x02026;, <italic>Y</italic><sub><italic>m</italic></sub>,</p>
<disp-formula id="E1"><mml:math id="M5"><mml:mrow><mml:mi>U</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>S</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>Y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>with</p>
<disp-formula id="E2"><mml:math id="M6"><mml:mrow><mml:mi>S</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mn>1</mml:mn></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mi>Y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x0003E;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mn>0</mml:mn></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mi>Y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mi>Y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x0003C;</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>If <italic>H</italic><sub>0</sub> is true, the expected value of <italic>U</italic> is 0 and its variance is <inline-formula><mml:math id="M7"><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>m</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:mfrac><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>N</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, where <italic>N</italic> &#x0003D; <italic>n</italic>&#x0002B;<italic>m</italic>, <italic>D</italic> is the number of distinct values in the data, and <italic>t</italic><sub><italic>r</italic></sub> is the number of observations that share the <italic>r</italic>th distinct value. The second term corrects the variance for the presence of ties in the data. If <inline-formula><mml:math id="M8"><mml:mi>Y</mml:mi><mml:mover class="stackrel"><mml:mrow><mml:mo>=</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:mrow></mml:mover><mml:mi>c</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>X</mml:mi><mml:mo>,</mml:mo></mml:math></inline-formula> <italic>c</italic>&#x02208;&#x0211D;, the distribution of <italic>U</italic> is stochastically increasing as a function of <italic>c</italic>. The power of the test depends on <italic>P</italic>(<italic>Y</italic>&#x0003E;<italic>X</italic>) and is high when this probability differs from 0.5.</p>
<p>The MWU test involves direct comparison of each data point in one group with each data point from the other group. As this includes comparisons of observations from different centers, it is impossible to compute the MWU statistic for a federated analysis. Two broad options are possible for federated analysis.</p>
<list list-type="bullet">
<list-item><p>Compute the MWU statistic separately for each center and then combine them across centers.</p></list-item>
<list-item><p>Generate a federated table summarizing the data from all the centers and then compute the MWU statistic on the federated table.</p></list-item>
</list>
<p>The next subsections present options for combining center-specific MWU statistics and the second analysis option, used in conjunction with our federated binning algorithm.</p></sec>
<sec>
<title>4.2. Sum of <italic>U</italic>-statistics</title>
<p>Denote by <italic>U</italic><sub><italic>l</italic></sub> the MWU from the <italic>l</italic>th center, based on <italic>n</italic><sub><italic>l</italic></sub> and <italic>m</italic><sub><italic>l</italic></sub> observations from the two groups, with <italic>N</italic><sub><italic>l</italic></sub> &#x0003D; <italic>n</italic><sub><italic>l</italic></sub>&#x0002B;<italic>m</italic><sub><italic>l</italic></sub>; and denote by <italic>V</italic><sub><italic>l</italic></sub> its variance under <italic>H</italic><sub>0</sub>. A simple way to form a federated test statistic is to sum the individual statistics over the centers and normalize them by their standard deviation, leading to</p>
<disp-formula id="E3"><label>(1)</label><mml:math id="M9"><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>L</mml:mi></mml:msubsup><mml:mrow><mml:msub><mml:mi>U</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:mstyle></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mstyle displaystyle='true'><mml:msubsup><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>L</mml:mi></mml:msubsup><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>0.5</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mover><mml:mo>&#x02192;</mml:mo><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:mrow></mml:mover><mml:mi>N</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></sec>
<sec>
<title>4.3. Weighted average of <italic>U</italic>-statistics</title>
<p>A simple generalization is to replace the sum of the statistics by a weighted sum, with an optimal choice of weights. It is convenient to do this using the normalized test statistics for each center, <inline-formula><mml:math id="M10"><mml:msub><mml:mrow><mml:mi>Z</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>U</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>/</mml:mo><mml:msubsup><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula>. The weighted test statistic is then</p>
<disp-formula id="E4"><label>(2)</label><mml:math id="M11"><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>w</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle='true'><mml:msub><mml:mo>&#x02211;</mml:mo><mml:mi>l</mml:mi></mml:msub><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:msub><mml:mi>Z</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow></mml:mstyle></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mstyle displaystyle='true'><mml:msub><mml:mo>&#x02211;</mml:mo><mml:mi>l</mml:mi></mml:msub><mml:mrow><mml:msubsup><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:mrow></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>0.5</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mover><mml:mo>&#x02192;</mml:mo><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:mrow></mml:mover><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
<p>The choice of weights can be made to maximize the power of the test when the null hypothesis is not true, using the fact that</p>
<disp-formula id="E5"><mml:math id="M12"><mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>w</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle='true'><mml:msub><mml:mo>&#x02211;</mml:mo><mml:mi>l</mml:mi></mml:msub><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:msub><mml:mi>Z</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow></mml:mstyle></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mstyle displaystyle='true'><mml:msub><mml:mo>&#x02211;</mml:mo><mml:mi>l</mml:mi></mml:msub><mml:mrow><mml:msubsup><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:mrow></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>.5</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mover><mml:mo>&#x02192;</mml:mo><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:mover><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mstyle displaystyle='true'><mml:msub><mml:mo>&#x02211;</mml:mo><mml:mi>l</mml:mi></mml:msub><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:msub><mml:mi>&#x003B4;</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow></mml:mstyle></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mstyle displaystyle='true'><mml:msub><mml:mo>&#x02211;</mml:mo><mml:mi>l</mml:mi></mml:msub><mml:mrow><mml:msubsup><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:mrow></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>.5</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where &#x003B4;<sub><italic>l</italic></sub> is the standardized effect in center <italic>l</italic>. For the MWU statistic, the standardized effect can be expressed as</p>
<disp-formula id="E6"><mml:math id="M13"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>E</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>Z</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>S</mml:mi><mml:mi>D</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>U</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M14"><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002B;</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>&#x0003E;</mml:mo><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M15"><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> in center <italic>l</italic>. Although the formula permits the probability difference to vary over centers, the natural basis for defining the weighted sum statistic is to assume a constant difference, in which case the optimal weights depend on the sample sizes and, if present, the extent of tied data. See Equation S1 in the <xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref> for derivation of the weights.</p></sec>
<sec>
<title>4.4. Fisher&#x00027;s method</title>
<p>Fisher&#x00027;s method [<xref ref-type="bibr" rid="B23">23</xref>] combines the <italic>p</italic>-values from independent samples. The corresponding statistic is <inline-formula><mml:math id="M16"><mml:msub><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>F</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mn>2</mml:mn><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:munderover><mml:mo class="qopname">log</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mover class="stackrel"><mml:mrow><mml:mo>&#x02192;</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:mover><mml:msubsup><mml:mrow><mml:mi>&#x003C7;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> where <italic>p</italic><sub><italic>l</italic></sub> is the <italic>p</italic>-value from the MWU test result in the <italic>l</italic>th center.</p></sec>
<sec>
<title>4.5. Federated table MWU statistic</title>
<p>We can compute the MWU statistic from the federated summary table generated by the algorithm described in Section 3. The table will have <italic>B</italic> bins whose frequencies are <italic>fx</italic><sub><italic>i</italic></sub> and <italic>fy</italic><sub><italic>i</italic></sub>. The frequencies sum to the total amount of data over all the centers, but need not be integers.</p>
<p>The MWU statistic for the federated table compares observations on the basis of their bins and is given by</p>
<disp-formula id="E7"><label>(3)</label><mml:math id="M17"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>U</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>f</mml:mi><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>f</mml:mi><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>c</italic><sub>0</sub>&#x0003C;<italic>c</italic><sub>1</sub> &#x0003C; &#x022EF; &#x0003C; <italic>c</italic><sub><italic>B</italic></sub> are the endpoints of the bins and</p>
<disp-formula id="E8"><mml:math id="M18"><mml:mrow><mml:mi>S</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mn>1</mml:mn></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x0003E;</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mn>0</mml:mn></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x0003C;</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>The variance of <italic>U</italic><sub><italic>fed</italic></sub> can be computed from the formula in Section 4.1, keeping in mind that all observations in the same bin are tied.</p></sec>
<sec>
<title>4.6. Comparison of the tests</title>
<p>A simulation study was used to compare the different federated MWU tests to an analysis of the combined data. Our goals are to assess how the federated analysis affects the power of the tests, and to use the power analysis to compare the testing methods. We also vary the simulation settings to examine how the results and comparisons are affected by the number of centers in the study and by heterogeneity across centers.</p>
<p>We simulated situations with 1,500 observations in each group, divided over 3, 5, or 10 centers, with the number of observations unbalanced among the centers (see <xref ref-type="table" rid="T1">Table 1</xref>).</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Observations per center.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Number of centers</bold></th>
<th valign="top" align="left"><bold>Number of observations in each group</bold></th>
<th valign="top" align="left"><bold>Each group total</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">3</td>
<td valign="top" align="left">698, 476, 326</td>
<td valign="top" align="left">1,500</td>
</tr> <tr>
<td valign="top" align="left">5</td>
<td valign="top" align="left">492, 368, 276, 208, 156</td>
<td valign="top" align="left">1,500</td>
</tr>
<tr>
<td valign="top" align="left">10</td>
<td valign="top" align="left">307, 250, 208, 172, 143, 118, 98, 81, 67, 56</td>
<td valign="top" align="left">1,500</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>For the Mann-Whitney test, only the order of the observations is important, so any distribution can be used to simulate the data. Our model generates control group observations at center <italic>l</italic> as</p>
<disp-formula id="E9"><mml:math id="M19"><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></disp-formula>
<p>and treatment group observations as</p>
<disp-formula id="E10"><mml:math id="M20"><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
<p>The possibility that centers may differ from one another is represented by <inline-formula><mml:math id="M21"><mml:msub><mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0007E;</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>. The difference between treatment and control at center <italic>l</italic> is <inline-formula><mml:math id="M22"><mml:msub><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0007E;</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x003B4;</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> where &#x003B4; is the overall difference, and &#x003C3;<sub>&#x003B2;</sub> represents heterogeneity of the treatment effect across centers. The terms &#x003F5;<sub><italic>il</italic></sub>, &#x003F5;<sub><italic>jl</italic></sub>&#x0007E;<italic>N</italic>(0, 1) are random errors. All random variables are independent of one another.</p>
<p>We simulated experiments with several different combinations of input parameters. We chose &#x003C3;<sub>&#x003B1;</sub>&#x02208;{0, 0.1, 0.2} and &#x003C3;<sub>&#x003B2;</sub>&#x02208;{0, 0.05, 0.06} to achieve between center variance, and &#x003B4;&#x02208;{0, 0.05, 0.1}. Including &#x003B4; &#x0003D; 0 allowed us to verify that the tests remain reliable when both groups have the same mean. Note, however, that the variance is slightly larger for the treatment group if &#x003C3;<sub>&#x003B2;</sub>&#x0003E;0, so that this setting does not fully match the null hypothesis of identical distributions.</p>
<p><xref ref-type="supplementary-material" rid="SM1">Supplementary Figure 1</xref> shows the distributions of <italic>p</italic>-values for all the tests in the null setting &#x003B4; &#x0003D; 0. The left panel includes heterogeneity across centers (&#x003C3;<sub>&#x003B1;</sub> &#x0003D; 0.1), but no effect heterogeneity, and shows a uniform distribution for all the tests, as desired. The right panel adds a small amount of effect heterogeneity (&#x003C3;<sub>&#x003B2;</sub> &#x0003D; 0.05. This results in a slightly wider spread of <italic>p</italic>-values for all the tests, so that actual type 1 errors are inflated from their nominal values. The fraction of <italic>p</italic>-values below 0.05 (0.01) was approximately 0.08 (0.025). The inflation was slightly weaker when more centers were included and slightly larger only for Fisher&#x00027;s test. The additional bias of Fisher&#x00027;s test is not surprising, as it is sensitive to the existence of an effect within a center, but not to having a consistent direction of the effect.</p>
<p><xref ref-type="fig" rid="F1">Figure 1</xref> compares methods when &#x003B4;&#x02260;0 across different parameters and numbers of centers. See also <xref ref-type="supplementary-material" rid="SM1">Supplementary Table S1</xref>. The federated table and weighted tests have <italic>p</italic>-value distributions that are very similar to those from combining all the data, indicating almost no loss of power. The sum test has higher <italic>p</italic>-values, hence consistently lower power. The <italic>p</italic>-values with Fisher&#x00027;s method are a bit higher when the treatment effect is consistent across centers (&#x003C3;<sub>&#x003B2;</sub> &#x0003D; 0). When the effect is not consistent, they are lower. However, as already seen, Fisher&#x00027;s test in this case fails to preserve type 1 error, with a bias toward low values.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Comparison of methods across different parameters and number of centers. Panels represent the number of centers, the <italic>Y</italic>-axis presents the <italic>p</italic>-values and the <italic>X</italic>-axis the parameters (&#x003B4;, &#x003C3;<sub>&#x003B1;</sub>, &#x003C3;<sub>&#x003B2;</sub>). The different methods are color-coded.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fams-09-1267034-g0001.tif"/>
</fig>
<p><xref ref-type="fig" rid="F2">Figure 2</xref> focuses on how closely the federated test results compare with those from the combined test (i.e., using the full data) by comparing the <italic>p</italic>-values of each method on the same simulated data set. The <italic>Y</italic> axis presents <italic>log</italic>(<italic>p</italic><sub><italic>iv</italic></sub>/<italic>p</italic><sub><italic>is</italic></sub>) where <italic>i</italic> represents the simulation number, <italic>s</italic> is the combined test and <italic>v</italic> is the federated test. A federated test that produces the same <italic>p</italic>-values as the combined test has no loss of power. The more tightly concentrated are these distributions around 0, the more nearly identical are the <italic>p</italic>-values of the federated method to those of the combined method.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>The figure shows log(<italic>p</italic><sub><italic>iv</italic></sub>/<italic>p</italic><sub><italic>is</italic></sub>) on the <italic>Y</italic> axis, where <italic>s</italic> is the combined data analysis. The panels correspond to the different parameter settings for (&#x003B4;, &#x003C3;<sub>&#x003B1;</sub>, &#x003C3;<sub>&#x003B2;</sub>). The number of centers is on the <italic>X</italic> axis and the methods are color-coded.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fams-09-1267034-g0002.tif"/>
</fig>
<p>Across all the settings, the weighted test most closely replicates the <italic>p</italic>-value of the combined test. The federated table is also similar, but more variable, especially when &#x003B4;&#x02260;0. In the top left panel, where <italic>H</italic><sub>0</sub> is true, all methods are similar to the combined test. However, adding treatment heterogeneity (top right panel) induces bias in the <italic>p</italic>-values from Fisher&#x00027;s test, which will reject the null hypothesis too often. It also increases the variance of the log ratio for that test and for the sum. In all the settings with center heterogeneity (&#x003C3;<sub>&#x003B1;</sub>&#x0003E;0), the sum test gave, typically, slightly higher <italic>p</italic>-values than the combined test, hence had lower power.</p>
<p>To assess the power of the tests as a function of the effect size, we simulated <italic>p</italic>-values over a set of 4 increasing values of &#x003B4;, when &#x003C3;<sub>&#x003B1;</sub> &#x0003D; 0.1 and &#x003C3;<sub>&#x003B2;</sub> &#x0003D; 0.05. <xref ref-type="fig" rid="F3">Figure 3</xref> compares the methods to the unconstrained test using log(<italic>p</italic><sub><italic>iv</italic></sub>/<italic>p</italic><sub><italic>is</italic></sub>) (<italic>Y</italic>-axis) where <italic>i</italic> represents the simulation number, <italic>s</italic> is the unconstrained method and <italic>v</italic> is the other method. Again the weighted test is most similar to the combined test, followed by the federated table. <xref ref-type="table" rid="T2">Table 2</xref> shows the median of the <italic>p</italic>-value distributions with 10 centers; smaller medians correspond to higher power for the test. The medians for the weighted test are consistently the lowest ones; with even the modest heterogeneity present here, they are lower even than those for the combined test. The test from our federated table has slightly higher medians throughout than does the combined test. Similar quantiles were found for 3 and for 5 centers, indicating that, for the settings we examined, the number of centers has little effect on power.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Comparison of methods across different parameters and number of centers. Rows correspond to the number of centers and segments within rows represent hyper-parameter configurations. The <italic>Y</italic>-axis represents the <italic>p</italic>-values. The methods are color-coded.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fams-09-1267034-g0003.tif"/>
</fig>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Comparison of the tests for different values of &#x003B4; with 10 centers when &#x003C3;<sub>&#x003B1;</sub> &#x0003D;.1, &#x003C3;<sub>&#x003B2;</sub> &#x0003D;.05, and the 0.5 quantile of the <italic>p</italic>-value distributions.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Quantile</bold></th>
<th valign="top" align="left" colspan="5"><bold>0.5</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<td valign="top" align="left">&#x003B4;</td>
<td valign="top" align="left"><bold>0.05</bold></td>
<td valign="top" align="left"><bold>0.0625</bold></td>
<td valign="top" align="left"><bold>0.075</bold></td>
<td valign="top" align="left"><bold>0.0875</bold></td>
<td valign="top" align="left"><bold>0.1</bold></td>
</tr>
<tr style="background-color:#dee1e1;color:&#x00023;ffffff">
<td valign="top" align="left" colspan="6"><bold>method</bold></td>
</tr> <tr>
<td valign="top" align="left">Combined</td>
<td valign="top" align="left">0.0921</td>
<td valign="top" align="left">0.0514</td>
<td valign="top" align="left">0.0228</td>
<td valign="top" align="left">0.0093</td>
<td valign="top" align="left">0.0037</td>
</tr> <tr>
<td valign="top" align="left">Federated</td>
<td valign="top" align="left">0.0949</td>
<td valign="top" align="left">0.0538</td>
<td valign="top" align="left">0.0234</td>
<td valign="top" align="left">0.0091</td>
<td valign="top" align="left">0.0038</td>
</tr> <tr>
<td valign="top" align="left">Fisher</td>
<td valign="top" align="left">0.0917</td>
<td valign="top" align="left">0.0536</td>
<td valign="top" align="left">0.0274</td>
<td valign="top" align="left">0.0119</td>
<td valign="top" align="left">0.0049</td>
</tr> <tr>
<td valign="top" align="left">Sum</td>
<td valign="top" align="left">0.1050</td>
<td valign="top" align="left">0.0646</td>
<td valign="top" align="left">0.0304</td>
<td valign="top" align="left">0.0157</td>
<td valign="top" align="left">0.0066</td>
</tr>
<tr>
<td valign="top" align="left">Weighted</td>
<td valign="top" align="left">0.0908</td>
<td valign="top" align="left">0.0510</td>
<td valign="top" align="left">0.0224</td>
<td valign="top" align="left">0.0090</td>
<td valign="top" align="left">0.0036</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Similar results were found for 3 and for 5 centers.</p>
</table-wrap-foot>
</table-wrap></sec></sec>
<sec id="s5">
<title>5. Estimation</title>
<p>This section considers the problem of quantile estimation when data are located in different centers. Quantile estimates are valuable for directing visual summaries of data distributions such as histograms or Kaplan-Meier plots. Standard methods for computing sample quantiles cannot be used, as they begin by ordering all the data, violating privacy. We propose and compare several methods for federated quantile estimation. Throughout we denote by <italic>F</italic>(<italic>x</italic>) the CDF and by <inline-formula><mml:math id="M23"><mml:msub><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> the <italic>p</italic>th quantile of the distribution.</p>
<sec>
<title>5.1. Federated estimates using the quantile loss</title>
<p>A quantile can be estimated as the solution to a minimization problem,</p>
<disp-formula id="E11"><label>(4)</label><mml:math id="M24"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo class="qopname">arg</mml:mo><mml:mo class="qopname">min</mml:mo></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0003C;</mml:mo><mml:mi>q</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi>q</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>p</mml:mi><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02265;</mml:mo><mml:mi>q</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi>q</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where the target function is the <italic>quantile loss function</italic>. The optimization can be carried out on federated data by returning function and gradient values from each center, proceeding iteratively to compute <inline-formula><mml:math id="M25"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. The need for an iterative algorithm to minimize the loss, has the drawback of communication inefficiency.</p>
<p>A more serious concern is that the quantile loss compromises privacy. The loss function within each center is piecewise linear with a change in derivative at each data value in the center. Thus the information from a collection of calls can be used to recover the original data values at the federated node.</p>
<p>Despite the privacy violation, we will include <inline-formula><mml:math id="M26"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> in the subsequent comparisons as a benchmark.</p>
<p>It is possible to exploit the loss function to compute approximate quantile estimators that are differentially private [<xref ref-type="bibr" rid="B24">24</xref>].</p></sec>
<sec>
<title>5.2. Estimating quantiles from the federated data using the Yeo-Johnson transformation</title>
<p>The binning algorithm we introduced in Section 3 can be used to compute a federated estimate of <italic>Q</italic><sub><italic>p</italic></sub> that is <italic>K</italic>-anonymous. Here, we apply the single group version of the algorithm which gives a summary table that has <italic>B</italic> bins with endpoints <italic>b</italic><sub>0</sub>&#x0003C;<italic>b</italic><sub>1</sub>&#x0003C;<italic>b</italic><sub>2</sub> &#x0003C; &#x022EF; &#x0003C; <italic>b</italic><sub><italic>B</italic></sub> and frequencies <italic>f</italic><sub><italic>x, k</italic></sub>. Let <inline-formula><mml:math id="M27"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denote the cumulative distribution for the federated table at <italic>b</italic><sub><italic>i</italic></sub>.</p>
<p>A naive estimate is the smallest bin limit with cumulative frequency greater than 100<italic>p%</italic> of the data. However, restricting <italic>Q</italic><sub><italic>p</italic></sub> to the set of bin limits is an obvious drawback, especially for quantiles in the tails of the distribution. A simple improvement is to interpolate the estimated CDF from one bin limit to the next. Linear interpolation corresponds to the assumption of a uniform distribution within each bin. That may be reasonable for bins in the center of the data. However, it is not likely to work well in the tails, especially in the most extreme bins. We did attempt to use linear interpolation, but the results were poor and are not reported here.</p>
<p>We propose here a more sophisticated interpolation method based on the Yeo-Johnson transformation (&#x0201C;YJ&#x0201D;) [<xref ref-type="bibr" rid="B25">25</xref>], a power transformation used to achieve a distribution that is closer to the normal. The approach extends the well-known Box-Cox [<xref ref-type="bibr" rid="B26">26</xref>] transformation to also handle variables that can take on negative values. The transformation is defined by</p>
<disp-formula id="E12"><label>(5)</label><mml:math id="M28"><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi>&#x003BB;</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:msup><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:mi>&#x003BB;</mml:mi></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:mtext>&#x000A0;if&#x000A0;</mml:mtext><mml:mi>&#x003BB;</mml:mi><mml:mo>&#x02260;</mml:mo><mml:mn>0</mml:mn><mml:mtext>&#x000A0;and&#x000A0;</mml:mtext><mml:mi>x</mml:mi><mml:mo>&#x02265;</mml:mo><mml:mn>0</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mi>log</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:mtext>&#x000A0;if&#x000A0;</mml:mtext><mml:mi>&#x003BB;</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mtext>&#x000A0;and&#x000A0;</mml:mtext><mml:mi>x</mml:mi><mml:mo>&#x02265;</mml:mo><mml:mn>0</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003BB;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003BB;</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:mtext>&#x000A0;if&#x000A0;</mml:mtext><mml:mi>&#x003BB;</mml:mi><mml:mo>&#x02260;</mml:mo><mml:mn>2</mml:mn><mml:mtext>&#x000A0;and&#x000A0;</mml:mtext><mml:mi>x</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mn>0</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mi>log</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:mtext>&#x000A0;if&#x000A0;</mml:mtext><mml:mi>&#x003BB;</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mtext>&#x000A0;and&#x000A0;</mml:mtext><mml:mi>x</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mn>0</mml:mn></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<sec>
<title>5.2.1. YJ table method</title>
<p>In this method, the goal is to find values of &#x003BB;, <italic>a</italic><sub>0</sub> and <italic>a</italic><sub>1</sub> for which the transformed bin limits approximately match a normal distribution with mean <italic>a</italic><sub>0</sub> and standard deviation <italic>a</italic><sub>1</sub>,</p>
<disp-formula id="E13"><label>(6)</label><mml:math id="M29"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mi>&#x003A6;</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>b</italic><sub><italic>k</italic></sub> is a bin limit, <inline-formula><mml:math id="M30"><mml:mover accent="true"><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula> is the estimator of the distribution function from the federated table and <italic>h</italic><sub>&#x003BB;</sub>(<italic>x</italic>) is the (&#x0201C;YJ&#x0201D;) [<xref ref-type="bibr" rid="B25">25</xref>] transformation. The quantile <italic>Q</italic><sub><italic>p</italic></sub> is then estimated by</p>
<disp-formula id="E14"><label>(7)</label><mml:math id="M31"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>Y</mml:mi><mml:mi>J</mml:mi><mml:mi>T</mml:mi><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x000E2;</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x000E2;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mi>&#x003A6;</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Given &#x003BB;, we can compute <italic>a</italic><sub>0</sub>, <italic>a</italic><sub>1</sub> using linear regression. To estimate &#x003BB;, we use the idea that an effective transformation <italic>h</italic><sub>&#x003BB;</sub> should have transformed quantiles that are linearly related to the YJ-estimated quantiles. This can be achieved by choosing &#x003BB; to maximize the correlation between them,</p>
<disp-formula id="E15"><label>(8)</label><mml:math id="M32"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder><mml:mrow><mml:mo class="qopname">argmax</mml:mo></mml:mrow><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo class="qopname">cor</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003A6;</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mo class="qopname">^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where the values of <italic>X</italic> we use are the interior bin limits <italic>b</italic><sub>1</sub>, &#x02026;, <italic>b</italic><sub><italic>B</italic>&#x02212;1</sub>.</p>
<p>Note that the range of the inverse transformation in 5 is given by</p>
<disp-formula id="E16"><label>(9)</label><mml:math id="M33"><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi>&#x003BB;</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>&#x0211D;</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mo>&#x0007C;</mml:mo><mml:mi>&#x003BB;</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x0007C;</mml:mo><mml:mo>,</mml:mo><mml:mi>&#x0221E;</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:mtext>&#x000A0;if&#x000A0;</mml:mtext><mml:mi>&#x003BB;</mml:mi><mml:mo>&#x0003E;</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mi>&#x0211D;</mml:mi></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:mtext>&#x000A0;if&#x000A0;</mml:mtext><mml:mn>0</mml:mn><mml:mo>&#x02264;</mml:mo><mml:mi>&#x003BB;</mml:mi><mml:mo>&#x02264;</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mi>&#x0221E;</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mo>&#x0007C;</mml:mo><mml:mi>&#x003BB;</mml:mi><mml:mo>&#x0007C;</mml:mo><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:mtext>&#x000A0;if&#x000A0;</mml:mtext><mml:mi>&#x003BB;</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mn>0</mml:mn></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>To ensure that the inverse transformation has values in &#x0211D; we set the constraint 0 &#x02264; &#x003BB; &#x02264; 2 for equation 8.</p>
<p>The YJ Table method is a &#x0201C;one pass&#x0201D; algorithm, calling the data only to produce the federated summary table. Thus it enjoys full communication efficiency.</p></sec>
<sec>
<title>5.2.2. YJ likelihood method</title>
<p>The parameters in the YJ transformation can also be estimated by maximum likelihood. Denoting by <italic>x</italic><sub><italic>il</italic></sub> the observations from center <italic>l</italic> and by N the total number of observations, the log likelihood is</p>
<disp-formula id="E17"><label>(10)</label><mml:math id="M36"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mo>-</mml:mo><mml:mi>N</mml:mi><mml:mo>/</mml:mo><mml:mn>2</mml:mn><mml:mo class="qopname">log</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo class="qopname">^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x003BB;</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munderover></mml:mstyle><mml:mo class="qopname">sign</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo class="qopname">log</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where</p>
<disp-formula id="E18"><label>(11)</label><mml:math id="M37"><mml:mrow><mml:msubsup><mml:mover accent='true'><mml:mi>&#x003C3;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mi>&#x003BB;</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>L</mml:mi></mml:munderover><mml:mrow><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi>&#x003BB;</mml:mi></mml:msub></mml:mrow></mml:mstyle></mml:mrow></mml:mstyle><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>L</mml:mi></mml:munderover><mml:mrow><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi>&#x003BB;</mml:mi></mml:msub></mml:mrow></mml:mstyle></mml:mrow></mml:mstyle><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
<p>For a fixed value of &#x003BB;, the log-likelihood requires only summary statistics from each center, so can be computed in a federated manner. This can be embedded in a simple optimization routine that maximizes the log likelihood over &#x003BB;.</p>
<p>As with the quantile loss, the YJ likelihood method employs an iterative algorithm, and thus is not communication efficient. However, unlike the quantile loss, the YJ log likelihood for each center is not a simple function of the data that can be immediately inverted to recover data values. Thus the privacy violations of the quantile loss do not occur here.</p>
<p>Once we have <inline-formula><mml:math id="M40"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula>, we can again use summary statistics from the centers to compute <inline-formula><mml:math id="M41"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow></mml:msub></mml:math></inline-formula>. The resulting quantile estimator is</p>
<disp-formula id="E19"><label>(12)</label><mml:math id="M42"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>Y</mml:mi><mml:mi>J</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mi>D</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mi>&#x003A6;</mml:mi></mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The likelihood maximization is iterative, so requires multiple communication steps with each center. By contrast, the methods based on the federated table are &#x0201C;one pass&#x0201D;, requiring just one call to each center. This communication inefficiency of the maximum likelihood method can be improved by submitting to each center a grid of possible &#x003BB; values. The centers then return the moments needed to compute the log likelihood for each value in the grid. The resulting estimate of &#x003BB; can either be the best value among those in the grid or the maximizer of an empirical fit to the relationship between the log likelihood and &#x003BB;. The result is an approximate, one pass MLE.</p></sec></sec>
<sec>
<title>5.3. Constructing summary tables from quantile estimates</title>
<p>Federated quantile estimates can be used to generate an alternative summary table, which presents a collection of quantiles. See <xref ref-type="table" rid="T3">Table 3</xref> for an example, with estimates from optimizing the quantile loss and the <italic>YJ</italic> likelihood.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>A summary table constructed from quantile estimates, based on 1,500 observations from 3 centers, with a mixed gamma distribution.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold><italic>p</italic></bold></th>
<th valign="top" align="left"><bold><italic>Q</italic><sub><italic>p</italic></sub></bold></th>
<th valign="top" align="left"><bold><inline-formula><mml:math id="M34"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></bold></th>
<th valign="top" align="left"><bold><inline-formula><mml:math id="M35"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mi>Y</mml:mi><mml:mi>J</mml:mi><mml:mi>D</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">0.02</td>
<td valign="top" align="left">1.027</td>
<td valign="top" align="left">1.137</td>
<td valign="top" align="left">1.076</td>
</tr> <tr>
<td valign="top" align="left">0.25</td>
<td valign="top" align="left">2.565</td>
<td valign="top" align="left">2.551</td>
<td valign="top" align="left">2.595</td>
</tr> <tr>
<td valign="top" align="left">0.50</td>
<td valign="top" align="left">3.716</td>
<td valign="top" align="left">3.700</td>
<td valign="top" align="left">3.695</td>
</tr> <tr>
<td valign="top" align="left">0.75</td>
<td valign="top" align="left">5.175</td>
<td valign="top" align="left">5.155</td>
<td valign="top" align="left">5.125</td>
</tr>
<tr>
<td valign="top" align="left">0.98</td>
<td valign="top" align="left">9.216</td>
<td valign="top" align="left">9.321</td>
<td valign="top" align="left">9.512</td>
</tr>
</tbody>
</table>
</table-wrap></sec>
<sec>
<title>5.4. Simulation results</title>
<p>We compared the three quantile estimators using a simulation configuration similar to that in the testing chapter. As the quantiles are univariate summaries, we generated data and estimated quantiles only in one group. Another important difference is that the form of the underlying distribution affects the estimation results. In particular, methods may vary when faced with long rather than short tails. To gain insight into this issue, we chose the Gamma as the base distribution for assessing the quality of quantile estimation.</p>
<p>Each simulated data set included 1,500 observations, spread across 3, 5, or 10 centers exactly as described in <xref ref-type="table" rid="T1">Table 1</xref>. The observations were generated from the following model: <italic>x</italic><sub><italic>il</italic></sub> &#x0003D; &#x003F5;<sub><italic>il</italic></sub>exp(&#x003B1;<sub><italic>l</italic></sub>) where <italic>x</italic><sub><italic>il</italic></sub> is observation <italic>i</italic> at center <italic>l</italic> with <inline-formula><mml:math id="M43"><mml:msub><mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0007E;</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> and &#x003F5;<sub><italic>il</italic></sub>&#x0007E;<italic>Gamma</italic>(<italic>r</italic>, 1) <italic>r</italic>&#x02208;{4, 10}. The skewness of Gamma is <inline-formula><mml:math id="M44"><mml:mfrac><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:math></inline-formula>, so the smaller value for <italic>r</italic> has a longer right tail.</p>
<p>For the Gamma data, heterogeneity across centers was induced using scale rather than location shifts. The value of &#x003C3;<sub>&#x003B1;</sub> was chosen to achieve between center heterogeneity similar in extent to that in Section 4. There the key term was the ratio &#x003C3;<sub>&#x003B1;</sub>/&#x003C3;<sub>&#x003F5;</sub>, which was taken to be 0, 0.1 or 0.2. With Gamma data, the standard deviation of the homogeneous data is proportional to the median, so the analogous choice is to set <inline-formula><mml:math id="M45"><mml:msub><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo class="qopname">log</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003D5;</mml:mi><mml:msub><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>5</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, with &#x003D5; similar to the values chosen above. We used only &#x003D5; &#x0003D; 0.1 in our simulations for quantile estimation.</p>
<p>For each combination of the parameters, 2,000 simulations were run. The true quantile <italic>Q</italic><sub><italic>p</italic></sub> for each simulation was computed from the mixture (over centers) distribution by solving the equation below with <italic>l</italic> as the center index.</p>
<disp-formula id="E20"><label>(13)</label><mml:math id="M46"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:msub><mml:mrow><mml:mi>&#x00393;</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>/</mml:mo><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>p</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>N</italic> is the number of observations from all centers, <italic>n</italic><sub><italic>l</italic></sub> the observations in center <italic>l</italic> and &#x00393;<sub><italic>r</italic></sub> is the standard Gamma CDF with shape parameter <italic>r</italic>. A dominant part of the quantile estimation errors is the natural variability of the underlying Gamma distribution. As the standard deviation for <italic>Gamma</italic>(<italic>r</italic>, 1) is <inline-formula><mml:math id="M47"><mml:msqrt><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msqrt></mml:math></inline-formula>, we summarized results via the normalized estimation error <inline-formula><mml:math id="M48"><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:msqrt><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msqrt></mml:math></inline-formula> where <inline-formula><mml:math id="M49"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the estimator of <italic>Q</italic><sub><italic>p</italic></sub>.</p>
<p>The simulation results for estimating <italic>Q</italic><sub>0.98</sub> are shown in <xref ref-type="fig" rid="F4">Figure 4</xref>. This quantile is presented separately, as it is the most challenging case, in the right tail of a right-skewed distribution. Results for <italic>Q</italic><sub>0.02</sub>,<italic>Q</italic><sub>0.25</sub>,<italic>Q</italic><sub>0.5</sub>, <italic>Q</italic><sub>0.75</sub> are depicted in <xref ref-type="fig" rid="F5">Figure 5</xref>. Further detail is provided in <xref ref-type="supplementary-material" rid="SM1">Supplementary Tables S2</xref>&#x02013;<xref ref-type="supplementary-material" rid="SM1">S5</xref>, which give, respectively, the estimated bias and standard deviation, the mean squared error (MSE), and the ratio of squared bias to variance for all the methods and all the quantiles.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Estimation errors - <italic>Q</italic><sub>0.98</sub>. The figure shows the standardized errors, <inline-formula><mml:math id="M38"><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:msqrt><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msqrt></mml:math></inline-formula> (on the Y axis) for estimating the 98th quantile. Each column represents the shape <italic>r</italic> and each row the number of centers. The methods are color-coded. The rest of the quantiles are depicted separately in <xref ref-type="fig" rid="F5">Figure 5</xref> due to having a different error scale.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fams-09-1267034-g0004.tif"/>
</fig>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Estimation errors - <italic>Q</italic><sub>0.02</sub>,<italic>Q</italic><sub>0.25</sub>,<italic>Q</italic><sub>0.5</sub>,<italic>Q</italic><sub>0.75</sub>. The figure shows the standardized errors, <inline-formula><mml:math id="M39"><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:msqrt><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:msqrt></mml:math></inline-formula> (on the Y axis) for estimating the quantiles (on the X axis). Each column represents the shape <italic>r</italic> and each row the number of centers. The methods are color-coded.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fams-09-1267034-g0005.tif"/>
</fig>
<p>The YJ data estimator achieved lower MSE than the quantile loss estimator. For the extreme quantiles, the decrease in MSE ranged from 14% to 44%. The &#x0201C;one pass&#x0201D; YJ table estimator was very accurate for estimating the median and the quartiles, but lost efficiency for the extreme quantiles with the more skewed of the two Gamma distributions and when the number of centers was large. In that setting, the estimator for <italic>Q</italic><sub>0.02</sub> suffered from negative bias and its MSE was almost 3 times as large as for the quantile loss estimator; the MSE for <italic>Q</italic><sub>0.98</sub> was about 80% larger.</p>
<p>For the settings we studied, variance was the dominant component of MSE. Bias was a substantial problem only in a small number of cases. The YJ methods had large positive bias for <italic>Q</italic><sub>0.98</sub> when <italic>r</italic> &#x0003D; 4; however, when <italic>r</italic> &#x0003D; 10, and the distribution is itself closer to normal, the bias was negligible.</p></sec></sec>
<sec id="s6">
<title>6. Summary</title>
<p>In this work we presented novel methods for federated data analysis and investigated their statistical properties. We proposed a simple algorithm for creating <italic>K</italic>-anonymous data tables in one- and two-group problems and we compared federated approaches for the non-parametric Mann-Whitney U (MWU) test and for estimating quantiles. Our federated data table is created in a &#x0201C;one pass&#x0201D; format, so that it is communication efficient.</p>
<p>For the MWU test, we found that the most powerful method was the weighted average of the MWU statistics from the individual centers, with weights reflecting the sample sizes. This statistic is also communication efficient, gives nearly identical <italic>p</italic>-values to those from the combined data and has the advantage of adjusting for inter-center heterogeneity, effectively treating each center as a block. The logged ratios of <italic>p</italic>-values from this method to those from the MWU test on the combined data were heavily concentrated around 0. With increasing effect sizes, the median <italic>p</italic>-value from the weighted MWU test was lower than that for the combined test, so actually increases power. The test based on our federated table was slightly less effective. The logged ratios of <italic>p</italic>-values were still strongly concentrated around 0, but with more spread than for the weighted average. With increasing sample size, the median <italic>p</italic>-values were slightly larger than those for the combined test. Thus, both of these tests will have almost identical power to the combined data test regardless of the level of significance desired.</p>
<p>For quantile estimation, the fully optimized YJ method consistently had the lowest MSE of the methods we compared. For the extreme quantiles, it improved by 14% to 44% over the quantile loss estimator. The &#x0201C;one pass&#x0201D; YJ table estimator had almost identical MSE for estimating the median and the quartiles, but lost efficiency for the extreme quantiles when the number of centers was large. The increase in MSE was more substantial (almost 80%) with the more skewed of the two Gamma distributions we studied. This is not surprising: our YJ method exploits a transformation to normality and is less successful when the distribution is further from the normal.</p>
<p>It is important that research on federated data analysis will relate to statistical efficiency and not just to algorithmic efficiency. Our work opens this avenue, but much more could be done. Here are some examples. One important extension is to consider the impact of federated analysis on a wider range of statistical inference procedures. Another needed direction is to consider alternative mechanisms for privacy protection and to compare them with respect to the loss in statistical efficiency. Our proposals raise a number of specific questions. The construction method for a federated summary table could be extended to multiple variables and to higher dimensions; our method creates the bins in a way fitted to a one-dimensional variable. This would be needed, for example, to produce a federated analog of a scatter plot. Our findings suggest that heterogeneity can harm the federated analysis. Methods are needed to identify heterogeneity and to account for it in the analysis. The investigation of quantile estimators could be extended to a wider class of distributions. Our implementation of the YJ method applies a single transformation to the distribution. For quantiles in the tails of the distribution, it might be better to use separate transformations in the left and in the right tails.</p>
<p>Our results, like those of [<xref ref-type="bibr" rid="B21">21</xref>], are encouraging for the use of federated statistical analysis. We show that the Mann-Whitney U test and quantile estimation can be used at close to full efficiency on federated data with the <italic>K</italic>-anonymity constraint (for <italic>K</italic> &#x0003D; 10). Similarly, Spath et al. [<xref ref-type="bibr" rid="B21">21</xref>] found little loss of efficiency for time-to-event analyses when differential privacy is applied. We do point to some potential problems, for example in coping with inter-center heterogeneity. At the same time, the challenges of federated statistical analysis can also stimulate more efficient methods; our use of the Yeo-Johnson transformation improved upon the standard quantile estimator for most of the settings examined. In any particular setting, we advise researchers to carefully assess the choice of methods for their analyses. As we show, efficiency also depends on how many centers are being federated, how diverse are the data across centers, and what statistical methods will be used for the analysis. The simulation framework that we describe and exploit here can be applied to assess and compare options.</p></sec>
<sec sec-type="data-availability" id="s7">
<title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found at: <ext-link ext-link-type="uri" xlink:href="https://github.com/oribech/federated-summary-table">https://github.com/oribech/federated-summary-table</ext-link>.</p></sec>
<sec sec-type="author-contributions" id="s8">
<title>Author contributions</title>
<p>DMS: Conceptualization, Funding acquisition, Methodology, Supervision, Writing&#x02014;original draft, Writing&#x02014;review &#x00026; editing. OB: Conceptualization, Formal analysis, Investigation, Methodology, Software, Writing&#x02014;original draft, Writing&#x02014;review &#x00026; editing. MM-K: Conceptualization, Funding acquisition, Methodology, Writing&#x02014;review &#x00026; editing.</p></sec>
</body>
<back>
<sec sec-type="funding-information" id="s9">
<title>Funding</title>
<p>The author(s) declare financial support was received for the research, authorship, and/or publication of this article. This research has received funding from the European Union&#x00027;s Horizon 2020 Framework Programme for Research and Innovation under the Specific Grant Agreement No. 785907 (Human Brain Project SGA2) and the Specific Grant Agreement No. 945539 (Human Brain Project SGA3).</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec sec-type="supplementary-material" id="s11">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fams.2023.1267034/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fams.2023.1267034/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.pdf" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/></sec>
<ref-list>
<title>References</title>
<ref id="B1">
<label>1.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Proietti</surname> <given-names>R</given-names></name> <name><surname>Rivera-Caravaca</surname> <given-names>JM</given-names></name> <name><surname>Lopez-Galvez</surname> <given-names>R</given-names></name> <name><surname>Harrison</surname> <given-names>SL</given-names></name> <name><surname>Buckley</surname> <given-names>BJR</given-names></name> <name><surname>Marin</surname> <given-names>F</given-names></name> <etal/></person-group>. <article-title>Clinical implications of different types of dementia in patients with atrial fibrillation: insights from a global federated health network analysis</article-title>. <source>Clin Cardiol.</source> (<year>2023</year>) <volume>46</volume>:<fpage>656</fpage>&#x02013;<lpage>62</lpage>. <pub-id pub-id-type="doi">10.1002/clc.24006</pub-id><pub-id pub-id-type="pmid">37038622</pub-id></citation></ref>
<ref id="B2">
<label>2.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shiri</surname> <given-names>I</given-names></name> <name><surname>Vafaei Sadr</surname> <given-names>A</given-names></name> <name><surname>Akhavan</surname> <given-names>A</given-names></name> <name><surname>Salimi</surname> <given-names>Y</given-names></name> <name><surname>Sanaat</surname> <given-names>A</given-names></name> <name><surname>Amini</surname> <given-names>M</given-names></name> <etal/></person-group>. <article-title>Decentralized collaborative multi-institutional PET attenuation and scatter correction using federated deep learning</article-title>. <source>Eur J Nuclear Med Mol Imaging</source> (<year>2023</year>) <volume>50</volume>:<fpage>1034</fpage>&#x02013;<lpage>50</lpage>. <pub-id pub-id-type="doi">10.1007/s00259-022-06053-8</pub-id><pub-id pub-id-type="pmid">36508026</pub-id></citation></ref>
<ref id="B3">
<label>3.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Annie</surname> <given-names>FH</given-names></name> <name><surname>Embrey</surname> <given-names>S</given-names></name> <name><surname>George</surname> <given-names>H</given-names></name> <name><surname>Gwinn</surname> <given-names>R</given-names></name> <name><surname>Mandapaka</surname> <given-names>S</given-names></name> <name><surname>Mukherjee</surname> <given-names>D</given-names></name> <etal/></person-group>. <article-title>Effect of sex differences in TAVR mortality using a federated database</article-title>. <source>J Am Coll Cardiol.</source> (<year>2021</year>) 77(18_Suppl _1):3370. <pub-id pub-id-type="doi">10.1016/S0735-1097(21)04724-0</pub-id></citation>
</ref>
<ref id="B4">
<label>4.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pati</surname> <given-names>S</given-names></name> <name><surname>Baid</surname> <given-names>U</given-names></name> <name><surname>Edwards</surname> <given-names>B</given-names></name> <name><surname>Sheller</surname> <given-names>M</given-names></name> <name><surname>Wang</surname> <given-names>SH</given-names></name> <name><surname>Reina</surname> <given-names>GA</given-names></name> <etal/></person-group>. <article-title>Federated learning enables big data for rare cancer boundary detection</article-title>. <source>Nat Commun.</source> (<year>2022</year>) <volume>13</volume>:<fpage>7346</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-022-33407-5</pub-id><pub-id pub-id-type="pmid">36470898</pub-id></citation></ref>
<ref id="B5">
<label>5.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ogier du Terrail</surname> <given-names>J</given-names></name> <name><surname>Leopold</surname> <given-names>A</given-names></name> <name><surname>Joly</surname> <given-names>C</given-names></name> <name><surname>Beguier</surname> <given-names>C</given-names></name> <name><surname>Andreux</surname> <given-names>M</given-names></name> <name><surname>Maussion</surname> <given-names>C</given-names></name> <etal/></person-group>. <article-title>Federated learning for predicting histological response to neoadjuvant chemotherapy in triple-negative breast cancer</article-title>. <source>Nat Med.</source> (<year>2023</year>) <volume>29</volume>:<fpage>135</fpage>&#x02013;<lpage>46</lpage>. <pub-id pub-id-type="doi">10.1038/s41591-022-02155-w</pub-id><pub-id pub-id-type="pmid">36658418</pub-id></citation></ref>
<ref id="B6">
<label>6.</label>
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Salles</surname> <given-names>A</given-names></name> <name><surname>Stahl</surname> <given-names>B</given-names></name> <name><surname>Bjaalie</surname> <given-names>J</given-names></name> <name><surname>Domingo-Ferrer</surname> <given-names>J</given-names></name> <name><surname>Rose</surname> <given-names>N</given-names></name> <name><surname>Rainey</surname> <given-names>S</given-names></name> <etal/></person-group>. <article-title>Opinion Action Plan on &#x02018;Data Protection Privacy&#x00027; (Human Brain Project).</article-title> (<year>2017</year>). Available online at: <ext-link ext-link-type="uri" xlink:href="https://sos-ch-dk-2.exo.io/public-website-production/filer_public/24/0e/240e2eaa-8a10-4a17-87bc-b056a3f0cc8c/opinion_on_data_protection_and_privacy_done_01.pdf">https://sos-ch-dk-2.exo.io/public-website-production/filer_public/24/0e/240e2eaa-8a10-4a17-87bc-b056a3f0cc8c/opinion_on_data_protection_and_privacy_done_01.pdf</ext-link></citation>
</ref>
<ref id="B7">
<label>7.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Samarati</surname> <given-names>P</given-names></name> <name><surname>Sweeney</surname> <given-names>L</given-names></name></person-group>. <source>Protecting Privacy When Disclosing Information: k-Anonymity and Its Enforcement Through Generalization and Suppression</source>. SRI International Computer Science Library (<year>1998</year>).</citation>
</ref>
<ref id="B8">
<label>8.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dwork</surname> <given-names>C</given-names></name></person-group>. <article-title>Differential privacy</article-title>. In:<person-group person-group-type="editor"><name><surname>Bugliesi</surname> <given-names>M</given-names></name> <name><surname>Preneel</surname> <given-names>B</given-names></name> <name><surname>Sassone</surname> <given-names>V</given-names></name> <name><surname>Wegener</surname> <given-names>I</given-names></name></person-group>, editors. <source>Automata, Languages and Programming</source>. Berlin; Heidelberg: Springer (<year>2006</year>). p. <fpage>1</fpage>&#x02013;<lpage>12</lpage>.</citation>
</ref>
<ref id="B9">
<label>9.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mann</surname> <given-names>HB</given-names></name> <name><surname>Whitney</surname> <given-names>DR</given-names></name></person-group>. <article-title>On a test of whether one of two random variables is stochastically larger than the other</article-title>. <source>Ann Math Stat.</source> (<year>1947</year>) <volume>18</volume>:<fpage>50</fpage>&#x02013;<lpage>60</lpage>. <pub-id pub-id-type="doi">10.1214/aoms/1177730491</pub-id></citation>
</ref>
<ref id="B10">
<label>10.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Q</given-names></name> <name><surname>Liu</surname> <given-names>Y</given-names></name> <name><surname>Chen</surname> <given-names>T</given-names></name> <name><surname>Tong</surname> <given-names>Y</given-names></name></person-group>. <article-title>Federated machine learning: concept and applications</article-title>. <source>arxiv preprint arxiv:1902.04885</source> (<year>2019</year>). <pub-id pub-id-type="doi">10.48550/ARXIV.1902.04885</pub-id></citation>
</ref>
<ref id="B11">
<label>11.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Q</given-names></name> <name><surname>Wen</surname> <given-names>Z</given-names></name> <name><surname>Wu</surname> <given-names>Z</given-names></name> <name><surname>Hu</surname> <given-names>S</given-names></name> <name><surname>Wang</surname> <given-names>N</given-names></name> <name><surname>Li</surname> <given-names>Y</given-names></name> <etal/></person-group>. <article-title>A survey on federated learning systems: vision, hype and reality for data privacy and protection</article-title>. <source>IEEE Trans Knowledge Data Eng.</source> (<year>2021</year>) <volume>35</volume>:<fpage>3347</fpage>&#x02013;<lpage>66</lpage>. <pub-id pub-id-type="doi">10.1109/2Ftkde.2021.3124599</pub-id></citation>
</ref>
<ref id="B12">
<label>12.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kairouz</surname> <given-names>P</given-names></name> <name><surname>McMahan</surname> <given-names>HB</given-names></name> <name><surname>Avent</surname> <given-names>B</given-names></name> <name><surname>Bellet</surname> <given-names>A</given-names></name> <name><surname>Bennis</surname> <given-names>M</given-names></name> <name><surname>Nitin Bhagoji</surname> <given-names>A</given-names></name> <etal/></person-group>. <article-title>Advances and open problems in federated learning</article-title>. <source>Found Trends Mach Learn.</source> (<year>2021</year>) <volume>14</volume>:<fpage>1</fpage>&#x02013;<lpage>210</lpage>. <pub-id pub-id-type="doi">10.1561/2200000083</pub-id></citation>
</ref>
<ref id="B13">
<label>13.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McMahan</surname> <given-names>HB</given-names></name> <name><surname>Moore</surname> <given-names>E</given-names></name> <name><surname>Ramage</surname> <given-names>D</given-names></name> <name><surname>Hampson</surname> <given-names>S</given-names></name> <name><surname>Bay</surname> <given-names>A</given-names></name></person-group>. <article-title>Communication-efficient learning of deep networks from decentralized data</article-title>. <source>arxiv preprint arxiv:1602.05629</source> (<year>2016</year>). <pub-id pub-id-type="doi">10.48550/ARXIV.1602.05629</pub-id></citation>
</ref>
<ref id="B14">
<label>14.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>T</given-names></name> <name><surname>Sahu</surname> <given-names>AK</given-names></name> <name><surname>Talwalkar</surname> <given-names>A</given-names></name> <name><surname>Smith</surname> <given-names>V</given-names></name></person-group>. <article-title>Federated learning: challenges, methods, and future directions</article-title>. <source>IEEE Signal Process Magaz.</source> (<year>2020</year>) <volume>37</volume>:<fpage>5060</fpage>. <pub-id pub-id-type="doi">10.1109/MSP.2020.2975749</pub-id></citation>
</ref>
<ref id="B15">
<label>15.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>X</given-names></name> <name><surname>Jiang</surname> <given-names>M</given-names></name> <name><surname>Zhang</surname> <given-names>X</given-names></name> <name><surname>Kamp</surname> <given-names>M</given-names></name> <name><surname>Dou</surname> <given-names>Q</given-names></name></person-group>. <article-title>Fed{bn}: federated learning on non-{iid} features via local batch normalization</article-title>. In: <source>International Conference on Learning Representations</source>. (<year>2021</year>).</citation>
</ref>
<ref id="B16">
<label>16.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hwang</surname> <given-names>H</given-names></name> <name><surname>Yang</surname> <given-names>S</given-names></name> <name><surname>Kim</surname> <given-names>D</given-names></name> <name><surname>Dua</surname> <given-names>R</given-names></name> <name><surname>Kim</surname> <given-names>JY</given-names></name> <name><surname>Yang</surname> <given-names>E</given-names></name> <etal/></person-group>. <article-title>Towards the practical utility of federated learning in the medical domain</article-title>. In:<person-group person-group-type="editor"><name><surname>Mortazavi</surname> <given-names>BJ</given-names></name> <name><surname>Sarker</surname> <given-names>T</given-names></name> <name><surname>Beam</surname> <given-names>A</given-names></name> <name><surname>Ho</surname> <given-names>JC</given-names></name></person-group>, editors. <source>Proceedings of the Conference on Health, Inference, and Learning</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>PMLR</publisher-name> (<year>2023</year>). p. <fpage>163</fpage>&#x02013;<lpage>81</lpage>.</citation>
</ref>
<ref id="B17">
<label>17.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nasirigerdeh</surname> <given-names>R</given-names></name> <name><surname>Torkzadehmahani</surname> <given-names>R</given-names></name> <name><surname>Matschinske</surname> <given-names>J</given-names></name> <name><surname>Frisch</surname> <given-names>T</given-names></name> <name><surname>List</surname> <given-names>M</given-names></name> <name><surname>Sp&#x000E4;th</surname> <given-names>J</given-names></name> <etal/></person-group>. <article-title>sPLINK: a federated, privacy-preserving tool as a robust alternative to meta-analysis in genome-wide association studies</article-title>. <source>bioRxiv</source>. (<year>2020</year>). <pub-id pub-id-type="doi">10.1101/2020.06.05.136382</pub-id><pub-id pub-id-type="pmid">35073941</pub-id></citation></ref>
<ref id="B18">
<label>18.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Duan</surname> <given-names>R</given-names></name> <name><surname>Boland</surname> <given-names>MR</given-names></name> <name><surname>Liu</surname> <given-names>Z</given-names></name> <name><surname>Liu</surname> <given-names>Y</given-names></name> <name><surname>Chang</surname> <given-names>HH</given-names></name> <name><surname>Xu</surname> <given-names>H</given-names></name> <etal/></person-group>. <article-title>Learning from electronic health records across multiple sites: A communication-efficient and privacy-preserving distributed algorithm</article-title>. <source>J Am Med Inform Assoc</source>. (<year>2019</year>) <volume>27</volume>:<fpage>376</fpage>&#x02013;<lpage>85</lpage>. <pub-id pub-id-type="doi">10.1093/jamia/ocz199</pub-id><pub-id pub-id-type="pmid">31816040</pub-id></citation></ref>
<ref id="B19">
<label>19.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Duan</surname> <given-names>R</given-names></name> <name><surname>Boland</surname> <given-names>MR</given-names></name> <name><surname>Moore</surname> <given-names>JH</given-names></name> <name><surname>Chen</surname> <given-names>Y</given-names></name></person-group>. <article-title>ODAL: A one-shot distributed algorithm to perform logistic regressions on electronic health records data from multiple clinical sites</article-title>. In: <source>Pacific Symposium on Biocomputing</source> (<year>2019</year>). p. <fpage>30</fpage>&#x02013;<lpage>41</lpage>.<pub-id pub-id-type="pmid">30864308</pub-id></citation></ref>
<ref id="B20">
<label>20.</label>
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Q</given-names></name> <name><surname>Ihler</surname> <given-names>A</given-names></name></person-group>. <article-title>Distributed estimation, information loss and exponential families</article-title>. In:<person-group person-group-type="editor"><name><surname>Ghahramani</surname> <given-names>Z</given-names></name> <name><surname>Welling</surname> <given-names>M</given-names></name> <name><surname>Cortes</surname> <given-names>C</given-names></name> <name><surname>Lawrence</surname> <given-names>N</given-names></name> <name><surname>Weinberger</surname> <given-names>KQ</given-names></name></person-group>, editors. <source>Advances in Neural Information Processing Systems.</source> Curran Associates, Inc. (<year>2014</year>). Available online at: <ext-link ext-link-type="uri" xlink:href="https://proceedings.neurips.cc/paper_files/paper/2014/file/303ed4c69846ab36c2904d3ba8573050-Paper.pdf">https://proceedings.neurips.cc/paper_files/paper/2014/file/303ed4c69846ab36c2904d3ba8573050-Paper.pdf</ext-link></citation>
</ref>
<ref id="B21">
<label>21.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Spath</surname> <given-names>J</given-names></name> <name><surname>Matschinske</surname> <given-names>J</given-names></name> <name><surname>Kamanu</surname> <given-names>FK</given-names></name> <name><surname>Murphy</surname> <given-names>SA</given-names></name> <name><surname>Zolotareva</surname> <given-names>O</given-names></name> <name><surname>Bakhtiari</surname> <given-names>M</given-names></name> <etal/></person-group>. <article-title>Privacy-aware multi-institutional time-to-event studies</article-title>. <source>PLoS Digit Health</source>. (<year>2022</year>) <volume>1</volume>:<fpage>e0000101</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pdig.0000101</pub-id><pub-id pub-id-type="pmid">36812603</pub-id></citation></ref>
<ref id="B22">
<label>22.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rosenblatt</surname> <given-names>JD</given-names></name> <name><surname>Nadler</surname> <given-names>B</given-names></name></person-group>. <article-title>On the optimality of averaging in distributed statistical learning</article-title>. <source>Inform Inference</source>. (<year>2016</year>) <volume>53</volume>:<fpage>79</fpage>&#x02013;<lpage>404</lpage>. <pub-id pub-id-type="doi">10.1093/imaiai/iaw013</pub-id></citation>
</ref>
<ref id="B23">
<label>23.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fisher</surname> <given-names>RA</given-names></name></person-group>. <source>Statistical Methods for Research Workers</source>. 4th ed. Oxford: Oliver &#x00026; Boyd (<year>1932</year>).</citation>
</ref>
<ref id="B24">
<label>24.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kaplan</surname> <given-names>H</given-names></name> <name><surname>Schnapp</surname> <given-names>S</given-names></name> <name><surname>Stemmer</surname> <given-names>U</given-names></name></person-group>. <article-title>Differentially private approximate quantiles</article-title>. In: <source>Proceedings of the 39th International Conference on Machine Learning</source>. (<year>2022</year>). p. <fpage>10751</fpage>&#x02013;<lpage>61</lpage>.</citation>
</ref>
<ref id="B25">
<label>25.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yeo</surname> <given-names>IK</given-names></name> <name><surname>Johnson</surname> <given-names>RA</given-names></name></person-group>. <article-title>A new family of power transformations to improve normality or symmetry</article-title>. <source>Biometrika.</source> (<year>2000</year>) <volume>87</volume>:<fpage>954</fpage>&#x02013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1093/biomet/87.4.954</pub-id></citation>
</ref>
<ref id="B26">
<label>26.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Box</surname> <given-names>GEP</given-names></name> <name><surname>Cox</surname> <given-names>DR</given-names></name></person-group>. <article-title>An analysis of transformations</article-title>. <source>J R Stat Soc Ser B Methodol</source>. (<year>1964</year>) <volume>26</volume>:<fpage>211</fpage>&#x02013;<lpage>52</lpage>.</citation>
</ref>
</ref-list>
</back>
</article>