<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurosci.</journal-id>
<journal-title>Frontiers in Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-453X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnins.2022.837646</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Unsupervised Black-Box Model Domain Adaptation for Brain Tumor Segmentation</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Liu</surname> <given-names>Xiaofeng</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1507601/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Yoo</surname> <given-names>Chaehwa</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1602641/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Xing</surname> <given-names>Fangxu</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1387513/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Kuo</surname> <given-names>C.-C. Jay</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>El Fakhri</surname> <given-names>Georges</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Kang</surname> <given-names>Je-Won</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Woo</surname> <given-names>Jonghye</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1367532/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Gordon Center for Medical Imaging, Massachusetts General Hospital and Harvard Medical School</institution>, <addr-line>Boston, MA</addr-line>, <country>United States</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Electronic and Electrical Engineering and Graduate Program in Smart Factory, Ewha Womans University</institution>, <addr-line>Seoul</addr-line>, <country>South Korea</country></aff>
<aff id="aff3"><sup>3</sup><institution>Department of Electrical and Computer Engineering, University of Southern California</institution>, <addr-line>Los Angeles, CA</addr-line>, <country>United States</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Tolga Cukur, Bilkent University, Turkey</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Ilkay Oksuz, Istanbul Technical University, Turkey; Andac Hamamci, Yeditepe University, Turkey</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Jonghye Woo <email>jwoo&#x00040;mgh.harvard.edu</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Brain Imaging Methods, a section of the journal Frontiers in Neuroscience</p></fn></author-notes>
<pub-date pub-type="epub">
<day>02</day>
<month>06</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>16</volume>
<elocation-id>837646</elocation-id>
<history>
<date date-type="received">
<day>16</day>
<month>12</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>22</day>
<month>02</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Liu, Yoo, Xing, Kuo, El Fakhri, Kang and Woo.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Liu, Yoo, Xing, Kuo, El Fakhri, Kang and Woo</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Unsupervised domain adaptation (UDA) is an emerging technique that enables the transfer of domain knowledge learned from a labeled source domain to unlabeled target domains, providing a way of coping with the difficulty of labeling in new domains. The majority of prior work has relied on both source and target domain data for adaptation. However, because of privacy concerns about potential leaks in sensitive information contained in patient data, it is often challenging to share the data and labels in the source domain and trained model parameters in cross-center collaborations. To address this issue, we propose a practical framework for UDA with a black-box segmentation model trained in the source domain only, without relying on source data or a white-box source model in which the network parameters are accessible. In particular, we propose a knowledge distillation scheme to gradually learn target-specific representations. Additionally, we regularize the confidence of the labels in the target domain via unsupervised entropy minimization, leading to performance gain over UDA without entropy minimization. We extensively validated our framework on a few datasets and deep learning backbones, demonstrating the potential for our framework to be applied in challenging yet realistic clinical settings.</p></abstract>
<kwd-group>
<kwd>unsupervised domain adaptation</kwd>
<kwd>black-box model</kwd>
<kwd>segmentation</kwd>
<kwd>brain tumor</kwd>
<kwd>MR image</kwd>
<kwd>knowledge distillation</kwd>
</kwd-group>
<contract-num rid="cn001">P41EB022544</contract-num>
<contract-num rid="cn001">R01DC018511</contract-num>
<contract-sponsor id="cn001">National Institutes of Health<named-content content-type="fundref-id">10.13039/100000002</named-content></contract-sponsor>
<counts>
<fig-count count="5"/>
<table-count count="5"/>
<equation-count count="7"/>
<ref-count count="55"/>
<page-count count="11"/>
<word-count count="8046"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Semantic segmentation provides the pixel-wise annotation of lesions or anatomical structures and has been an important prerequisite for early diagnosis and treatment planning (Liu et al., <xref ref-type="bibr" rid="B23">2020a</xref>,<xref ref-type="bibr" rid="B26">b</xref>; He et al., <xref ref-type="bibr" rid="B11">2022</xref>). Because of the high cost of manual delineations, there is a large demand for automatic segmentation tools for clinical practice. For the past several years, with the development of data-driven deep learning, the performance of segmentation tasks has been substantially improved (Liu et al., <xref ref-type="bibr" rid="B29">2020c</xref>, <xref ref-type="bibr" rid="B32">2021h</xref>). For example, U-Net and its follow-up backbones achieved outstanding performance compared with their predecessors, in many natural and medical image analysis tasks, including the brain tumor localization and segmentation from magnetic resonance (MR) images (MRI) (Liu et al., <xref ref-type="bibr" rid="B37">2020d</xref>; He et al., <xref ref-type="bibr" rid="B11">2022</xref>).</p>
<p>The performance of a pre-trained deep learning model, however, can be substantially degraded, when its training distribution (i.e., source domain) differs from a testing distribution (i.e., target domain). This is because the majority of deep learning architectures assume that the source and target data distributions are independent and identically distributed (<italic>i</italic>.<italic>i</italic>.<italic>d</italic>.) and thus invariant across domains. This assumption, however, is deemed unrealistic in many clinical settings. For example, tumors with different grades are likely to exhibit different data distributions, due to varying degrees of tumor severity and growth patterns (Liu et al., <xref ref-type="bibr" rid="B34">2021j</xref>). In addition, in cross-center collaborations, data acquired even with the same vendor and with the same acquisition protocol can be substantially different from one another. Furthermore, under many multimodal MR image segmentation scenarios, cross-modality domain shifts, e.g., T2-weighted to T1-weighted MRI, can arise, leading to large performance degradation.</p>
<p>To accommodate the difference in distributions between training and testing data, a possible solution is to fine-tune developed models with supervised training, which requires pixel-wise ground truth labeling in the target domain. Since it is costly to annotate high-quality labeled data in new target domains, unsupervised domain adaptation (UDA) has been developed (Liu et al., <xref ref-type="bibr" rid="B28">2021e</xref>) to adapt the model trained in a labeled source domain to different and unlabeled target domains. In the conventional UDA, segmentation models have been trained using both source and target data, but only the source data are labeled at the adaptation stage. Promising results have been reported by means of co-training models with source domain data, primarily by enforcing the similar feature distribution of source and target domains with maximum mean discrepancy minimization (Long et al., <xref ref-type="bibr" rid="B38">2015</xref>), adversarial training (Liu et al., <xref ref-type="bibr" rid="B22">2021a</xref>), and self-training (Zou et al., <xref ref-type="bibr" rid="B55">2019</xref>).</p>
<p>Although UDA offers a promising solution to the problem of domain shift, because of privacy concerns about sensitive patient data being leaked, it is often challenging to access data and their labels in the source domain and trained model parameters in cross-center collaborations (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>). Cross-center data sharing usually requires sophisticated anonymous processing and ethics approvals, which can hinder fast deployment. In addition, large-scale and well-labeled medical datasets can be a valuable core competence for both research and commercial institutes. To address this issue, Liu et al. (<xref ref-type="bibr" rid="B35">2021k</xref>) have proposed a source-free or source-relaxed UDA approach (i.e., white-box domain adaptation) for segmentation. In that work, an off-the-shelf segmentation model was adapted to a target domain via a pre-trained model in a source domain, by transferring its batch normalization statistics. Recently, a deep inversion technique (Yin et al., <xref ref-type="bibr" rid="B51">2020</xref>) has shown that original training data can be recovered from knowledge used during white-box domain adaptation, which may leak confidential information and raise privacy concerns over patient data (Zhang et al., <xref ref-type="bibr" rid="B52">2021</xref>). In addition, source-free UDA usually relies on the same network structure as in the trained source domain, which is not flexible to update state-of-the-art or lightweight backbones to achieve better performance or implementation on memory-limited mobile devices.</p>
<p>This study aims to overcome these limitations by developing a black-box domain adaptation approach, in which we opt to restrict the use of knowledge from a source segmentation model, and do not rely on the network parameters. As a result, we provide stricter protection of medical data privacy. In addition, public release of large-scaled trained and packaged models can be easily applied to task-specific adaptation, such as segmentation and classification. To the best of our knowledge, this is the first attempt at achieving UDA for deep segmentation networks using black-box domain adaptation. Our prior work showed an initial network design and concept (Liu et al., <xref ref-type="bibr" rid="B36">2022</xref>). Building upon that work, the present study describes refined network architectures and provides extensive validations on a few different datasets and network backbones. The black-box setting provides a more effective way to protect privacy, compared with white-box domain adaptation approaches (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>) or conventional UDA approaches (Zou et al., <xref ref-type="bibr" rid="B55">2019</xref>). To our knowledge, no prior work has yet been reported on recovering data from a &#x0201C;black-box&#x0201D; model. Recently, Zhang et al. (<xref ref-type="bibr" rid="B52">2021</xref>) proposed to use black-box UDA for classification, with class-wise noise rate estimation and category-wise sampling. That work presented iterative learning with noisy labels, in which the black-box predictions were considered noisy labels. However, that work cannot be directly applied to the segmentation task to perform pixel-wise classification. Additionally, a few attempts have been made to carry out black-box domain adaptation (Liu et al., <xref ref-type="bibr" rid="B36">2022</xref>), although they could be used in more challenging yet realistic clinical scenarios.</p>
</sec>
<sec id="s2">
<title>2. Related Work</title>
<sec>
<title>2.1. Semantic Segmentation</title>
<p>The fully convolutional network (FCN) (Long et al., <xref ref-type="bibr" rid="B38">2015</xref>) was a pioneering work of deep semantic segmentation. Then, the Pyramid Scene Parsing Network (PSPNet) (Zhao et al., <xref ref-type="bibr" rid="B53">2017</xref>) was proposed to exploit the spatial feature at different scales of FCN. Recently, U-Net (Ronneberger et al., <xref ref-type="bibr" rid="B42">2015</xref>) has been widely used as the backbone of many segmentation networks, which have skip connections between the encoder and decoder to adaptively learn the correlations at different resolution scales. Other than the conventional convolutional layers used in vanilla U-Net, more advanced versions of U-Net, including ResNet (He et al., <xref ref-type="bibr" rid="B10">2016</xref>) and MobileNet (Howard et al., <xref ref-type="bibr" rid="B13">2017</xref>) have also been proposed to further boost performance or efficiency. Our black-box UDA framework is agnostic to any segmentation network, where the network used in source and target domains can be different to fit into specific requirements in implementation.</p>
</sec>
<sec>
<title>2.2. Unsupervised Domain Adaptation</title>
<p>Unsupervised domain adaptation (He et al., <xref ref-type="bibr" rid="B8">2020a</xref>,<xref ref-type="bibr" rid="B9">b</xref>; Liu et al., <xref ref-type="bibr" rid="B25">2021c</xref>,<xref ref-type="bibr" rid="B28">e</xref>) has been an important technology to alleviate the problem of domain shift and costly labeling in a new domain. Conventional approaches have utilized both source and target domain data for training (Liu et al., <xref ref-type="bibr" rid="B22">2021a</xref>,<xref ref-type="bibr" rid="B27">d</xref>,<xref ref-type="bibr" rid="B31">g</xref>,<xref ref-type="bibr" rid="B33">i</xref>,<xref ref-type="bibr" rid="B34">j</xref>). Recently, source-free UDA (Bateson et al., <xref ref-type="bibr" rid="B1">2020</xref>; Liang et al., <xref ref-type="bibr" rid="B21">2020</xref>; Wang et al., <xref ref-type="bibr" rid="B49">2020</xref>) has been proposed, which uses a pre-trained model rather than co-training the network with source and target domain data. We note that domain generalization (Liu et al., <xref ref-type="bibr" rid="B24">2021b</xref>), a closely related but different task, assumes that there are no target domain data in its training. A recent work (Liu et al., <xref ref-type="bibr" rid="B30">2021f</xref>) explored shared or domain-specific batch-normalization statistics to achieve domain alignment.</p>
</sec>
<sec>
<title>2.3. Model Transfer</title>
<p>Early works (Joachims et al., <xref ref-type="bibr" rid="B15">1999</xref>; Duan et al., <xref ref-type="bibr" rid="B4">2009</xref>) for adapting a model with parameters attempted to transfer a trained source classifier with a subset of labeled samples, which is only applicable for semi-supervised adaptation tasks. Kuzborskij and Orabona (<xref ref-type="bibr" rid="B19">2013</xref>) proposed a detailed theoretical analysis of hypothesis transfer learning for linear regression, which is the basis for subsequent UDA solutions that do not rely on source data at the adaptation stage (Chidlovskii et al., <xref ref-type="bibr" rid="B3">2016</xref>). In the deep learning era, Liang et al. (<xref ref-type="bibr" rid="B21">2020</xref>) proposed to fix the last few layers by turning the feature extraction parts into information maximization and pseudo-label-based self-training. Recently, Li et al. (<xref ref-type="bibr" rid="B20">2020</xref>) proposed using conditional generative adversarial networks (GAN) to generate images at the adaptation stage. Similarly, Kundu et al. (<xref ref-type="bibr" rid="B18">2020</xref>) utilized GAN to explore conditional entropy. However, all of the above methods require knowledge of the network parameters, which thereby can be regarded as white-box source-free UDA.</p>
</sec>
<sec>
<title>2.4. Knowledge Distillation</title>
<p>Knowledge distillation is proposed to transfer knowledge learned by a teacher model to a student model. Typically, the teacher model has larger backbones with more parameters, while the student one is typically a more compact model. Therefore, it is possible to efficiently compact a model with little sacrifice of performance. The conventional solution used a distillation loss function to enforce the consistency between the outputs of teacher and student models with the same input sample (Hinton et al., <xref ref-type="bibr" rid="B12">2015</xref>). Essentially, the knowledge distillation is an adaptive label smoothing regularization (Szegedy et al., <xref ref-type="bibr" rid="B46">2016</xref>). Kim et al. (<xref ref-type="bibr" rid="B16">2020</xref>) showed that the previous prediction can teach the network with a self-knowledge distillation scheme, which can be potentially used for semi-supervised learning. A recent work (Samuli and Timo, <xref ref-type="bibr" rid="B44">2017</xref>) assembled the prediction along with the training as a teacher model prediction. Rather than using the average teacher model predictions, Tarvainen and Valpola (<xref ref-type="bibr" rid="B47">2017</xref>) used the averaged previous model parameters as a teacher model.</p>
</sec>
</sec>
<sec sec-type="methods" id="s3">
<title>3. Methodology</title>
<p>Image segmentation partitions medical images into coherent regions for different lesions or anatomical structures, and is essential for many computer-aided diagnosis systems. A typical solution would be to formulate the segmentation task as a pixel-wise classification. <italic>f</italic><sub><italic>s</italic></sub> takes an encoder and decoder structure to map an input image, e.g., an MRI slice in the BraTS database <inline-formula><mml:math id="M1"><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mn>128</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>128</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, to its corresponding segmentation map <inline-formula><mml:math id="M2"><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mn>128</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>128</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, where <italic>C</italic> is the number of classes.</p>
<p>Considering the potential distribution shift between two domains, we assume that there are a source domain <italic>p</italic><sub><italic>s</italic></sub>(<italic>x, y</italic>) and a target domain <italic>p</italic><sub><italic>t</italic></sub>(<italic>x, y</italic>), where <italic>x</italic> indicates the to be segmented image and <italic>y</italic> is its corresponding label of the segmentation map. In the setting of black-box UDA segmentation, we have a segmentation network <italic>f</italic><sub><italic>s</italic></sub> trained with a labeled source domain set <inline-formula><mml:math id="M3"><mml:mrow><mml:msub><mml:mi mathvariant='-tex-caligraphic'>D</mml:mi><mml:mi mathvariant='-tex-caligraphic'>S</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x0007B;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>&#x0007D;</mml:mo></mml:mrow></mml:math></inline-formula> drawn <italic>i</italic>.<italic>i</italic>.<italic>d</italic>. from <italic>p</italic><sub><italic>s</italic></sub>(<italic>x, y</italic>), where <italic>f</italic><sub><italic>s</italic></sub> is fixed and accessed only through a nontransparent API during the adaptation stage. At the adaptation stage, we only have access to a black-box <italic>f</italic><sub><italic>s</italic></sub> and an unlabeled target domain set <inline-formula><mml:math id="M4"><mml:mrow><mml:msub><mml:mi mathvariant='-tex-caligraphic'>D</mml:mi><mml:mi mathvariant='-tex-caligraphic'>T</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x0007B;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x0007D;</mml:mo></mml:mrow></mml:math></inline-formula> drawn <italic>i</italic>.<italic>i</italic>.<italic>d</italic>. from the marginal distribution <italic>p</italic><sub><italic>t</italic></sub>(<italic>x</italic>), to train a target domain network <italic>f</italic><sub><italic>t</italic></sub> to achieve a good segmentation performance in the target domain. It is noteworthy that the backbones of <italic>f</italic><sub><italic>s</italic></sub> and <italic>f</italic><sub><italic>t</italic></sub> do not need to be the same. The network structure details may also not be available at the adaptation stage.</p>
<p>In this work, we propose a practical solution to black-box UDA for segmentation with a noise-aware knowledge distillation scheme using pseudo labels with exponential mixup decay (EMD). The framework is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Illustration of our black-box UDA framework using knowledge distillation with exponential mixup decay (EMD) pseudo label and unsupervised entropy minimization. Only the red shaded parts are used in testing in the target domain. Note that <italic>f</italic><sub><italic>s</italic></sub> and <italic>f</italic><sub><italic>t</italic></sub> can have different backbones.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-837646-g0001.tif"/>
</fig>
<sec>
<title>3.1. Supervised Source Domain Training</title>
<p>A good source model is a basis for the target domain adaptation performance. The UDA is motivated by the following theorem (Kouw, <xref ref-type="bibr" rid="B17">2018</xref>):</p>
<p><bold>Theorem 1</bold> For a hypothesis <italic>h</italic></p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M5"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02264;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>d</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003F5;</mml:mi><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M6"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M7"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> denote the expected loss with hypothesis <italic>h</italic> in the source and target domains, respectively, and <italic>d</italic>[&#x000B7;] measures the divergence of the marginal distributions of <italic>x</italic> between two domains (Salimans et al., <xref ref-type="bibr" rid="B43">2016</xref>). We note that the last term &#x003F5; &#x0003D; min[<italic>E</italic><sub><italic>x</italic>&#x0007E;<sub><italic>p</italic></sub><sub><italic>s</italic></sub></sub>|<italic>p</italic><sub><italic>s</italic></sub>(<italic>y</italic>|<italic>x</italic>) &#x02212; <italic>p</italic><sub><italic>t</italic></sub>(<italic>y</italic>|<italic>x</italic>)|, <italic>E</italic><sub><italic>x</italic>&#x0007E;<sub><italic>p</italic></sub><sub><italic>t</italic></sub></sub>|<italic>p</italic><sub><italic>s</italic></sub>(<italic>y</italic>|<italic>x</italic>) &#x02212; <italic>p</italic><sub><italic>t</italic></sub>(<italic>y</italic>|<italic>x</italic>)|] is usually a small value and does not affect the performance. A small <inline-formula><mml:math id="M8"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> is essential to achieve a low <inline-formula><mml:math id="M9"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>, i.e., high accuracy in the target domain.</p>
<p>For the supervision of training, a cross-entropy (CE) loss is usually used for optimization. Specifically, the pixel-wise CE loss can be formulated as:</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M10"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:munderover></mml:mstyle><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msubsup><mml:mrow><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo class="qopname">log</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>H</italic><sub>0</sub> and <italic>W</italic><sub>0</sub> are the height and width of an image, <italic>n</italic> indexes the pixel, and <italic>i</italic> indexes class labels. We note that other loss functions can also be used, e.g., dice loss, IoU loss, and boundary loss (Jadon, <xref ref-type="bibr" rid="B14">2020</xref>). The supervised source domain training is independent of the adaptation stage in our black-box UDA setting. We note that we do not re-train or fine-tune the fixed black-box source segmentation model.</p>
</sec>
<sec>
<title>3.2. Knowledge Distillation With Exponential Mixup Decay</title>
<p>Following knowledge distillation (Yin et al., <xref ref-type="bibr" rid="B51">2020</xref>), the well-trained source model can act as a teacher to provide its pixel-wise softmax histogram prediction of each image. The target domain model <italic>f</italic><sub><italic>t</italic></sub> is trained to imitate the source model <italic>f</italic><sub><italic>s</italic></sub>. The consistency of their predictions can be enforced with the Kullback-Leibler (KL) divergence between their pixel-wise softmax histogram distributions. In the conventional knowledge distillation, we assume that there is no domain shift, and the predictions of the teacher model can be reliable and simply be used as ground truth. However, due to the domain shift, the prediction of <italic>f</italic><sub><italic>s</italic></sub> in the target domain can be noisy. Simply using it as ground truth cannot outperform the source models, which is not expected in the UDA setting.</p>
<p>Accordingly, we resort to a self-training scheme (Liu et al., <xref ref-type="bibr" rid="B34">2021j</xref>) to construct the pseudo label for target domain training. Considering that the source model predictions can be a relatively reliable supervision signal, compared with unsupervised objectives in the initial epochs, we propose adjusting the contribution of the supervision signals as the training progresses. Specifically, to achieve the gradual translation to the target domain, we mix up the source and target domain predictions, i.e., <italic>f</italic><sub><italic>s</italic></sub>(<italic>x</italic><sub><italic>t</italic></sub>) and <italic>f</italic><sub><italic>t</italic></sub>(<italic>x</italic><sub><italic>t</italic></sub>), and adjust their ratio for the pseudo label <inline-formula><mml:math id="M11"><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> with EMD:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M12"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x003BB;</mml:mi><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>&#x003BB;</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msup><mml:mtext class="textrm" mathvariant="normal">exp</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mi>I</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>n</italic> indexes the pixel, and <italic>f</italic><sub><italic>s</italic></sub>(<sub><italic>x</italic><sub><italic>t</italic></sub>)<italic>n</italic></sub> and <italic>f</italic><sub><italic>t</italic></sub>(<sub><italic>x</italic><sub><italic>t</italic></sub>)<italic>n</italic></sub> are the histogram distributions of the softmax output of the <italic>n</italic>-th pixel of the predictions <italic>f</italic><sub><italic>s</italic></sub>(<italic>x</italic><sub><italic>t</italic></sub>) and <italic>f</italic><sub><italic>t</italic></sub>(<italic>x</italic><sub><italic>t</italic></sub>), respectively. &#x003BB; is the target adaptation momentum parameter with the exponential decay with respect to iteration <italic>I</italic>. &#x003BB;<sup>0</sup> is the initial weight of <italic>f</italic><sub><italic>s</italic></sub>(<italic>x</italic><sub><italic>t</italic></sub>), which is empirically set to 1. Therefore, along with the increase in iteration <italic>I</italic>, we have smaller &#x003BB;, which adjusts the contribution of the source model prediction to be large at the start of the training and to be smaller at the later training epochs. The loss knowledge distillation with the EMD pseudo label can be formulated as:</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M14"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>H</italic><sub>0</sub> and <italic>W</italic><sub>0</sub> are the height and width of the image. We note that the KL divergence is a measure of how a probability distribution, e.g., the histogram distribution of <italic>f</italic><sub><italic>t</italic></sub>(<sub><italic>x</italic><sub><italic>t</italic></sub>)<italic>n</italic></sub>, is different from a reference probability distribution, e.g., the histogram distribution of <inline-formula><mml:math id="M15"><mml:msub><mml:mrow><mml:msubsup><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Minimizing the KL divergence explicitly enforces the similarity of the two distributions of the predictions. Therefore, the weight of &#x003BB; can be smoothly decreased along with the training, and <italic>f</italic><sub><italic>t</italic></sub> gradually represents the target data.</p>
</sec>
<sec>
<title>3.3. Self-Entropy Minimization</title>
<p>In addition to the supervision signal provided by the source domain black-box model, we opt to explore unsupervised learning protocols for unlabeled target domain data. Unsupervised learning has a long history, and there are a number of possible solutions for segmentation. Among them, entropy minimization (Grandvalet and Bengio, <xref ref-type="bibr" rid="B6">2005</xref>) can be an efficient unsupervised training scheme for deep learning-based segmentation. Since it does not need a modification to the networks, it can be a simple add-on loss function on top of our framework. For implementation, the entropy for pixel segmentation can be formulated as the averaged entropy of the pixel-wise softmax prediction, given by</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M16"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>E</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:munderover></mml:mstyle><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mtext class="textrm" mathvariant="normal">log</mml:mtext><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Minimizing <inline-formula><mml:math id="M17"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>E</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> leads to the output <italic>f</italic><sub><italic>t</italic></sub>(<sub><italic>x</italic><sub><italic>t</italic></sub>)<italic>n</italic></sub> close to a one-hot distribution, i.e., confident prediction. The unsupervised learning is combined collaboratively with the black-box source model supervision to update the target model.</p>
</sec>
<sec>
<title>3.4. Overall Training Protocol</title>
<p>In summary, our training objective can be formulated as</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M18"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>K</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>E</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003B1; is used to balance between the knowledge distillation with the EMD pseudo label and the entropy minimization scheme. Since the entropy minimization may lead to a trivial solution in that the prediction of any unlabeled target samples is the same one-hot prediction (Grandvalet and Bengio, <xref ref-type="bibr" rid="B6">2005</xref>), we adopt a simple yet effective solution to stabilize the training, by linearly decreasing the hyper-parameter &#x003B1; from 5 to 0 along with the training.</p>
</sec>
</sec>
<sec id="s4">
<title>4. Experiments and Results</title>
<sec>
<title>4.1. Dataset and Data Split</title>
<p>We evaluated our approach on the BraTS2018 database (Menze et al., <xref ref-type="bibr" rid="B39">2014</xref>). In this work, we used a total of 75 patients who have low-grade gliomas (LGG), and a total of 210 patients who have high-grade gliomas (HGG) (Menze et al., <xref ref-type="bibr" rid="B39">2014</xref>) as shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. As a preprocessing step, all of the imaging modalities for each subject were registered with each other, including T1-weighted (T1), T1-contrast enhanced (T1ce), T2-weighted (T2), and T2 Fluid Attenuated Inversion Recovery (FLAIR) MRI. In addition, the voxel-wise labels for the enhancing tumor (EnhT), the peritumoral edema (ED), and the necrotic and non-enhancing tumor core (CoreT) were provided. The whole tumor includes the EnhT, ED, and CoreT. More information about the database can be found in Menze et al. (<xref ref-type="bibr" rid="B39">2014</xref>). The source and target domains have the same classes, e.g., CoreT, EnhT, ED, and background.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Examples of MRI slices of LGG and HGG samples. Each sample has four MR modalities, i.e., T1-weighted MRI, T1ce MRI, FLAIR MRI, and T2-weighted MRI.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-837646-g0002.tif"/>
</fig>
<p>Following the previous white-box source free UDA (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>) and UDA with source data (Shanis et al., <xref ref-type="bibr" rid="B45">2019</xref>), there are two evaluation protocols for UDA, i.e., cross-subtype and cross-modality UDA segmentation. In the cross-subtype setting, we used the HGG subjects as the source domain, and the LGG subjects as the target domain, which have different size and position distributions (Shanis et al., <xref ref-type="bibr" rid="B45">2019</xref>). The slices of the four modalities were concatenated as a 4-channel input with a spatial size of 128&#x000D7;128. The training of adaptation used the LGG training set. We followed the training and testing split as in Liu et al. (<xref ref-type="bibr" rid="B35">2021k</xref>). In the cross-modality setting, we used T1- or T2-weighted MRI as the source or target domain, which has a larger domain shift compared with the cross-subtype setting. Each input sample had a single modality slice with the spatial size of 128&#x000D7;128.</p>
</sec>
<sec>
<title>4.2. Training Protocol and Evaluation Metrics</title>
<p>For the cross-subtype UDA, we experimented on both HGG-to-LGG and LGG-to-HGG tasks. For the HGG-to-LGG task, our training set had a total of 210 labeled HGG subjects as the source domain and a total of 55 unlabeled LGG subjects as the target domain. The remaining 5 and 15 LGG subjects were used as the validation and testing sets, respectively. For the LGG-to-HGG task, our training set had a total of 75 labeled LGG subjects as the source domain and a total of 160 unlabeled HGG subjects as the target domain. The remaining 10 and 40 HGG subjects were used as the validation and testing sets, respectively.</p>
<p>For the cross-modality UDA, we experimented on both T2-to-T1 and T1-to-T2 tasks. For the T2-to-T1 task, our training set had a total of 55 labeled T2 subjects as the source domain and a total of 55 unlabeled T1 subjects as the target domain. The remaining 5 and 15 T1 subjects were used as the validation and testing sets, respectively. For the T1-to-T2 task, our training set had a total of 55 labeled T1 subjects as the source domain and a total of 55 unlabeled T2 subjects as the target domain. The remaining 5 and 15 T2 subjects were used as the validation and testing sets, respectively.</p>
<p>We trained <italic>f</italic><sub><italic>s</italic></sub> using our prior work (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>) and did not have access to its network parameters and source domain data at the adaptation stage. All of the networks used were based on 2D U-Net as in the previous works of UDA using the BraTS18 database (Shanis et al., <xref ref-type="bibr" rid="B45">2019</xref>). Each subject contained a total of 155 MRI slices for each modality. We used U-Net with 15 convolutional layers alongside batch normalization. Of note, for <italic>f</italic><sub><italic>s</italic></sub>, we can use different segmentation backbones as the source domain model. We evaluated two settings that use the same U-Net as <italic>f</italic><sub><italic>s</italic></sub>, and a 15-layer MobileNet-based U-Net. It is of note that MobileNet-based U-Net requires 10&#x000D7; fewer parameters, which is attributed to its separable convolutional operations. It is, therefore, easier for training, requiring much fewer parameters, which has been demonstrated in segmentation tasks on natural images. We used the validation set to tune our parameters. For both source domain only pre-training and adaptation, we used 100 epochs.</p>
<p>The target network at the adaptation stage was trained using Adam as an optimizer with &#x003B2;<sub>1</sub> &#x0003D; 0.9 and &#x003B2;<sub>2</sub> &#x0003D; 0.99. The training was performed on four NVIDIA TITAN Xp GPUs with the PyTorch deep learning toolbox (Paszke et al., <xref ref-type="bibr" rid="B40">2017</xref>), which took about 5 h for the cross-subtype task and 4 h for the cross-modality task.</p>
<p>The U-Net with ResNet-15 took about 5 h for the cross-subtype task and 4 h for the cross-modality task. In contrast, the Mobilenet-based U-Net took about 5 h for the cross-subtype task and 4 h for the cross-modality task. For testing, the ResNet and MobileNet based U-Net took about 15 ms and 8 ms for each slice, respectively.</p>
<p>The small size of MobileNet makes it possible to implement MobileNet on some memory-restricted portable devices, e.g., smartphones. We note that the use of MobileNet is to show that we do not need to use and know the same network and the network details, respectively, in the &#x0201C;black-box&#x0201D; case.</p>
<p>For evaluation, we adopted two metrics including Dice similarity coefficient (DSC) and Hausdorff distance (HD) metrics (Zou et al., <xref ref-type="bibr" rid="B54">2020</xref>). The DSC or S&#x000F8;rensen-Dice index, measures the similarity between two sets of data, e.g., pixel set in the image. DSC has been a widely used metric for evaluating image segmentation models. Specifically, it can be formulated as</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M19"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>D</mml:mi><mml:mi>S</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x01EF9;</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mo>|</mml:mo><mml:mi>&#x01EF9;</mml:mi><mml:mo>&#x02229;</mml:mo><mml:mi>y</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mi>&#x01EF9;</mml:mi><mml:mo>|</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:mo>|</mml:mo><mml:mi>y</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The HD between two point sets is defined by the sum of all minimum distances from all points from a point set to another, divided by the number of points in a point set. HD is more sensitive than DSC in terms of the segmentation boundary. In our image segmentation task, the point sets represent the voxels of the ground truth and the segmentation result, respectively, which indicates the maximum HD between the labeled boundary and the predicted boundary.</p>
</sec>
<sec>
<title>4.3. Evaluation Results</title>
<p>The segmentation results of different methods are shown in <xref ref-type="fig" rid="F3">Figure 3</xref>. BBUDA and BBUDA-Ent indicate our black-box UDA framework and the ablation study without entropy minimization, respectively. We can see that the predictions of our proposed BBUDA outperform the no adaptation model by a large margin. The better performance of BBUDA over BBUDA-Ent demonstrates the effectiveness of our entropy minimization. In addition, BBUDA&#x0002B;MobileNet indicates using the MobileNet-based U-Net as a segmentor, which has a different structure than the source domain model.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Examples of our segmentation results from an LGG MRI slice with different methods in the HGG to LGG UDA task. In addition, BBUDA-Ent represents an ablation study of the entropy minimization. We use white, dark gray, and gray color to indicate the CoreT, EnhT, and ED, respectively. Of note, OSUDA (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>) with the white-box source model for adaptation is considered an &#x0201C;upper bound.&#x0201D;</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-837646-g0003.tif"/>
</fig>
<p>For the cross-subtype UDA task, the quantitative evaluation results of the HGG-to-LGG and LGG-to-HGG tasks are shown in <xref ref-type="table" rid="T1">Tables 1</xref>, <xref ref-type="table" rid="T2">2</xref>, respectively. Our proposed BBUDA achieved the state-of-the-art performance for the black-box source-free UDA segmentation, approaching the performance of the white-box OSUDA (Bateson et al., <xref ref-type="bibr" rid="B1">2020</xref>; Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>) with the source model parameters, which can be considered an &#x0201C;upper-bound.&#x0201D; We note that the labeling ratio consistency assumption in CRUDA (Bateson et al., <xref ref-type="bibr" rid="B1">2020</xref>) does not hold in this HGG to LGG transfer task, which thus leads to inferior performance. The qualitative evaluation results are shown in <xref ref-type="fig" rid="F3">Figure 3</xref>.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Quantitative comparisons w.r.t. DSC and HD of HGG to LGG black-box.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="left"><bold>Source</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>Dice score [%]</bold> &#x02191;</th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>Hausdorff distance [mm]</bold> &#x02193;</th>
</tr>
</thead>
<tbody>
<tr>
<td/>
<td valign="top" align="left"><bold>model</bold></td>
<td valign="top" align="left"><bold>WholeT</bold></td>
<td valign="top" align="left"><bold>EnhT</bold></td>
<td valign="top" align="left"><bold>CoreT</bold></td>
<td valign="top" align="left"><bold>WholeT</bold></td>
<td valign="top" align="left"><bold>EnhT</bold></td>
<td valign="top" align="left"><bold>CoreT</bold></td>
</tr>
<tr>
<td valign="top" align="left">Source only (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>)</td>
<td valign="top" align="left">no UDA</td>
<td valign="top" align="left">79.29</td>
<td valign="top" align="left">30.09</td>
<td valign="top" align="left">44.11</td>
<td valign="top" align="left">38.7</td>
<td valign="top" align="left">46.1</td>
<td valign="top" align="left">40.2</td>
</tr>
<tr>
<td valign="top" align="left">BBUDA</td>
<td valign="top" align="left">black-box</td>
<td valign="top" align="left">82.21</td>
<td valign="top" align="left">31.33</td>
<td valign="top" align="left">46.64</td>
<td valign="top" align="left">28.6</td>
<td valign="top" align="left">26.4</td>
<td valign="top" align="left">28.1</td>
</tr>
<tr>
<td valign="top" align="left">BBUDA-Ent</td>
<td valign="top" align="left">black-box</td>
<td valign="top" align="left">81.84</td>
<td valign="top" align="left">31.26</td>
<td valign="top" align="left">45.75</td>
<td valign="top" align="left">29.4</td>
<td valign="top" align="left">27.5</td>
<td valign="top" align="left">29.0</td>
</tr>
<tr>
<td valign="top" align="left">CRUDA (Bateson et al., <xref ref-type="bibr" rid="B1">2020</xref>)</td>
<td valign="top" align="left">white-box</td>
<td valign="top" align="left">79.85</td>
<td valign="top" align="left">31.05</td>
<td valign="top" align="left">43.92</td>
<td valign="top" align="left">31.7</td>
<td valign="top" align="left">29.5</td>
<td valign="top" align="left">30.2</td>
</tr>
<tr>
<td valign="top" align="left">OSUDA (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>)</td>
<td valign="top" align="left">white-box</td>
<td valign="top" align="left">83.62</td>
<td valign="top" align="left">32.15</td>
<td valign="top" align="left">46.88</td>
<td valign="top" align="left">27.2</td>
<td valign="top" align="left">23.4</td>
<td valign="top" align="left">26.3</td>
</tr>
<tr>
<td valign="top" align="left">BBUDA&#x0002B;MobileNet</td>
<td valign="top" align="left">black-box</td>
<td valign="top" align="left">81.84</td>
<td valign="top" align="left">31.25</td>
<td valign="top" align="left">46.16</td>
<td valign="top" align="left">28.9</td>
<td valign="top" align="left">26.8</td>
<td valign="top" align="left">28.3</td>
</tr>
<tr>
<td valign="top" align="left">OSUDA&#x0002B;MobileNet</td>
<td valign="top" align="left">white-box</td>
<td valign="top" align="left">82.67</td>
<td valign="top" align="left">32.09</td>
<td valign="top" align="left">46.59</td>
<td valign="top" align="left">28.1</td>
<td valign="top" align="left">24.2</td>
<td valign="top" align="left">26.7</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Source only model indicates directly using the source domain trained model in the pre-training step without the subsequent adaptation steps</italic>.</p>
<p><italic>Of note, OSUDA (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>) with the white-box source model for adaptation is considered an &#x0201C;upper bound.&#x0201D;</italic></p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Quantitative comparisons w.r.t. DSC and HD of LGG to HGG black-box.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="left"><bold>Source</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>Dice score [%]</bold> &#x02191;</th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>Hausdorff distance [mm]</bold> &#x02193;</th>
</tr>
</thead>
<tbody>
<tr>
<td/>
<td valign="top" align="left"><bold>model</bold></td>
<td valign="top" align="left"><bold>WholeT</bold></td>
<td valign="top" align="left"><bold>EnhT</bold></td>
<td valign="top" align="left"><bold>CoreT</bold></td>
<td valign="top" align="left"><bold>WholeT</bold></td>
<td valign="top" align="left"><bold>EnhT</bold></td>
<td valign="top" align="left"><bold>CoreT</bold></td>
</tr>
<tr>
<td valign="top" align="left">Source only (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>)</td>
<td valign="top" align="left">no UDA</td>
<td valign="top" align="left">81.45</td>
<td valign="top" align="left">34.36</td>
<td valign="top" align="left">40.30</td>
<td valign="top" align="left">36.7</td>
<td valign="top" align="left">41.6</td>
<td valign="top" align="left">37.2</td>
</tr>
<tr>
<td valign="top" align="left">BBUDA</td>
<td valign="top" align="left">black-box</td>
<td valign="top" align="left">85.47</td>
<td valign="top" align="left">39.56</td>
<td valign="top" align="left">45.18</td>
<td valign="top" align="left">26.7</td>
<td valign="top" align="left">33.8</td>
<td valign="top" align="left">29.6</td>
</tr>
<tr>
<td valign="top" align="left">BBUDA-Ent</td>
<td valign="top" align="left">black-box</td>
<td valign="top" align="left">84.92</td>
<td valign="top" align="left">38.64</td>
<td valign="top" align="left">44.73</td>
<td valign="top" align="left">27.1</td>
<td valign="top" align="left">34.6</td>
<td valign="top" align="left">31.3</td>
</tr>
<tr>
<td valign="top" align="left">CRUDA (Bateson et al., <xref ref-type="bibr" rid="B1">2020</xref>)</td>
<td valign="top" align="left">white-box</td>
<td valign="top" align="left">87.62</td>
<td valign="top" align="left">40.17</td>
<td valign="top" align="left">49.65</td>
<td valign="top" align="left">23.9</td>
<td valign="top" align="left">22.7</td>
<td valign="top" align="left">23.9</td>
</tr>
<tr>
<td valign="top" align="left">OSUDA (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>)</td>
<td valign="top" align="left">white-box</td>
<td valign="top" align="left">89.75</td>
<td valign="top" align="left">44.21</td>
<td valign="top" align="left">50.34</td>
<td valign="top" align="left">22.2</td>
<td valign="top" align="left">19.3</td>
<td valign="top" align="left">21.6</td>
</tr>
<tr>
<td valign="top" align="left">BBUDA&#x0002B;MobileNet</td>
<td valign="top" align="left">black-box</td>
<td valign="top" align="left">82.14</td>
<td valign="top" align="left">31.02</td>
<td valign="top" align="left">46.13</td>
<td valign="top" align="left">27.2</td>
<td valign="top" align="left">33.0</td>
<td valign="top" align="left">28.2</td>
</tr>
<tr>
<td valign="top" align="left">OSUDA&#x0002B;MobileNet</td>
<td valign="top" align="left">white-box</td>
<td valign="top" align="left">83.36</td>
<td valign="top" align="left">31.84</td>
<td valign="top" align="left">46.62</td>
<td valign="top" align="left">23.5</td>
<td valign="top" align="left">21.6</td>
<td valign="top" align="left">22.8</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Source only model indicates directly using the source domain trained model in the pre-training step without the subsequent adaptation steps</italic>.</p>
<p><italic>Of note, OSUDA (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>) with the white-box source model for adaptation is considered an &#x0201C;upper bound.&#x0201D;</italic></p>
</table-wrap-foot>
</table-wrap>
<p>For the cross-modality UDA task, we provide the quantitative evaluation results of T2 to T1 and T1 to T2 in <xref ref-type="table" rid="T3">Tables 3</xref>, <xref ref-type="table" rid="T4">4</xref>, respectively. In addition, the qualitative evaluations are shown in <xref ref-type="fig" rid="F4">Figures 4</xref>, <xref ref-type="fig" rid="F5">5</xref>. Our proposed BBUDA improved performance in the target domain and outperformed the source model by a large margin. Both the DSC and HD metrics of our framework approached those of the &#x0201C;white-box&#x0201D; model. The sensitivity study of &#x003B1; is provided in <xref ref-type="table" rid="T5">Table 5</xref>. We found that decreasing the value of &#x003B1; yielded better performance than using a constant value of &#x003B1;.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Comparison of T2 to T1 black-box UDA.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Source</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>Dice score [%]</bold> &#x02191;</th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>Hausdorff distance [mm]</bold> &#x02193;</th>
</tr>
</thead>
<tbody>
<tr>
<td/>
<td valign="top" align="center"><bold>model</bold></td>
<td valign="top" align="center"><bold>WholeT</bold></td>
<td valign="top" align="center"><bold>EnhT</bold></td>
<td valign="top" align="center"><bold>CoreT</bold></td>
<td valign="top" align="center"><bold>WholeT</bold></td>
<td valign="top" align="center"><bold>EnhT</bold></td>
<td valign="top" align="center"><bold>CoreT</bold></td>
</tr>
<tr>
<td valign="top" align="left">Source only (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>)</td>
<td valign="top" align="center">no UDA</td>
<td valign="top" align="center">54.50</td>
<td valign="top" align="center">29.62</td>
<td valign="top" align="center">23.18</td>
<td valign="top" align="center">42.7</td>
<td valign="top" align="center">46.4</td>
<td valign="top" align="center">44.5</td>
</tr>
<tr>
<td valign="top" align="left">BBUDA</td>
<td valign="top" align="center">black-box</td>
<td valign="top" align="center">78.35</td>
<td valign="top" align="center">36.17</td>
<td valign="top" align="center">39.28</td>
<td valign="top" align="center">34.6</td>
<td valign="top" align="center">38.5</td>
<td valign="top" align="center">36.8</td>
</tr>
<tr>
<td valign="top" align="left">OSUDA (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>)</td>
<td valign="top" align="center">white-box</td>
<td valign="top" align="center">79.24</td>
<td valign="top" align="center">36.43</td>
<td valign="top" align="center">40.10</td>
<td valign="top" align="center">32.3</td>
<td valign="top" align="center">36.7</td>
<td valign="top" align="center">35.4</td>
</tr>
<tr>
<td valign="top" align="left">BBUDA&#x0002B;MobileNet</td>
<td valign="top" align="center">black-box</td>
<td valign="top" align="center">77.62</td>
<td valign="top" align="center">35.36</td>
<td valign="top" align="center">38.45</td>
<td valign="top" align="center">36.4</td>
<td valign="top" align="center">39.5</td>
<td valign="top" align="center">37.1</td>
</tr>
<tr>
<td valign="top" align="left">OSUDA&#x0002B;MobileNet</td>
<td valign="top" align="center">white-box</td>
<td valign="top" align="center">78.37</td>
<td valign="top" align="center">36.28</td>
<td valign="top" align="center">39.84</td>
<td valign="top" align="center">35.2</td>
<td valign="top" align="center">39.0</td>
<td valign="top" align="center">36.3</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The source model is trained with T2-weighted MRI slices, and the testing input is T1-weighted MRI slices. OSUDA (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>) with the white-box source model for UDA training is regarded as an &#x0201C;upper bound.&#x0201D;</italic></p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Comparison of T1 to T2 black-box UDA.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="left"><bold>Source</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>Dice score [%]</bold> &#x02191;</th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>Hausdorff distance [mm]</bold> &#x02193;</th>
</tr>
</thead>
<tbody>
<tr>
<td/>
<td valign="top" align="left"><bold>model</bold></td>
<td valign="top" align="left"><bold>WholeT</bold></td>
<td valign="top" align="left"><bold>EnhT</bold></td>
<td valign="top" align="left"><bold>cCoreT</bold></td>
<td valign="top" align="left"><bold>WholeT</bold></td>
<td valign="top" align="left"><bold>EnhT</bold></td>
<td valign="top" align="left"><bold>CoreT</bold></td>
</tr>
<tr>
<td valign="top" align="left">Source only (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>)</td>
<td valign="top" align="left">no UDA</td>
<td valign="top" align="left">52.61</td>
<td valign="top" align="left">20.44</td>
<td valign="top" align="left">22.69</td>
<td valign="top" align="left">45.3</td>
<td valign="top" align="left">47.8</td>
<td valign="top" align="left">40.9</td>
</tr>
<tr>
<td valign="top" align="left">BBUDA</td>
<td valign="top" align="left">black-box</td>
<td valign="top" align="left">76.26</td>
<td valign="top" align="left">37.30</td>
<td valign="top" align="left">39.57</td>
<td valign="top" align="left">39.5</td>
<td valign="top" align="left">42.3</td>
<td valign="top" align="left">33.7</td>
</tr>
<tr>
<td valign="top" align="left">OSUDA (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>)</td>
<td valign="top" align="left">white-box</td>
<td valign="top" align="left">77.47</td>
<td valign="top" align="left">39.64</td>
<td valign="top" align="left">40.06</td>
<td valign="top" align="left">38.4</td>
<td valign="top" align="left">41.7</td>
<td valign="top" align="left">32.3</td>
</tr>
<tr>
<td valign="top" align="left">BBUDA&#x0002B;MobileNet</td>
<td valign="top" align="left">black-box</td>
<td valign="top" align="left">76.64</td>
<td valign="top" align="left">38.25</td>
<td valign="top" align="left">38.74</td>
<td valign="top" align="left">40.6</td>
<td valign="top" align="left">43.2</td>
<td valign="top" align="left">34.8</td>
</tr>
<tr>
<td valign="top" align="left">OSUDA&#x0002B;MobileNet</td>
<td valign="top" align="left">white-box</td>
<td valign="top" align="left">77.32</td>
<td valign="top" align="left">39.37</td>
<td valign="top" align="left">38.92</td>
<td valign="top" align="left">40.1</td>
<td valign="top" align="left">42.0</td>
<td valign="top" align="left">32.5</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>OSUDA (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>) with the white-box source model for UDA training is regarded as an &#x0201C;upper bound.&#x0201D;</italic></p>
</table-wrap-foot>
</table-wrap>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Examples of our segmentation results from T1-weighted MRI with different methods in the T2 to T1 UDA task. In addition, BBUDA-Ent represents an ablation study of the entropy minimization. We use white, dark gray, and gray color to indicate the CoreT, EnhT, and ED, respectively. OIt is of note that OSUDA (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>) with the white-box source model for adaptation is considered an &#x0201C;upper bound.&#x0201D;</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-837646-g0004.tif"/>
</fig>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Examples of our segmentation results from T2-weighted MRI with different methods in the T1 to T2 UDA task. In addition, BBUDA-Ent represents an ablation study of the entropy minimization. We use white, dark gray, and gray color to indicate the CoreT, EnhT, and ED, respectively. OSUDA (Liu et al., <xref ref-type="bibr" rid="B35">2021k</xref>) with the white-box source model for adaptation is considered an &#x0201C;upper bound.&#x0201D;</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-16-837646-g0005.tif"/>
</fig>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Sensitivity analysis of the hyperparameter &#x003B1;.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>&#x003B1;</bold></th>
<th valign="top" align="left"><bold>DSC of WholeT</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">5 &#x02192; 0</td>
<td valign="top" align="left">76.26</td>
</tr>
<tr>
<td valign="top" align="left">10 &#x02192; 0</td>
<td valign="top" align="left">76.13</td>
</tr>
<tr>
<td valign="top" align="left">1 &#x02192; 0</td>
<td valign="top" align="left">76.08</td>
</tr>
<tr>
<td valign="top" align="left">0</td>
<td valign="top" align="left">72.59</td>
</tr>
<tr>
<td valign="top" align="left">5</td>
<td valign="top" align="left">75.87</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In addition to the HGG to LGG setting, we also proposed to adapt from the LGG to HGG setting. Similar to the HGG to LGG task, there were also domain gaps w.r.t. tumor types and the label proportion of each class. The results are shown in <xref ref-type="table" rid="T2">Table 2</xref>. Our proposed BBUDA achieved superior performance consistently.</p>
</sec>
</sec>
<sec sec-type="discussion" id="s5">
<title>5. Discussion</title>
<p>This work presented a UDA framework for black-box segmentation networks. The performance of the brain tumor segmentation is promising, where our framework achieved performance on par with the white-box adaptation. Therefore, our system has the potential to be applied to well-trained segmentation models in a source domain with target domain data in a range of clinical sites to combat the problem of domain shift, without the need for data sharing. Therefore, our approach enables fast, accurate, and automated lesion contouring to facilitate subsequent clinical decision processes.</p>
<p>Both the source domain data and network parameters are not accessible in our setting, which is a stricter requirement than typical UDA to guarantee data security, i.e., the privacy of patient data. A well-trained model usually requires large-scaled and well-labeled source domain data. However, the application of the system to different test-time institutes is likely to suffer from significant domain shifts, because of differences in study populations, imaging devices, imaging parameter settings, and subtype proportions, which lead to a significant performance drop. As a consequence, ways to make trained models available for collaborating target institutes are a key issue for the successful deployment of developed models. Among others, cross-institute data sharing can be a major difficulty in real-world applications. Recent deep inversion technologies further imposed restrictions on network parameter sharing. Therefore, our proposed framework can potentially alleviate the concerns over cross-institute medical data sharing.</p>
<p>The proposed knowledge distillation scheme in UDA has demonstrated its effectiveness in both cross-subtype and cross-modality tasks. The consistent loss, e.g., KL divergence, works as an efficient way to distill the knowledge in the trained black-box model. In addition, a previous study of the knowledge distillation in a single domain (Guo et al., <xref ref-type="bibr" rid="B7">2020</xref>; Vu et al., <xref ref-type="bibr" rid="B48">2021</xref>) also has shown that the student model learned with the distillation can be more general (Wang et al., <xref ref-type="bibr" rid="B50">2021</xref>). Thus, our framework can be a viable solution to train a target domain model with a decent generalization ability.</p>
<p>The hyper-parameter &#x003B1; plays an important role in balancing between the knowledge distillation and the unsupervised learning objective. The prediction of the source domain model can provide a good initialization. The target model training with only the knowledge distillation, however, can hardly outperform the source model, i.e., teacher. Therefore, it is important to utilize the unlabeled target domain data to further improve the performance in the target domain. To this end, we linearly decreased &#x003B1; from 5 to 0 for all of our experiments. While changing the start value did not affect the performance significantly, the linearly decreasing scheme is an essential step to achieving our goal. We note that setting &#x003B1; &#x0003D; 0 is equivalent to only using the knowledge distillation objective. Instead, keeping &#x003B1; as a constant in the training cannot adjust their contribution at different training stages.</p>
<p>In the present work, we were able to obtain decent segmentation results in the cross-modality segmentation task, especially on EnhT. The enhanced core shown in the brain tumor MR images is due to Blood Brain Barrier disruption in high-grade glioma. A contrast agent (e.g., gadolinium) injected into the blood stream of a patient can pass to the brain parenchyma, appearing as bright regions in the post-contrast T1-weighted MR images. Recent literature shows that this information is also encoded in non-contrast MRI images (e.g., T1, T2, and FLAIR) to some extent (Ferles and Barkhof, <xref ref-type="bibr" rid="B5">2021</xref>). In Preetha et al. (<xref ref-type="bibr" rid="B41">2021</xref>), which is one of the recent studies on T1ce synthesis, using a 3D CNN based on U-Net architecture, they reported a median Dice overlap of 28% between segmentations on synthetic and real T1ce. Although the datasets are different, they used multiple modalities as the input channels and performed training and testing in similar domains. Further investigation on synthesis and segmentation is subject to future work.</p>
<p>The backbones of the source and target models can be different. We only require that the input and output have a similar data structure. For example, in the cross-modality brain tumor segmentation task, our framework takes either T1-weighted or T2-weighted MRI slices as input and predicts the corresponding segmentation maps. The typical choice of the segmentation network would be FCN, PSPNet, or U-Net with ResNet or MobileNet backbones. More advanced backbones may provide better accuracy or reduce the computational cost aimed at different target applications. In addition, the backbones in some commercial black-box models may not be publicly available; our framework, therefore, enables the flexible use of a variety of backbones.</p>
<p>Several aspects are not fully explored in the present work. First, while we showed promising performance for brain tumor segmentation tasks with MRI, the developed framework is applied to other body parts using a variety of imaging modalities. Second, more advanced knowledge distillation and unsupervised learning methods could be analyzed beyond the current simple yet efficient framework. In addition, in the present work, we only considered a scenario, in which the source and target domains have the same segmentation classes, e.g., EnhT, ED, and CoreT in the BraTS2018 database, which is the most common case in real-world applications. Incorporating open-set UDA or out-of-distribution methods (Che et al., <xref ref-type="bibr" rid="B2">2021</xref>; Liu et al., <xref ref-type="bibr" rid="B22">2021a</xref>) can potentially lead to novel subtype discoveries.</p>
</sec>
<sec sec-type="conclusions" id="s6">
<title>6. Conclusion</title>
<p>This work proposed black-box UDA for segmentation under a realistic and meaningful scenario, presenting a practical and efficient knowledge distillation scheme with EMD pseudo labels. In particular, it provides a novel mechanism for smoothly transferring the segmentation in the source domain to the target domain with EMD to construct the pseudo label. Furthermore, unsupervised entropy minimization was incorporated into our model to improve segmentation performance. Experimental results, performed on the cross-subtype (e.g., HGG to LGG) and cross-modality (e.g., T1 to T2) adaptation tasks, demonstrated that our proposed BBUDA outperformed the source model, by a large margin, and importantly, the DSC and HD metrics of our framework were comparable to those of the white-box UDA approaches. In this work, while we only investigated brain tumor segmentation under the cross-subtype or cross-modality settings, the model could be broadly applicable to any segmentation UDA tasks using different modalities. In addition, more advanced knowledge distillation and unsupervised learning methods could be easily added to further augment performance.</p>
</sec>
<sec sec-type="data-availability" id="s7">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material. Further inquiries can be directed to the first author or corresponding author.</p>
</sec>
<sec id="s8">
<title>Ethics Statement</title>
<p>Ethical review and approval were not required for the study on human participants in accordance with the local legislation and institutional requirements. This research study was conducted retrospectively using human subject data made available in open access by BraTS18.</p>
</sec>
<sec id="s9">
<title>Author Contributions</title>
<p>XL was involved in the conceptualization, implementation, programming, and manuscript writing. CY conceived the project and was involved in writing the manuscript. FX conceived the project and was involved in writing the manuscript. C-CK conceived and supervised the project. GEF conceived and supervised the project. J-WK conceived the project and was involved in writing the manuscript. JW conceived and supervised the project and was involved in the conceptualization and writing the manuscript. All authors read and approved the final manuscript.</p>
</sec>
<sec sec-type="funding-information" id="s10">
<title>Funding</title>
<p>This work was partially supported by NIH R01DC018511 and P41EB022544.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s11">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bateson</surname> <given-names>M.</given-names></name> <name><surname>Kervadec</surname> <given-names>H.</given-names></name> <name><surname>Dolz</surname> <given-names>J.</given-names></name> <name><surname>Lombaert</surname> <given-names>H.</given-names></name> <name><surname>Ayed</surname> <given-names>I. B.</given-names></name></person-group> (<year>2020</year>). <article-title>Source-relaxed domain adaptation for image segmentation</article-title>, in <source>International Conference on Medical Image Computing and Computer-Assisted Intervention</source> (<publisher-loc>Lima</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>490</fpage>&#x02013;<lpage>499</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-59710-8_48</pub-id><pub-id pub-id-type="pmid">34734216</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Che</surname> <given-names>T.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Ge</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>R.</given-names></name> <name><surname>Xiong</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Deep verifier networks: verification of deep discriminative models with deep generative models</article-title>, in <source>AAAI</source> (<publisher-loc>New York, NY</publisher-loc>).</citation>
</ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chidlovskii</surname> <given-names>B.</given-names></name> <name><surname>Clinchant</surname> <given-names>S.</given-names></name> <name><surname>Csurka</surname> <given-names>G.</given-names></name></person-group> (<year>2016</year>). <article-title>Domain adaptation in the absence of source domain data</article-title>, in <source>Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source> (<publisher-loc>San Francisco, CA</publisher-loc>), <fpage>451</fpage>&#x02013;<lpage>460</lpage>. <pub-id pub-id-type="doi">10.1145/2939672.2939716</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Duan</surname> <given-names>L.</given-names></name> <name><surname>Tsang</surname> <given-names>I. W.</given-names></name> <name><surname>Xu</surname> <given-names>D.</given-names></name> <name><surname>Maybank</surname> <given-names>S. J.</given-names></name></person-group> (<year>2009</year>). <article-title>Domain transfer svm for video concept detection</article-title>, in <source>2009 IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Miami, FL</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1375</fpage>&#x02013;<lpage>1381</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2009.5206747</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ferles</surname> <given-names>A.</given-names></name> <name><surname>Barkhof</surname> <given-names>F.</given-names></name></person-group> (<year>2021</year>). <article-title>Seeing more with less: virtual gadolinium-enhanced glioma imaging</article-title>. <source>Lancet Digital Health</source> <volume>3</volume>, <fpage>e754</fpage>&#x02013;<lpage>e755</lpage>. <pub-id pub-id-type="doi">10.1016/S2589-7500(21)00219-3</pub-id><pub-id pub-id-type="pmid">34688601</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Grandvalet</surname> <given-names>Y.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2005</year>). <article-title>Semi-supervised learning by entropy minimization</article-title>, in <source>NIPS</source> (<publisher-loc>Vancouver, BC</publisher-loc>).</citation>
</ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>Q.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Wu</surname> <given-names>Y.</given-names></name> <name><surname>Yu</surname> <given-names>Z.</given-names></name> <name><surname>Liang</surname> <given-names>D.</given-names></name> <name><surname>Hu</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Online knowledge distillation via collaborative learning</article-title>, in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Seattle</publisher-loc>), <fpage>11020</fpage>&#x02013;<lpage>11029</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.01103</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>He</surname> <given-names>G.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Fan</surname> <given-names>F.</given-names></name> <name><surname>You</surname> <given-names>J.</given-names></name></person-group> (<year>2020a</year>). <article-title>Classification-aware semi-supervised domain adaptation</article-title>, in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops</source> (<publisher-loc>Seattle</publisher-loc>), <fpage>964</fpage>&#x02013;<lpage>965</lpage>. <pub-id pub-id-type="doi">10.1109/CVPRW50498.2020.00490</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>He</surname> <given-names>G.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Fan</surname> <given-names>F.</given-names></name> <name><surname>You</surname> <given-names>J.</given-names></name></person-group> (<year>2020b</year>). <article-title>Image2audio: facilitating semi-supervised audio emotion recognition with facial expression image</article-title>, in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops</source> (<publisher-loc>Seattle</publisher-loc>), <fpage>912</fpage>&#x02013;<lpage>913</lpage>. <pub-id pub-id-type="doi">10.1109/CVPRW50498.2020.00464</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep residual learning for image recognition</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>, <fpage>770</fpage>&#x02013;<lpage>778</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id><pub-id pub-id-type="pmid">32166560</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>He</surname> <given-names>X.</given-names></name> <name><surname>Xu</surname> <given-names>W.</given-names></name> <name><surname>Yang</surname> <given-names>C.</given-names></name> <name><surname>Mao</surname> <given-names>J.</given-names></name> <name><surname>Chen</surname> <given-names>S.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name></person-group> (<year>2022</year>). <article-title>Deep convolutional neural network with a multi-scale attention feature fusion module for segmentation of multimodal brain tumor</article-title>. <source>Front. Neurosci</source>. <volume>15</volume>, <fpage>782968</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2021.782968</pub-id><pub-id pub-id-type="pmid">34899175</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G.</given-names></name> <name><surname>Vinyals</surname> <given-names>O.</given-names></name> <name><surname>Dean</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Distilling the knowledge in a neural network</article-title>. <source>arXiv preprint arXiv:1503.02531</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1503.02531</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Howard</surname> <given-names>A. G.</given-names></name> <name><surname>Zhu</surname> <given-names>M.</given-names></name> <name><surname>Chen</surname> <given-names>B.</given-names></name> <name><surname>Kalenichenko</surname> <given-names>D.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Weyand</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>MobileNets: efficient convolutional neural networks for mobile vision applications</article-title>. <source>arXiv preprint arXiv:1704.04861</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1704.04861</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jadon</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>A survey of loss functions for semantic segmentation</article-title>, in <source>2020 IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB)</source> (<publisher-loc>Virtual</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1109/CIBCB48159.2020.9277638</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Joachims</surname> <given-names>T.</given-names></name></person-group> (<year>1999</year>). <article-title>Transductive inference for text classification using support vector machines</article-title>, in <source>ICML</source>, <fpage>200</fpage>&#x02013;<lpage>209</lpage>.<pub-id pub-id-type="pmid">19623491</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>K.</given-names></name> <name><surname>Ji</surname> <given-names>B.</given-names></name> <name><surname>Yoon</surname> <given-names>D.</given-names></name> <name><surname>Hwang</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Self-knowledge distillation: a simple way for better generalization</article-title>. <source>arXiv preprint arXiv:2006.12000</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2006.12000</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kouw</surname> <given-names>W. M.</given-names></name></person-group> (<year>2018</year>). <article-title>An introduction to domain adaptation and transfer learning</article-title>. <source>arXiv preprint arXiv:1812.11806</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1812.11806</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kundu</surname> <given-names>J. N.</given-names></name> <name><surname>Venkat</surname> <given-names>N.</given-names></name> <name><surname>Babu</surname> <given-names>R. V.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Universal source-free domain adaptation</article-title>, in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Virtual</publisher-loc>), <fpage>4544</fpage>&#x02013;<lpage>4553</lpage>.</citation>
</ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kuzborskij</surname> <given-names>I.</given-names></name> <name><surname>Orabona</surname> <given-names>F.</given-names></name></person-group> (<year>2013</year>). <article-title>Stability and hypothesis transfer learning</article-title>, in <source>International Conference on Machine Learning</source> (<publisher-loc>Atlanta</publisher-loc>: <publisher-name>PMLR</publisher-name>), <fpage>942</fpage>&#x02013;<lpage>950</lpage>.</citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>R.</given-names></name> <name><surname>Jiao</surname> <given-names>Q.</given-names></name> <name><surname>Cao</surname> <given-names>W.</given-names></name> <name><surname>Wong</surname> <given-names>H.-S.</given-names></name> <name><surname>Wu</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Model adaptation: unsupervised domain adaptation without source data</article-title>, in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>, <fpage>9641</fpage>&#x02013;<lpage>9650</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00966</pub-id><pub-id pub-id-type="pmid">34874854</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liang</surname> <given-names>J.</given-names></name> <name><surname>Hu</surname> <given-names>D.</given-names></name> <name><surname>Feng</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Do we really need to access the source data? source hypothesis transfer for unsupervised domain adaptation</article-title>, in <source>International Conference on Machine Learning</source> (<publisher-loc>PMLR</publisher-loc>), <fpage>6028</fpage>&#x02013;<lpage>6039</lpage>.</citation>
</ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Guo</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Xing</surname> <given-names>F.</given-names></name> <name><surname>You</surname> <given-names>J.</given-names></name> <name><surname>Kuo</surname> <given-names>C.-C. J.</given-names></name> <etal/></person-group>. (<year>2021a</year>). <article-title>Adversarial unsupervised domain adaptation with conditional and label shift: infer, align and iterate</article-title>, in <source>Proceedings of the IEEE/CVF International Conference on Computer Vision</source> (<publisher-loc>Virtual</publisher-loc>), <fpage>10367</fpage>&#x02013;<lpage>10376</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV48922.2021.01020</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Han</surname> <given-names>Y.</given-names></name> <name><surname>Bai</surname> <given-names>S.</given-names></name> <name><surname>Ge</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>T.</given-names></name> <name><surname>Han</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2020a</year>). <article-title>Importance-aware semantic segmentation in self-driving with discrete wasserstein training</article-title>, in <source>Proceedings of the AAAI Conference on Artificial Intelligence</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>11629</fpage>&#x02013;<lpage>11636</lpage>. <pub-id pub-id-type="doi">10.1609/aaai.v34i07.6831</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Hu</surname> <given-names>B.</given-names></name> <name><surname>Jin</surname> <given-names>L.</given-names></name> <name><surname>Han</surname> <given-names>X.</given-names></name> <name><surname>Xing</surname> <given-names>F.</given-names></name> <name><surname>Ouyang</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2021b</year>). <article-title>Domain generalization under conditional and label shifts via variational Bayesian inference</article-title>, in <source>IJCAI</source> (<publisher-loc>Virtual</publisher-loc>). <pub-id pub-id-type="doi">10.24963/ijcai.2021/122</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Hu</surname> <given-names>B.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Lu</surname> <given-names>J.</given-names></name> <name><surname>You</surname> <given-names>J.</given-names></name> <name><surname>Kong</surname> <given-names>L.</given-names></name></person-group> (<year>2021c</year>). <article-title>Energy-constrained self-training for unsupervised domain adaptation</article-title>, in <source>2020 25th International Conference on Pattern Recognition (ICPR)</source> (<publisher-loc>IEEE</publisher-loc>), <fpage>7515</fpage>&#x02013;<lpage>7520</lpage>. <pub-id pub-id-type="doi">10.1109/ICPR48806.2021.9413284</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Ji</surname> <given-names>W.</given-names></name> <name><surname>You</surname> <given-names>J.</given-names></name> <name><surname>Fakhri</surname> <given-names>G. E.</given-names></name> <name><surname>Woo</surname> <given-names>J.</given-names></name></person-group> (<year>2020b</year>). <article-title>Severity-aware semantic segmentation with reinforced wasserstein training</article-title>, in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Virtual</publisher-loc>), <fpage>12566</fpage>&#x02013;<lpage>12575</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.01258</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Ge</surname> <given-names>Y.</given-names></name> <name><surname>Ye</surname> <given-names>P.</given-names></name> <name><surname>You</surname> <given-names>J.</given-names></name> <name><surname>Lu</surname> <given-names>J.</given-names></name></person-group> (<year>2021d</year>). <article-title>Recursively conditional gaussian for ordinal unsupervised domain adaptation</article-title>, in <source>Proceedings of the IEEE/CVF International Conference on Computer Vision</source>, <fpage>764</fpage>&#x02013;<lpage>773</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV48922.2021.00080</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Hu</surname> <given-names>B.</given-names></name> <name><surname>Ji</surname> <given-names>W.</given-names></name> <name><surname>Xing</surname> <given-names>F.</given-names></name> <name><surname>Lu</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2021e</year>). <article-title>Subtype-aware unsupervised domain adaptation for medical diagnosis</article-title>, in <source>AAAI</source> (<publisher-loc>Virtual</publisher-loc>).</citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Lu</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Bai</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>You</surname> <given-names>J.</given-names></name></person-group> (<year>2020c</year>). <article-title>Wasserstein loss with alternative reinforcement learning for severity-aware semantic segmentation</article-title>. <source>IEEE trans. Intell. Transp. Syst</source>. <pub-id pub-id-type="doi">10.1109/tits.2020.3014137</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Xing</surname> <given-names>F.</given-names></name> <name><surname>El Fakhri</surname> <given-names>G.</given-names></name> <name><surname>Woo</surname> <given-names>J.</given-names></name></person-group> (<year>2021f</year>). <article-title>Adapting off-the-shelf source segmenter for target medical image segmentation</article-title>, in <source>MICCAI</source> (<publisher-loc>Virtual</publisher-loc>). <pub-id pub-id-type="doi">10.1007/978-3-030-87196-3_51</pub-id><pub-id pub-id-type="pmid">34734216</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Xing</surname> <given-names>F.</given-names></name> <name><surname>El Fakhri</surname> <given-names>G.</given-names></name> <name><surname>Woo</surname> <given-names>J.</given-names></name></person-group> (<year>2021g</year>). <article-title>A unified conditional disentanglement framework for multimodal brain MR image translation</article-title>, in <source>2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI)</source> (<publisher-loc>Virtual</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>10</fpage>&#x02013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1109/ISBI48211.2021.9433897</pub-id><pub-id pub-id-type="pmid">34567419</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Xing</surname> <given-names>F.</given-names></name> <name><surname>Gaggin</surname> <given-names>H. K.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Kuo</surname> <given-names>C.-C. J.</given-names></name> <name><surname>Fakhri</surname> <given-names>G. E.</given-names></name> <etal/></person-group>. (<year>2021h</year>). <article-title>Segmentation of cardiac structures via successive subspace learning with SAAB transform from cine mri</article-title>. <source>arXiv preprint arXiv:2107.10718</source>. <pub-id pub-id-type="doi">10.1109/EMBC46164.2021.9629770</pub-id><pub-id pub-id-type="pmid">34892002</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Xing</surname> <given-names>F.</given-names></name> <name><surname>Prince</surname> <given-names>J. L.</given-names></name> <name><surname>Carass</surname> <given-names>A.</given-names></name> <name><surname>Stone</surname> <given-names>M.</given-names></name> <name><surname>El Fakhri</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2021i</year>). <article-title>Dual-cycle constrained bijective vae-gan for tagged-to-cine magnetic resonance image synthesis</article-title>, in <source>2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI)</source> (<publisher-loc>Virtual</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1448</fpage>&#x02013;<lpage>1452</lpage>. <pub-id pub-id-type="doi">10.1109/ISBI48211.2021.9433852</pub-id><pub-id pub-id-type="pmid">34707796</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Xing</surname> <given-names>F.</given-names></name> <name><surname>Stone</surname> <given-names>M.</given-names></name> <name><surname>Zhuo</surname> <given-names>J.</given-names></name> <name><surname>Reese</surname> <given-names>T.</given-names></name> <name><surname>Prince</surname> <given-names>J. L.</given-names></name> <etal/></person-group>. (<year>2021j</year>). <article-title>Generative self-training for cross-domain unsupervised tagged-to-cine MRI synthesis</article-title>, in <source>International Conference on Medical Image Computing and Computer-Assisted Intervention</source> (<publisher-loc>Springer</publisher-loc>), <fpage>138</fpage>&#x02013;<lpage>148</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-87199-4_13</pub-id><pub-id pub-id-type="pmid">34734217</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Xing</surname> <given-names>F.</given-names></name> <name><surname>Yang</surname> <given-names>C.</given-names></name> <name><surname>El Fakhri</surname> <given-names>G.</given-names></name> <name><surname>Woo</surname> <given-names>J.</given-names></name></person-group> (<year>2021k</year>). <article-title>Adapting off-the-shelf source segmenter for target medical image segmentation</article-title>, in <source>International Conference on Medical Image Computing and Computer-Assisted Intervention</source> (<publisher-loc>Virtual</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>549</fpage>&#x02013;<lpage>559</lpage>.<pub-id pub-id-type="pmid">34734216</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Yoo</surname> <given-names>C.</given-names></name> <name><surname>Xing</surname> <given-names>F.</given-names></name> <name><surname>Kuo</surname> <given-names>C.-C. J.</given-names></name> <name><surname>El Fakhri</surname> <given-names>G.</given-names></name> <name><surname>Woo</surname> <given-names>J.</given-names></name></person-group> (<year>2022</year>). <article-title>Unsupervised domain adaptation for segmentation with black-box source model</article-title>, in <source>SPIE Medical Imaging 2022: Image Processing</source>. <pub-id pub-id-type="doi">10.1117/12.2607895</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Bai</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>You</surname> <given-names>J.</given-names></name></person-group> (<year>2020d</year>). <article-title>Reinforced wasserstein training for severity-aware semantic segmentation in autonomous driving</article-title>. <source>arXiv preprint arXiv:2008.04751</source>.</citation>
</ref>
<ref id="B38">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Long</surname> <given-names>J.</given-names></name> <name><surname>Shelhamer</surname> <given-names>E.</given-names></name> <name><surname>Darrell</surname> <given-names>T.</given-names></name></person-group> (<year>2015</year>). <article-title>Fully convolutional networks for semantic segmentation</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Boston</publisher-loc>), <fpage>3431</fpage>&#x02013;<lpage>3440</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2015.7298965</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Menze</surname> <given-names>B. H.</given-names></name> <name><surname>Jakab</surname> <given-names>A.</given-names></name> <name><surname>Bauer</surname> <given-names>S.</given-names></name> <name><surname>Kalpathy-Cramer</surname> <given-names>J.</given-names></name> <name><surname>Farahani</surname> <given-names>K.</given-names></name> <name><surname>Kirby</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>The multimodal brain tumor image segmentation benchmark (BRATS)</article-title>. <source>IEEE Trans. Med. Imaging</source> <volume>34</volume>, <fpage>1993</fpage>&#x02013;<lpage>2024</lpage>. <pub-id pub-id-type="doi">10.1109/TMI.2014.2377694</pub-id><pub-id pub-id-type="pmid">25494501</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Paszke</surname> <given-names>A.</given-names></name> <name><surname>Gross</surname> <given-names>S.</given-names></name> <name><surname>Chintala</surname> <given-names>S.</given-names></name> <name><surname>Chanan</surname> <given-names>G.</given-names></name> <name><surname>Yang</surname> <given-names>E.</given-names></name> <name><surname>DeVito</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Automatic differentiation in pytorch</article-title>, in <source>NIPS 2017 Workshop</source> (<publisher-loc>Long Beach</publisher-loc>).</citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Preetha</surname> <given-names>C. J.</given-names></name> <name><surname>Meredig</surname> <given-names>H.</given-names></name> <name><surname>Brugnara</surname> <given-names>G.</given-names></name> <name><surname>Mahmutoglu</surname> <given-names>M. A.</given-names></name> <name><surname>Foltyn</surname> <given-names>M.</given-names></name> <name><surname>Isensee</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Deep-learning-based synthesis of post-contrast t1-weighted mri for tumour response assessment in neuro-oncology: a multicentre, retrospective cohort study</article-title>. <source>Lancet Digit. Health</source> <volume>3</volume>, <fpage>e784</fpage>&#x02013;<lpage>e794</lpage>. <pub-id pub-id-type="doi">10.1016/S2589-7500(21)00205-3</pub-id><pub-id pub-id-type="pmid">34688602</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ronneberger</surname> <given-names>O.</given-names></name> <name><surname>Fischer</surname> <given-names>P.</given-names></name> <name><surname>Brox</surname> <given-names>T.</given-names></name></person-group> (<year>2015</year>). <article-title>U-net: Convolutional networks for biomedical image segmentation</article-title>, in <source>International Conference on Medical Image Computing and Computer-Assisted Intervention</source> (<publisher-loc>Springer</publisher-loc>), <fpage>234</fpage>&#x02013;<lpage>241</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-24574-4_28</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Salimans</surname> <given-names>T.</given-names></name> <name><surname>Goodfellow</surname> <given-names>I.</given-names></name> <name><surname>Zaremba</surname> <given-names>W.</given-names></name> <name><surname>Cheung</surname> <given-names>V.</given-names></name> <name><surname>Radford</surname> <given-names>A.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name></person-group> (<year>2016</year>). <article-title>Improved techniques for training GANs</article-title>, in <source>NIPS</source>.</citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Samuli</surname> <given-names>L.</given-names></name> <name><surname>Timo</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). <article-title>Temporal ensembling for semi-supervised learning</article-title>, in <source>International Conference on Learning Representations (ICLR)</source>, <fpage>6</fpage>.</citation>
</ref>
<ref id="B45">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Shanis</surname> <given-names>Z.</given-names></name> <name><surname>Gerber</surname> <given-names>S.</given-names></name> <name><surname>Gao</surname> <given-names>M.</given-names></name> <name><surname>Enquobahrie</surname> <given-names>A.</given-names></name></person-group> (<year>2019</year>). <article-title>Intramodality domain adaptation using self ensembling and adversarial training</article-title>, in <source>Domain Adaptation and Representation Transfer and Medical Image Learning with Less Labels and Imperfect Data</source>, <fpage>28</fpage>&#x02013;<lpage>36</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-33391-1_4</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Szegedy</surname> <given-names>C.</given-names></name> <name><surname>Vanhoucke</surname> <given-names>V.</given-names></name> <name><surname>Ioffe</surname> <given-names>S.</given-names></name> <name><surname>Shlens</surname> <given-names>J.</given-names></name> <name><surname>Wojna</surname> <given-names>Z.</given-names></name></person-group> (<year>2016</year>). <article-title>Rethinking the inception architecture for computer vision</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>, <fpage>2818</fpage>&#x02013;<lpage>2826</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.308</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tarvainen</surname> <given-names>A.</given-names></name> <name><surname>Valpola</surname> <given-names>H.</given-names></name></person-group> (<year>2017</year>). <article-title>Mean teachers are better role models: Weight-averaged consistency targets improve Semi-supervised deep learning results</article-title>, in <source>Advances in Neural Information Processing Systems</source> (<publisher-loc>Long Beach</publisher-loc>) (2017), <fpage>30</fpage>.</citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vu</surname> <given-names>D.-Q.</given-names></name> <name><surname>Le</surname> <given-names>N.</given-names></name> <name><surname>Wang</surname> <given-names>J.-C.</given-names></name></person-group> (<year>2021</year>). <article-title>Teaching yourself: a self-knowledge distillation approach to action recognition</article-title>. <source>IEEE Access</source> <volume>9</volume>, <fpage>105711</fpage>&#x02013;<lpage>105723</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2021.3099856</pub-id></citation>
</ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>D.</given-names></name> <name><surname>Shelhamer</surname> <given-names>E.</given-names></name> <name><surname>Liu</surname> <given-names>S.</given-names></name> <name><surname>Olshausen</surname> <given-names>B.</given-names></name> <name><surname>Darrell</surname> <given-names>T.</given-names></name></person-group> (<year>2020</year>). <article-title>Fully test-time adaptation by entropy minimization</article-title>. <source>arXiv preprint arXiv:2006.10726</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2006.10726</pub-id></citation>
</ref>
<ref id="B50">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Chau</surname> <given-names>L.-P.</given-names></name> <name><surname>Kot</surname> <given-names>A. C.</given-names></name></person-group> (<year>2021</year>). <article-title>Embracing the dark knowledge: domain generalization using regularized knowledge distillation</article-title>, in <source>Proceedings of the 29th ACM International Conference on Multimedia</source>, <fpage>2595</fpage>&#x02013;<lpage>2604</lpage>. <pub-id pub-id-type="doi">10.1145/3474085.3475434</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yin</surname> <given-names>H.</given-names></name> <name><surname>Molchanov</surname> <given-names>P.</given-names></name> <name><surname>Alvarez</surname> <given-names>J. M.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Mallya</surname> <given-names>A.</given-names></name> <name><surname>Hoiem</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Dreaming to distill: data-free knowledge transfer via deepinversion</article-title>, in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>, <fpage>8715</fpage>&#x02013;<lpage>8724</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00874</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Jia</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name></person-group> (<year>2021</year>). <article-title>Unsupervised domain adaptation of black-box source models</article-title>. <source>arXiv preprint arXiv:2101.02839</source>.</citation>
</ref>
<ref id="B53">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>H.</given-names></name> <name><surname>Shi</surname> <given-names>J.</given-names></name> <name><surname>Qi</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Jia</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>Pyramid scene parsing network</article-title>, in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>, <fpage>2881</fpage>&#x02013;<lpage>2890</lpage>. <pub-id pub-id-type="doi">10.48550/arXiv.1612.01105</pub-id><pub-id pub-id-type="pmid">33390119</pub-id></citation></ref>
<ref id="B54">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zou</surname> <given-names>D.</given-names></name> <name><surname>Zhu</surname> <given-names>Q.</given-names></name> <name><surname>Yan</surname> <given-names>P.</given-names></name></person-group> (<year>2020</year>). <article-title>Unsupervised domain adaptation with dualscheme fusion network for medical image segmentation</article-title>, in <source>IJCAI</source>, <fpage>3291</fpage>&#x02013;<lpage>3298</lpage>. <pub-id pub-id-type="doi">10.24963/ijcai.2020/455</pub-id></citation>
</ref>
<ref id="B55">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zou</surname> <given-names>Y.</given-names></name> <name><surname>Yu</surname> <given-names>Z.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Kumar</surname> <given-names>B.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>Confidence regularized self-training</article-title>, in <source>Proceedings of the IEEE/CVF International Conference on Computer Vision</source>, <fpage>5982</fpage>&#x02013;<lpage>5991</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2019.00608</pub-id></citation>
</ref>
</ref-list>
</back>
</article>