<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frai.2022.868232</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Credit Risk Modeling Using Transfer Learning and Domain Adaptation</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Suryanto</surname> <given-names>Hendra</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Mahidadia</surname> <given-names>Ashesh</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1554653/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Bain</surname> <given-names>Michael</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Guan</surname> <given-names>Charles</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Guan</surname> <given-names>Ada</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Rich Data Corporation</institution>, <addr-line>Sydney, NSW</addr-line>, <country>Australia</country></aff>
<aff id="aff2"><sup>2</sup><institution>School of Computer Science and Engineering, The University of New South Wales</institution>, <addr-line>Sydney, NSW</addr-line>, <country>Australia</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Sameena Shah, JP Morgan Chase &#x00026; Co, United States</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Bertrand Kian Hassani, University College London, United Kingdom; Jyotir Moy Chatterjee, Lord Buddha Education Foundation (LBEF), Nepal</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Ashesh Mahidadia <email>ashesh.mahidadia&#x00040;richdataco.com</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Artificial Intelligence in Finance, a section of the journal Frontiers in Artificial Intelligence</p></fn></author-notes>
<pub-date pub-type="epub">
<day>03</day>
<month>05</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>5</volume>
<elocation-id>868232</elocation-id>
<history>
<date date-type="received">
<day>02</day>
<month>02</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>23</day>
<month>03</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Suryanto, Mahidadia, Bain, Guan and Guan.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Suryanto, Mahidadia, Bain, Guan and Guan</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>In the domain of credit risk assessment lenders may have limited or no data on the historical lending outcomes of credit applicants. Typically this disproportionately affects Micro, Small, and Medium Enterprises (MSMEs), for which credit may be restricted or too costly, due to the difficulty of predicting the Probability of Default (PD). However, if data from other related credit risk domains is available <italic>Transfer Learning</italic> may be applied to successfully train models, e.g., from the credit card lending and debt consolidation (CD) domains to predict in the small business lending domain. In this article, we report successful results from an approach using transfer learning to predict the probability of default based on the novel concept of <italic>Progressive Shift Contribution</italic> (PSC) from source to target domain. Toward real-world application by lenders of this approach, we further address two key questions. The first is to explain transfer learning models, and the second is to adjust features when the source and target domains differ. To address the first question, we apply Shapley values to investigate how and why transfer learning improves model accuracy, and also propose and test a domain adaptation approach to address the second. These results show that adaptation improves model accuracy in addition to the improvement from transfer learning. We extend this by proposing and testing a combined strategy of feature selection and adaptation to convert values of source domain features to better approximate values of target domain features. Our approach includes a strategy to choose features for adaptation and an algorithm to adapt the values of these features. In this setting, transfer learning appears to improve model accuracy by increasing the contribution of less predictive features. Although the percentage improvements are small, such improvements in real world lending could be of significant economic importance.</p></abstract>
<kwd-group>
<kwd>credit risk</kwd>
<kwd>transfer learning</kwd>
<kwd>domain adaptation</kwd>
<kwd>explainable AI</kwd>
<kwd>deep learning</kwd>
</kwd-group>
<counts>
<fig-count count="6"/>
<table-count count="8"/>
<equation-count count="29"/>
<ref-count count="27"/>
<page-count count="16"/>
<word-count count="10111"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>In 2014, 42% of all adults worldwide reported borrowing in the previous year (excluding credit cards). Adults in underdeveloped nations borrow three times as much from family members and friends as from financial institutions. Borrowing through an institution offers advantages over borrowing from family or friends since it provides access to adequate funds and, presumably, better credit conditions under regulation (World Bank, <xref ref-type="bibr" rid="B24">2017</xref>). Access to formal credit has become an issue for young adults in developed countries too. A recent survey by Bankrate found that 58% of millennials (born between 1981 and 1996) in the United States have been denied at least one financial product because of their credit score (BankRate, <xref ref-type="bibr" rid="B1">2019</xref>). Consequently, fintech-based financial products, such as Alipay, Affirm, Klarna, Paypal Credit, and Afterpay have become popular with millennials and Generation Z (born between 1997 and 2010), although with these platforms credit is typically provided in very restricted domains, such as for retail purchases.</p>
<p>Applications for unsecured consumer loans such as credit cards and debt consolidation (CD) loans are common. They are typically scored by algorithms that are mostly based on a person&#x00027;s credit score, income, spending, and other factors such as job and housing stability. This area is now a crowded and competitive marketplace that has been helped by recent fintech activities (especially in the US, UK, and China). These activities have amassed historical data and, consequently, reliable and accurate scoring models. Small company financing is a relatively new sector for fintechs; it is riskier, more diversified, more difficult to forecast outcomes, and lacks data. Although public datasets on this type of lending are scarce (apart from exceptions like the Lending Club data used in this article<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref>), it is known that the quantity of historical lending outcomes for small business loans is typically far lower than for other lending types, making it very difficult to develop an accurate and stable model using traditional supervised learning. Since small company lending has less competition and larger margins than consumer lending, finding strategies to forecast loan outcomes and service this market is potentially more profitable for lenders. MSMEs are also one of the most powerful generators of economic development, innovation, and employment. MSMEs typically cite a lack of access to capital as a major stumbling block to expansion. Providing possibilities for MSMEs in developing markets is a critical step toward economic growth and poverty reduction. There are 65 million unmet financial requirements in developing nations (or 40% of formal MSMEs). Forum (<xref ref-type="bibr" rid="B5">2019</xref>) estimates that the MSME financing gap in developing nations is $5.2 trillion, which is 1.4 times the present level of MSME lending.</p>
<p>However, resolving these concerns presents various difficulties. When a lender enters new market segments, a new credit risk model is necessary to evaluate loan applicants&#x00027; credit risk. The current strategy relies on expert rules, in which a credit risk expert develops business rules based on data and their expertise and knowledge. Lenders begin by collecting sufficient labeled data with an expert model in order to develop a supervised learning model. A comparison is made between the expert model and the supervised learning model. The superior model is adopted if one model outperforms the other by a significant margin. In another way, if both models work well together, they can be combined into an &#x0201C;ensemble&#x0201D; model. Lenders in commercial lending systems usually charge more money or limit the amount of credit they can give out because there are not enough labeled data to test the expert models. Consequently, many individuals and small enterprises are shut out of these &#x0201C;conventional&#x0201D; financing systems. When sensitive and personal data can only be accessed on-site by authorized people, it may be hard to get a suitable expert to analyze such data.</p>
<p>Transfer learning is proposed as a connection between alternative lending data and standard credit history evaluation, such as credit bureau ratings. We examine how transfer learning may help with credit scoring accuracy in this research. We investigate a domain with no or limited prior lending results, such as providing credit to unbanked or underbanked populations, or micro and small firms, where historical data is scarce. Lenders currently depend mostly on expert rules for credit scoring. Lenders demand a hefty fee or refuse to grant credit because of the considerable unpredictability of such scoring methods. Transfer learning from adjacent areas might help fill in the gaps in information and increase financial inclusion. Transferring CD loan knowledge to riskier small business loans, or utility bill payments to loan repayments, e.g., might result in a more accurate scoring model. The aim of this research is to address the following issue: can we use this &#x0201C;alternative lending&#x0201D; data to enhance credit behavior prediction, and hence credit access, for those with low traditional credit histories?</p>
<p>We explore how transfer learning could be used in the early stages of a credit risk model deployment when there is a lack of historical labeled data. The stability and accuracy of model performance in the credit risk domain are business goals in order to anticipate the chance of default. We provide a method for combining the results of transferred models from related credit risk domains with new models based on newly acquired labeled data from new domains. We get a greater level of precision while also maintaining the overall model&#x00027;s stability using this method. Experiments on real-world commercial data (which we are unable to discuss in this article) indicated that utilizing a gradually transitioned strategy, combining the transferred models and the new models can achieve these aims. We reproduce in this article versions of our experiments using publicly accessible Lending Club data to allow us to publish the results while still complying with the privacy standards of our client&#x00027;s data.</p>
<p>We focus on two scenarios in which a big dataset of current loan products is used to improve the credit risk model for new loan products with a considerably smaller dataset. In the first scenario, Lending Club data is used to resemble a lender that already possesses (CD) data and is ready to start lending to small companies. The second scenario likewise makes use of Lending Club data to resemble a lender that already has a credit card loan product and now wants to expand into vehicle loans.</p>
<p>For pre-processing, we pick 16 variables as inputs, and the output to be predicted is whether or not a loan will be approved or not. We transform the loan status to a binary result to simplify the model. Here, 1 indicates a defaulted loan, a charged-off loan, or a late loan payment; 0 indicates a paid-off debt. Current loans that are not yet due are not included in this calculation. The details of the pre-processing are detailed in Data Availability Statement.</p>
<p>Furthermore, to adopt a transfer learning approach in the real world, the following two questions need to be addressed. The first issue is to explain transferred models. Many jurisdictions require credit decisions to be explained for anti-discrimination and human rights purposes. For example, the European Union&#x00027;s General Data Protection Regulation (GDPR) requires &#x0201C;meaningful information about the logic involved&#x0201D; in automated decisions, providing an explanation that enables a data subject to exercise their rights under the GDPR and human rights law (Selbst and Powles, <xref ref-type="bibr" rid="B19">2017</xref>). SHAP (Lundberg and Lee, <xref ref-type="bibr" rid="B11">2017</xref>), based on Shapley values, is one of the most popular methods for explaining machine learned models. In this article, we apply SHAP to analyze the contribution of features and the impact of the transfer.</p>
<p>The second question is how to handle the difference between features from source and target domains. For instance, a source domain could be for a small short-term alternative loan, but the target domain may be for large and long-term installment loans. Key features, such as loan amount, loan terms, and interest cover. can differ between these domains. We could use the progressive shifting contribution network proposed in Suryanto et al. (<xref ref-type="bibr" rid="B21">2019</xref>) which combines source and target domain feature learning to improve model accuracy, but a key question that remains is: Can we adapt the features <italic>before</italic> transfer learning to get more accurate models?</p>
<p>We develop an approach to this question in this article based on three perspectives. First, we use a Kolmogorov Smirnov (KS) test to quantify the difference between source and target domains and use domain adaptation to treat only features that differ substantially between the domains before training. Second, after we find candidate features to be adapted, based on their KS differences, we include other features that are highly correlated with the candidate features and test the accuracy of models adapting these feature combinations. Finally, we exclude from adaptation those features related to a borrower&#x00027;s credit history where the adaptation would incorrectly impact the classification.</p>
<p>The remainder of the article is structured as follows. We cover related research in Section 2 and key aspects of the problem in Section 3. We describe our transfer learning approach in detail in Section 4, and in Section 5, we address two critical issues: domain adaptation and explainability. Discussion and conclusion are contained in the final two sections (6 and 7).</p>
</sec>
<sec id="s2">
<title>2. Related Studies</title>
<p>Transfer learning pre-dates deep learning (DL). Given the difficulty of defining features in image processing, many approaches were pioneered in that area. For example, the transfer of parameters from a trained SVM model was proposed by Yang et al. (<xref ref-type="bibr" rid="B26">2007</xref>). This can also be applied to unsupervised learning &#x02014; a domain adaptation technique known as transfer component analysis was given by Pan et al. (<xref ref-type="bibr" rid="B13">2011</xref>). In their survey article, Pan et al. (<xref ref-type="bibr" rid="B14">2010</xref>) suggested four categories for transfer learning: transfer of instances, transfer of feature representations, transfer of parameter values, and transfer of relational knowledge.</p>
<p>Transfer learning requires training on a source domain and (re)training on a target domain, on which the class labels are to be predicted (for classification tasks). In this article, we are mainly concerned with the transfer of feature representations. The work on transductive transfer learning (Pan et al., <xref ref-type="bibr" rid="B13">2011</xref>) has some similarities with our approach, but in that work, both the source and target classification tasks must be the same. A further difference in our approach is that we aim to progressively optimize for the right hyperparameter setting to balance the relative weight of the source to target the transfer of feature representations.</p>
<p>A central issue in transfer learning is the relation of the source to the target domain. This relationship can be affected by the relative heterogeneity of the data in the domain, and issues such as symmetry in the transfer of features, which can also impact the transfer of parameter values and relational knowledge, and the selection of the source domain, as discussed in Weiss et al. (<xref ref-type="bibr" rid="B23">2016</xref>). It is also important to be aware of and mitigate the potential for transfer learning to <italic>decrease</italic> performance on the target domain (which is known as &#x0201C;negative transfer&#x0201D;). Several approaches have been proposed to address this (Gao et al., <xref ref-type="bibr" rid="B6">2008</xref>; Chattopadhyay et al., <xref ref-type="bibr" rid="B3">2011</xref>; Sun et al., <xref ref-type="bibr" rid="B20">2013</xref>; Xiao et al., <xref ref-type="bibr" rid="B25">2014</xref>). However, in this study, we instead focussed on optimizing the architecture of target network models to enable the successful transfer of features. Addressing the risk of negative transfer in our approach is left for future study.</p>
<p>While the terms &#x0201C;transfer learning&#x0201D; and &#x0201C;domain adaptation&#x0201D; have been used interchangeably, we use transfer learning when the focus is the modeling configuration, and domain adaptation when the focus is on transforming the data. There are only a few published studies on domain adaptation for credit risk, e.g., Huang and Chen (<xref ref-type="bibr" rid="B9">2018</xref>) proposed domain adaptation for transforming the data distribution. In other domains, approaches such as Balanced Distribution Adaptation (Wang et al., <xref ref-type="bibr" rid="B22">2017</xref>) and adapting without target labels have been used (Huang and Chen, <xref ref-type="bibr" rid="B9">2018</xref>; Zhang et al., <xref ref-type="bibr" rid="B27">2018</xref>; Kouw and Loog, <xref ref-type="bibr" rid="B10">2019</xref>).</p>
<p>In this article, we use a method of Progressive Network configuration for transfer learning, similar to Rusu et al. (<xref ref-type="bibr" rid="B17">2016</xref>). The contribution of this article is a strategy to apply domain adaptation to the source data when target data with labels is limited and to apply both domain adaptation and transfer learning to credit risk. Neyshabur et al. (<xref ref-type="bibr" rid="B12">2020</xref>) investigated what is being transferred, which are general features in the lower layers and more specific features in the higher layers. Our Progressive Network configuration facilitates the search to find from which layers we retrain the network. In credit risk, where the data labels are scarce, the approach such as transferring learned knowledge from self-supervised tasks to downstream tasks could improve the performance of the network (Han et al., <xref ref-type="bibr" rid="B8">2021</xref>).</p>
</sec>
<sec id="s3">
<title>3. Credit Risk</title>
<p>Across their loan portfolios, lenders strive to maximize the risk-return ratio. The cornerstone of this optimization is the accurate and consistent measurement of credit risk. Expected Loss (<italic>EL</italic>) is a standard metric used by lenders to assess credit risk. <italic>EL</italic> is mostly governed by the Probability of Default (<italic>PD</italic>) in an unsecured loan scenario. <italic>PD</italic> is calculated using credit scoring models. The characteristics of the loan applicant and their application are usually inputs to a credit rating algorithm. To demonstrate our techniques, we use attributes from <ext-link ext-link-type="uri" xlink:href="https://www.lendingclub.com/info/download-data.action">lendingclub.com</ext-link> data. The most often used metrics in credit risk assessment include the Gini Index, KS statistics, Lift, Mahalanobis distance, and information statistics (&#x00158;ez&#x000E1;&#x0010D; and &#x00158;ez&#x000E1;&#x0010D;, <xref ref-type="bibr" rid="B16">2011</xref>). The Gini Index (abbreviated as "Gini") is used in this study.</p>
<p>A score ranging from 0 to 1 is generated by the scoring model. It is a likelihood of default that has been calculated. Typically, a portion of the data is pre-allocated for score calibration. With the resulting <italic>PD</italic> and loan application data as inputs, lenders utilize a set of decision processes and rules to make the best conclusion possible. An eligibility test is usually the first step in the decision-making process. The <italic>PD</italic> is calculated for all potential candidates and then used to divide them into decision groups. The interest rate, e.g., may differ depending on the decision group, as could the loan amount as a proportion of net income.</p>
<p>The focus of this study is on credit rating for unsecured loans. The Area Under Receiver Operating Curve (<italic>AUC</italic>) or <italic>GiniROC</italic>, which is 2<italic>AUC</italic>&#x02212;1 (Flach, <xref ref-type="bibr" rid="B4">2003</xref>), is used to evaluate the performance of our credit scoring model. The Gini Index, which is employed as the splitting criteria in CART (Breiman, <xref ref-type="bibr" rid="B2">1996</xref>), is based on the same principle as GiniROC. Gini and GiniROC, on the other hand, serve distinct purposes. The measure GiniROC is used to assess model quality based on <italic>PD</italic> without the requirement to transform <italic>PD</italic> into binary classifications because the threshold for doing such classifications is established in the credit decisioning.</p>
<sec>
<title>3.1. Credit Scoring</title>
<p>Credit scoring generates a <italic>PD</italic> value that is used to estimate whether a loan will be repaid or defaulted. There are other possible consequences in real-world settings, such as late or incomplete payment. We need a measure to assess the model&#x00027;s quality without setting a threshold to transform the <italic>PD</italic> into a classification in credit scoring. We can use a statistic like <italic>Fscore</italic> after we have these classifications. This judgment is delayed in credit risk to the credit-decisioning process when expert rules/guidelines are used to determine whether or not the loan is accepted.</p>
</sec>
<sec>
<title>3.2. Credit Decisioning</title>
<p>Credit decisioning uses <italic>PD</italic> to accept or deny a loan application. A mapping table is used to map ranges of <italic>PD</italic> to decisions when converting from <italic>PD</italic> to decisions. It may also change the loan amount, interest rate, and period, in addition to approving or declining the loan. Because the data is typically scarce and/or the search space is too big for supervised learning models to be built, this model is typically based on expert rules.</p>
</sec>
</sec>
<sec id="s4">
<title>4. Transfer Learning</title>
<p>In this section, we first outline our transfer learning framework and then summarize the experimental results (derived from those of Suryanto et al., <xref ref-type="bibr" rid="B21">2019</xref>). The key idea of this part of our study is the notion of Progressive Shift Contribution (<italic>PSC</italic>) where a combination of network architecture operators and training methods enable a number of variants of transfer learning to be implemented and evaluated. Essentially, the idea behind <italic>PSC</italic> is to generate different transfer learning models in which the target domain contribution gradually grows while the source domain contribution drops. We first describe a base model network architecture, explained in Section 4.2, then how to create a variety of network topologies from this in Section 4.3.1. We empirically evaluate the effectiveness of this approach to transfer learning algorithms by testing each variant model on different source and target domain datasets, with the results shown in Section 4.4.</p>
<sec>
<title>4.1. Model Development</title>
<p>A total of six different network architectures were developed, which we refer to in the following as Models 1&#x02013;6.</p>
<p>Only source data is used to train Model No. 1. We progressively shifted the domain contribution from source to target domain in Model Nos. 2&#x02013;5. Only target domain data is used to train Model No. 6, the last model. The ratio of trained layers using the target domain to trained layers using the source domain indicates the contribution variations between the source and target domains. This approach can be generalized to a network configuration of any size. The algorithm&#x00027;s specifics will be explored in further depth in Section 4.3.1. All experiments&#x00027; source code and data are accessible in Data Availability Statement.</p>
<p>To find the optimum network configuration, we transferred the <italic>PSC</italic> from the source domain to the target domain and evaluated Gini performance using test data from the target domain. The performance of the model is primarily determined by the following factors:</p>
<p>a) methodologies used for modeling (e.g., generalized linear model, gradient boosting machine, deep learning), hyper parameters<xref ref-type="fn" rid="fn0002"><sup>2</sup></xref>, b) data signal strength, and c) feature engineering. We can express the link between <italic>Gini</italic>, which we denote by &#x1D53E;, and the above factors as follows:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mo>&#x1D53E;</mml:mo><mml:mo>=</mml:mo><mml:mi>&#x1D524;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi mathvariant="-tex-caligraphic">T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D530;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x1D530;<sub>&#x1D522;</sub> represents the source domain&#x00027;s test data, &#x1D510;<sub>&#x1D522;</sub> represents the model trained on the source domain&#x00027;s training data, <inline-formula><mml:math id="M2"><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:math></inline-formula>() represents th activity of testing a model on the test data that produces the test results, and &#x1D524;() represents a function to compute the Gini coefficient of the results. The variable &#x1D510;<sub>&#x1D522;</sub> is defined as follows:</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M3"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>0</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D513;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D531;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x1D510;<sub><italic>0</italic></sub> is a configuration of a deep neural network with initial random weights, &#x1D513;<sub>&#x1D522;</sub> is a set of hyper parameters for training &#x1D510;<sub>&#x1D522;</sub>, &#x1D531;<sub>&#x1D522;</sub> is the source domain&#x00027;s training data, &#x1D509;<sub>&#x1D522;</sub> is a collection of features obtained from &#x1D531;<sub>&#x1D522;</sub>. <inline-formula><mml:math id="M4"><mml:msub><mml:mi mathvariant="-tex-caligraphic">T</mml:mi><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>() is an activity to train a model with these four factors. The outcome of <inline-formula><mml:math id="M5"><mml:msub><mml:mi mathvariant="-tex-caligraphic">T</mml:mi><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>() is a model that has been trained using the above four factors.</p>
<p>We now describe how <italic>PSC</italic> is performed. To begin, we define a function <inline-formula><mml:math id="M6"><mml:mi mathvariant="-tex-caligraphic">S</mml:mi></mml:math></inline-formula>() that partitions &#x1D510;<sub>&#x1D522;</sub> into two segments: &#x1D510;<sub>&#x1D51B;<sub>&#x1D522;</sub></sub> and &#x1D510;<sub>&#x1D509;<sub>&#x1D522;</sub></sub>. &#x1D510;<sub>&#x1D51B;<sub>&#x1D522;</sub></sub> denotes the segment in which the layers were trained using &#x1D531;<sub>&#x1D522;</sub> and the layers are not re-trainable. &#x1D510;<sub>&#x1D509;<sub>&#x1D522;</sub></sub> is the segment in which the layers were also trained using &#x1D531;<sub>&#x1D522;</sub>, however, these layers can be re-trained using the target domain&#x00027;s training data &#x1D531;<sub>&#x1D52B;</sub>.</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D51B;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi mathvariant="-tex-caligraphic">S</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p><inline-formula><mml:math id="M8"><mml:mi mathvariant="-tex-caligraphic">S</mml:mi></mml:math></inline-formula>()&#x00027;s inverse function is <inline-formula><mml:math id="M9"><mml:mi mathvariant="-tex-caligraphic">C</mml:mi></mml:math></inline-formula>(), which combines &#x1D510;<sub>&#x1D51B;<sub>&#x1D522;</sub></sub> and &#x1D510;<sub>&#x1D509;<sub>&#x1D522;</sub></sub> to form &#x1D510;<sub>&#x1D522;</sub>.</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M10"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="-tex-caligraphic">C</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D51B;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>For the target domain &#x1D510;<sub>&#x1D531;<italic>&#x1D52F;&#x1D51E;&#x1D52B;&#x1D530;&#x1D523;&#x1D522;&#x1D52F;</italic></sub>, we developed a model, that incorporates data from both the source and target domains, by transferring the structure and weights of the &#x1D510;<sub>&#x1D51B;<sub>&#x1D522;</sub></sub> layers and retraining the &#x1D510;<sub>&#x1D509;<sub>&#x1D522;</sub></sub> layers.</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M11"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D513;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D531;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In the end, we put together the target model &#x1D510;<sub>&#x1D509;<sub>&#x1D52B;</sub></sub> and the model &#x1D510;<sub>&#x1D51B;<sub>&#x1D522;</sub></sub>. The result is the transferred model, which is called &#x1D510;<sub>&#x1D531;<italic>&#x1D52F;&#x1D51E;&#x1D52B;&#x1D530;&#x1D523;&#x1D522;&#x1D52F;</italic></sub>.</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M12"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D531;</mml:mi><mml:mi>&#x1D52F;</mml:mi><mml:mi>&#x1D51E;</mml:mi><mml:mi>&#x1D52B;</mml:mi><mml:mi>&#x1D530;</mml:mi><mml:mi>&#x1D523;</mml:mi><mml:mi>&#x1D522;</mml:mi><mml:mi>&#x1D52F;</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="-tex-caligraphic">C</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D51B;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The overarching objective is to maximize &#x1D53E;<sub>&#x1D54B;</sub> by tracking &#x1D53E;<sub>&#x1D54B;</sub> as the <italic>PSC</italic> moves from source to target domain data. We evaluate the performance of all six network topologies in Section 4.1 to get the highest &#x1D53E;<sub>&#x1D54B;</sub>. In the following equation, &#x1D530;<sub>&#x1D52B;</sub> represents the target domain&#x00027;s test data.</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M13"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mo>&#x1D53E;</mml:mo></mml:mrow><mml:mrow><mml:mo>&#x1D54B;</mml:mo></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x1D524;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D531;</mml:mi><mml:mi>&#x1D52F;</mml:mi><mml:mi>&#x1D51E;</mml:mi><mml:mi>&#x1D52B;</mml:mi><mml:mi>&#x1D530;</mml:mi><mml:mi>&#x1D523;</mml:mi><mml:mi>&#x1D522;</mml:mi><mml:mi>&#x1D52F;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D530;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
</sec>
<sec>
<title>4.2. The Base Model</title>
<p><xref ref-type="fig" rid="F1">Figure 1</xref> illustrates the initial basic model. It contains 16 input nodes on the input layer, three hidden layers with 32 nodes on each, and one output node on the output layer. The configuration of the network is chosen using a hyper parameter search to get a configuration that is close to optimal. By and large, the base models were constructed in accordance with network architectures.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Network u: the base model.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-868232-g0001.tif"/>
</fig>
</sec>
<sec>
<title>4.3. The Comparison Model (Model u)</title>
<p><xref ref-type="fig" rid="F1">Figure 1</xref> illustrates the network configuration u on which the comparison model is based. Only the target domain data is used to train this model, no data from the source domain is utilized. The model developed using target domain data is defined similarly to Equation 2:</p>
<disp-formula id="E8"><label>(8)</label><mml:math id="M14"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>&#x1D510;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D532;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D510;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D532;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>0</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D513;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D531;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x1D510;(&#x1D532;)<sub>&#x1D52B;</sub> is a model based on network configuration u that was built using data from the target domain. The starting model &#x1D510;(&#x1D532;)<sub><italic>0</italic></sub> is based on network configuration u with all weights randomly initialized, and the parameters, training data, and features required to build the model &#x1D510;(&#x1D532;)<sub>&#x1D52B;</sub> are &#x1D513;<sub>&#x1D52B;</sub>, &#x1D531;<sub>&#x1D52B;</sub>, and &#x1D509;<sub>&#x1D52B;</sub>.</p>
<disp-formula id="E9"><label>(9)</label><mml:math id="M15"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mo>&#x1D53E;</mml:mo><mml:mo>=</mml:mo><mml:mi>&#x1D524;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D510;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D532;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D530;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In the above equation, <italic>s</italic><sub><italic>n</italic></sub> denotes the target domain&#x00027;s test data.</p>
<sec>
<title>4.3.1. PSC Models</title>
<p>Six models were introduced in Section 4.1, where the <italic>PSC</italic> changes between the source and target domain data. We added an extra parameter to the split function provided in Equation 3 to determine the fraction of <italic>PSC</italic>. This parameter&#x00027;s value is one of the following: <italic>v</italic>, <italic>N</italic><sub>1</sub>, <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub>, <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub><italic>N</italic><sub>3</sub>, or <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub><italic>N</italic><sub>3</sub><italic>N</italic><sub>4</sub>. Based on the varied <italic>PSC</italic> from the source to the target domain, each value results in a distinct network configuration. For each of these five values, we created five <italic>PSC</italic> models. In addition, we include the baseline Comparison Model mentioned in Section 4.3. In the next sections, we go over Models 2 through 6.</p>
</sec>
<sec>
<title>4.3.2. Model <italic>v</italic></title>
<p>Model <italic>v</italic> is exclusively constructed from the source domain data only. To generate Model <italic>v</italic>, we first trained model &#x1D510;(&#x1D533;)<sub>&#x1D522;</sub> using Equation 10 and the network configuration shown in <xref ref-type="fig" rid="F2">Figure 2</xref>.</p>
<disp-formula id="E11"><label>(10)</label><mml:math id="M16"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>&#x1D510;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D533;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D510;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D533;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>0</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D513;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D531;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>On target domain data, the model was evaluated, and a Gini coefficient was computed.</p>
<disp-formula id="E12"><label>(11)</label><mml:math id="M17"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mo>&#x1D53E;</mml:mo><mml:mo>=</mml:mo><mml:mi>&#x1D524;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D510;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D533;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D530;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Network v.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-868232-g0002.tif"/>
</fig>
</sec>
<sec>
<title>4.3.3. Model <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub><italic>N</italic><sub>3</sub><italic>N</italic><sub>4</sub></title>
<p>Four parallel networks were used to develop this model, each with three hidden layers connected to the input and output layers. To begin, we replicated the hidden layers of network v into networks <italic>N</italic><sub>1</sub>, <italic>N</italic><sub>2</sub>, <italic>N</italic><sub>3</sub>, and <italic>N</italic><sub>4</sub>. We explain the transformation conceptually using Equation 12.</p>
<disp-formula id="E14"><label>(12)</label><mml:math id="M18"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>&#x1D510;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="italic"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>1</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>2</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>3</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>4</mml:mi></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="-tex-caligraphic">T</mml:mi><mml:mi mathvariant="-tex-caligraphic">R</mml:mi><mml:mi mathvariant="-tex-caligraphic">A</mml:mi><mml:mi mathvariant="-tex-caligraphic">N</mml:mi><mml:mi mathvariant="-tex-caligraphic">S</mml:mi><mml:mi mathvariant="-tex-caligraphic">F</mml:mi><mml:mi mathvariant="-tex-caligraphic">O</mml:mi><mml:mi mathvariant="-tex-caligraphic">R</mml:mi><mml:mi mathvariant="-tex-caligraphic">M</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D510;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D533;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The networks <italic>N</italic><sub>1</sub>, <italic>N</italic><sub>2</sub>, <italic>N</italic><sub>3</sub>, and <italic>N</italic><sub>4</sub> were configured in the manner depicted in <xref ref-type="fig" rid="F3">Figure 3</xref>.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Network <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub><italic>N</italic><sub>3</sub><italic>N</italic><sub>4</sub>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-868232-g0003.tif"/>
</fig>
<p>After the structure and weights were set (as shown in <xref ref-type="fig" rid="F3">Figure 3</xref>), we then set the following as trainable, using the target domain data: The 3rd hidden layer of Network <italic>N</italic><sub>2</sub>, the 2nd and 3rd hidden layers of Network <italic>N</italic><sub>3</sub>, and all hidden layers of Network <italic>N</italic><sub>4</sub>. The next three steps are indicated in numbers 1, 2, 3 within ellipses in <xref ref-type="fig" rid="F3">Figure 3</xref>:</p>
<p>Following the establishment of the structure and weights (as seen in <xref ref-type="fig" rid="F3">Figure 3</xref>), we designated the following hidden layers as trainable using the target domain data: <italic>N</italic><sub>2</sub><italic>HL</italic><sub>3</sub>, <italic>N</italic><sub>3</sub><italic>HL</italic><sub>2</sub>, <italic>N</italic><sub>3</sub><italic>HL</italic><sub>3</sub>, <italic>N</italic><sub>4</sub><italic>HL</italic><sub>1</sub>, <italic>N</italic><sub>4</sub><italic>HL</italic><sub>2</sub>, and <italic>N</italic><sub>4</sub><italic>HL</italic><sub>3</sub>. The next three steps are as below:</p>
<list list-type="order">
<list-item><p>The source domain&#x00027;s training data (&#x1D531;<sub>&#x1D522;</sub>) is used to derive weights for all four networks (<italic>N</italic><sub>1</sub>, <italic>N</italic><sub>2</sub>, <italic>N</italic><sub>3</sub>, and <italic>N</italic><sub>4</sub>). Depending on the configuration, some hidden layers in networks <italic>N</italic><sub>1</sub>, <italic>N</italic><sub>2</sub>, <italic>N</italic><sub>3</sub>, and <italic>N</italic><sub>4</sub> and the output layer are configured to be re-trainable, using the target domain&#x00027;s training data (&#x1D531;<sub>&#x1D52B;</sub>).</p></list-item>
<list-item><p>Train re-trainable layers using the target domain&#x00027;s training data (&#x1D531;<sub>&#x1D52B;</sub>).</p></list-item>
<list-item><p>Evaluate the whole parallel network (<italic>N</italic><sub>1</sub>, <italic>N</italic><sub>2</sub>, <italic>N</italic><sub>3</sub>, and <italic>N</italic><sub>4</sub>) using the target domain&#x00027;s test data (&#x1D530;<sub>&#x1D52B;</sub>), and then compute the Gini coefficient from the result.</p></list-item>
</list>
<p>The following three equations can describe the evolution of Model <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub><italic>N</italic><sub>3</sub><italic>N</italic><sub>4</sub>:</p>
<disp-formula id="E15"><label>(13)</label><mml:math id="M19"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D51B;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="italic"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>1</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>2</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>3</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>4</mml:mi></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="italic"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>1</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>2</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>3</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>4</mml:mi></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi mathvariant="-tex-caligraphic">S</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D510;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="italic"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>1</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>2</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>3</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>4</mml:mi></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E16"><label>(14)</label><mml:math id="M20"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="italic"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>1</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>2</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>3</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>4</mml:mi></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="italic"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>1</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>2</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>3</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>4</mml:mi></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D513;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D531;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E17"><label>(15)</label><mml:math id="M21"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>&#x1D510;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="italic"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>1</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>2</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>3</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>4</mml:mi></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="-tex-caligraphic">C</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D51B;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="italic"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>1</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>2</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>3</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>4</mml:mi></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="italic"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>1</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>2</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>3</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>4</mml:mi></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Model <italic>N</italic><sub><italic>1</italic></sub><italic>N</italic><sub><italic>2</italic></sub><italic>N</italic><sub><italic>3</italic></sub><italic>N</italic><sub><italic>4</italic></sub> was trained using the source domain data and six hidden layers and the output layer was retrained using the target domain data.</p>
</sec>
<sec>
<title>4.3.4. Model <italic>N</italic><sub>1</sub></title>
<p>After deleting Networks <italic>N</italic><sub>2</sub>, <italic>N</italic><sub>3</sub>, and <italic>N</italic><sub>4</sub> from the model <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub><italic>N</italic><sub>3</sub><italic>N</italic><sub>4</sub>, we derive Model <italic>N</italic><sub>1</sub>. <xref ref-type="fig" rid="F4">Figure 4</xref> illustrates this network configuration. All hidden layers were trained on source domain data and only the output layer was retrained on target domain data in Model <italic>N</italic><sub>1</sub>. Equations 16, 17, and 18 illustrate the evolution of Model <italic>N</italic><sub>1</sub>.</p>
<disp-formula id="E18"><label>(16)</label><mml:math id="M22"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D51B;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi mathvariant="-tex-caligraphic">S</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D510;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E19"><label>(17)</label><mml:math id="M23"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D513;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D531;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E20"><label>(18)</label><mml:math id="M24"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>&#x1D510;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="-tex-caligraphic">C</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D51B;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Network <italic>N</italic><sub>1</sub> and Network <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-868232-g0004.tif"/>
</fig>
</sec>
<sec>
<title>4.3.5. Model <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub></title>
<p>Model <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub> is derived from Model <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub><italic>N</italic><sub>3</sub><italic>N</italic><sub>4</sub> by excluding Networks <italic>N</italic><sub>3</sub> and <italic>N</italic><sub>4</sub>. <xref ref-type="fig" rid="F4">Figure 4</xref> illustrates this network configuration. All hidden layers were trained on the source domain data in Model <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub>. The target domain data was used to retrain one hidden layer (<italic>N</italic><sub>2</sub><italic>HL</italic><sub>3</sub>) and the output layer. The evolution of the model <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub> is illustrated in Equations 19, 20, and 21.</p>
<disp-formula id="E21"><label>(19)</label><mml:math id="M25"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D51B;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi mathvariant="-tex-caligraphic">S</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D510;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E22"><label>(20)</label><mml:math id="M26"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D513;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D531;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E23"><label>(21)</label><mml:math id="M27"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>&#x1D510;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="-tex-caligraphic">C</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D51B;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
</sec>
<sec>
<title>4.3.6. Model <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub><italic>N</italic><sub>3</sub></title>
<p>Model <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub><italic>N</italic><sub>3</sub> is derived from Model <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub><italic>N</italic><sub>3</sub><italic>N</italic><sub>4</sub>, excluding Network <italic>N</italic><sub>4</sub>. <xref ref-type="fig" rid="F5">Figure 5</xref> illustrates this network configuration. All hidden layers were trained in Model <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub><italic>N</italic><sub>3</sub> using source domain data. The target domain data was used to retrain three hidden layers (<italic>N</italic><sub>2</sub><italic>HL</italic><sub>3</sub>, <italic>N</italic><sub>3</sub><italic>HL</italic><sub>2</sub>, and <italic>N</italic><sub>3</sub><italic>HL</italic><sub>3</sub>) and the output layer. Equations 22, 23, and 24 illustrate the evolution of the model <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub><italic>N</italic><sub>3</sub>.</p>
<disp-formula id="E24"><label>(22)</label><mml:math id="M28"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D51B;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi mathvariant="-tex-caligraphic">S</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D510;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E25"><label>(23)</label><mml:math id="M29"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D513;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D531;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D52B;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E26"><label>(24)</label><mml:math id="M30"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>&#x1D510;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>s</mml:mi><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="-tex-caligraphic">C</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D51B;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Network <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub><italic>N</italic><sub>3</sub>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-868232-g0005.tif"/>
</fig>
</sec>
</sec>
<sec>
<title>4.4. Experiments</title>
<p>We utilized data from <ext-link ext-link-type="uri" xlink:href="https://www.lendingclub.com">lendingclub.com</ext-link> in our studies, which is similar to our original client&#x00027;s (confidential) data, and spans the years 2007&#x02013;2018 (refer to Data Availability Statement). We began by developing base models and training them without using transfer learning. A grid search was then used to identify a close-to-optimal set of hyper-parameters for each neural network architecture. To evaluate our transfer learning approach, we selected data on <italic>CD</italic> as the source domain and Small Business (<italic>SB</italic>) as the target domain. Our objective is to evaluate the implementation of transfer learning from <italic>CD</italic> to <italic>SB</italic>, to exemplify a real-world situation where data in the target domain (here <italic>SB</italic>) is scarce, but data in the source domain (here <italic>CD</italic>) is more readily available.</p>
<p>To validate the performance of <italic>DL</italic> on the <italic>CD</italic> and <italic>SB</italic> datasets, we also compared it to that of Gradient Boosting Machines (<italic>GBM</italic>). The Gini coefficients on the <italic>CD</italic> data were equal (0.43, with a s.d. of 0.01). On the <italic>SB</italic> data, they were almost identical, at 0.30 with 0.05 s.d. for <italic>GBM</italic> compared to 0.31 with 0.02 s.d. for <italic>DL</italic>, indicating that there is no statistically significant difference in performance between the two methods. Consequently, in the following sections, our experiments concentrate exclusively on <italic>DL</italic>.</p>
<sec>
<title>4.4.1. Datasets</title>
<p>From the data acquired from <ext-link ext-link-type="uri" xlink:href="https://www.lendingclub.com">lendingclub.com</ext-link>, ten datasets were generated. For the first source domain dataset, <italic>CD</italic>4, we randomly selected 1,00,000 records from 9,40,948 records on loans for the purpose of paying off credit cards and consolidating debt. The rate of bad debt in this sample was 21%. The first target domain dataset, <italic>SB</italic>4, had 13,794 records pertaining to loans made for the purpose of investing in small businesses. This is a more risky loan type, where the default rate is 30%. These two datasets were not subjected to any outlier screening.</p>
<p>We then constructed the source domain datasets <italic>CD</italic>1, <italic>CD</italic>2, and <italic>CD</italic>3 as subsets of the dataset <italic>CD</italic>4, with different time range filters applied. Target domain datasets <italic>SB</italic>1, <italic>SB</italic>2, and <italic>SB</italic>3 were also constructed as subsets of <italic>SB</italic>4 similarly to those for the source domain. These restrictions allow investigation of any effects due to changes in the time range of the datasets used.</p>
<p>Additionally, we identified a further pair of source and target domain datasets, as follows. The source domain <italic>CCD</italic> is a subset of the <italic>CD</italic>4 dataset that was filtered to extract Credit Card Loans. Similarly, the Car Loan data target domain dataset was a set of car loans extracted from Lending Club datasets. Both these datasets both spanned the time range 2007&#x02013;2018, as in <italic>CD</italic>4 and <italic>SB</italic>4 (for further details refer to the link provided in Data Availability Statement).</p>
<p>The data in <xref ref-type="table" rid="T1">Table 1</xref> was used in all of the experiments. They were performed using a 10-fold cross-validation procedure that was done five times. The basic model for the transfer learning was created using the datasets <italic>CD</italic>1, <italic>CD</italic>2, <italic>CD</italic>3, <italic>CD</italic>4, and <italic>CCD</italic>, and the network configuration <italic>u</italic>, as illustrated in <xref ref-type="fig" rid="F1">Figure 1</xref> and specified in Equation 2. The intensity of the signal associated with the result being predicted from the data was one element that affected model performance.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>The datasets used in the transfer learning studies are listed below.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>ID</bold></th>
<th valign="top" align="left"><bold>Dataset</bold></th>
<th valign="top" align="left"><bold>Period</bold></th>
<th valign="top" align="left"><bold>Size</bold></th>
<th valign="top" align="left"><bold>Type</bold></th>
<th valign="top" align="left"><bold>Gini (s.d.)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">CD1</td>
<td valign="top" align="left"><italic>CreditCard</italic>/<italic>DebtConsolidation</italic></td>
<td valign="top" align="left">2007&#x02013;2011</td>
<td valign="top" align="left">23,813</td>
<td valign="top" align="left">Source</td>
<td valign="top" align="left">0.364(0.023)</td>
</tr>
<tr>
<td valign="top" align="left">SB1</td>
<td valign="top" align="left"><italic>SmallBusinessLoan</italic></td>
<td valign="top" align="left">2007&#x02013;2011</td>
<td valign="top" align="left">1,831</td>
<td valign="top" align="left">Target</td>
<td valign="top" align="left">0.272(0.067)</td>
</tr>
<tr>
<td valign="top" align="left">CD2</td>
<td valign="top" align="left"><italic>CreditCard</italic>/<italic>DebtConsolidation</italic></td>
<td valign="top" align="left">2007&#x02013;2014</td>
<td valign="top" align="left">100,000</td>
<td valign="top" align="left">Source</td>
<td valign="top" align="left">0.417(0.016)</td>
</tr>
<tr>
<td valign="top" align="left">SB2</td>
<td valign="top" align="left"><italic>SmallBusinessLoan</italic></td>
<td valign="top" align="left">2007&#x02013;2014</td>
<td valign="top" align="left">6,686</td>
<td valign="top" align="left">Target</td>
<td valign="top" align="left">0.274(0.040)</td>
</tr>
<tr>
<td valign="top" align="left">CD3</td>
<td valign="top" align="left"><italic>CreditCard</italic>/<italic>DebtConsolidation</italic></td>
<td valign="top" align="left">2007&#x02013;2016</td>
<td valign="top" align="left">100,000</td>
<td valign="top" align="left">Source</td>
<td valign="top" align="left">0.447(0.013)</td>
</tr>
<tr>
<td valign="top" align="left">SB3</td>
<td valign="top" align="left"><italic>SmallBusinessLoan</italic></td>
<td valign="top" align="left">2007&#x02013;2016</td>
<td valign="top" align="left">12,114</td>
<td valign="top" align="left">Target</td>
<td valign="top" align="left">0.331(0.032)</td>
</tr>
<tr>
<td valign="top" align="left">CD4</td>
<td valign="top" align="left"><italic>CreditCard</italic>/<italic>DebtConsolidation</italic></td>
<td valign="top" align="left">2007&#x02013;2018</td>
<td valign="top" align="left">100,000</td>
<td valign="top" align="left">Source</td>
<td valign="top" align="left">0.448(0.012)</td>
</tr>
<tr>
<td valign="top" align="left">SB4</td>
<td valign="top" align="left"><italic>SmallBusinessLoan</italic></td>
<td valign="top" align="left">2007&#x02013;2018</td>
<td valign="top" align="left">13,794</td>
<td valign="top" align="left">Target</td>
<td valign="top" align="left">0.351(0.024)</td>
</tr>
<tr>
<td valign="top" align="left">CCD</td>
<td valign="top" align="left"><italic>CreditCard</italic></td>
<td valign="top" align="left">2007&#x02013;2018</td>
<td valign="top" align="left">100,000</td>
<td valign="top" align="left">Source</td>
<td valign="top" align="left">0.463(0.014)</td>
</tr>
<tr>
<td valign="top" align="left">CAR</td>
<td valign="top" align="left"><italic>CarLoan</italic></td>
<td valign="top" align="left">2007&#x02013;2018</td>
<td valign="top" align="left">12,734</td>
<td valign="top" align="left">Target</td>
<td valign="top" align="left">0.436(0.036)</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The &#x0201C;Type&#x0201D; column shows whether the dataset is used as the source or target for the transfer learning process; the results given are cross-validation runs&#x00027; means and standard deviations (s.d.)</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>As seen in <xref ref-type="table" rid="T1">Table 1</xref>, there is an effect of both dataset size (on the target domain, the larger the dataset, and the higher the Gini) and the date range (on the source domain, the more time covered by the dataset, the higher the Gini) on the baseline performance.</p>
<p>As in Equation 9, the Gini values for <italic>SB</italic>1, <italic>SB</italic>2, <italic>SB</italic>3, <italic>SB</italic>4, and <italic>CAR</italic> (as shown in <xref ref-type="table" rid="T1">Tables 1</xref>, <xref ref-type="table" rid="T2">2</xref>) are computed from the test result of model <italic>M</italic>(<italic>u</italic>)<sub><italic>n</italic></sub> by applying the function <italic>g</italic>() to the test results (for source domain datasets we use model <italic>v</italic>).</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Experimental Results for six models with progressively shifted contribution, built on the source and target datasets described in <xref ref-type="table" rid="T1">Table 1</xref> (all of the source:Credit Card/Debt Consolidation (CD), target:Small Business (SB) Loan datasets, plus the source:Credit Card, target:Car Loan datasets); results shown as <italic>means (s.d.)</italic>; models with the highest performance in each column are denoted by the symbol &#x0002A;.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th/>
<th valign="top" align="center" colspan="5"><bold>Source/Target</bold></th>
</tr>
</thead>
<tbody>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="center"><bold>Model</bold></td>
<td valign="top" align="left"><bold><italic>CD</italic>1/<italic>SB</italic>1</bold></td>
<td valign="top" align="left"><bold><italic>CD</italic>2/<italic>SB</italic>2</bold></td>
<td valign="top" align="left"><bold><italic>CD</italic>3/<italic>SB</italic>3</bold></td>
<td valign="top" align="left"><bold><italic>CD</italic>4/<italic>SB</italic>4</bold></td>
<td valign="top" align="left"><bold><italic>CCD</italic>/<italic>CAR</italic></bold></td>
</tr> <tr>
<td valign="top" align="left"><italic>M</italic>(<italic>v</italic>)<sub><italic>e</italic></sub></td>
<td valign="top" align="left">0.157(0.022)</td>
<td valign="top" align="left">0.236(0.051)</td>
<td valign="top" align="left">&#x02212;0.191(0.260)</td>
<td valign="top" align="left">0.196(0.026)</td>
<td valign="top" align="left">0.262(0.355)</td>
</tr>
<tr>
<td valign="top" align="left"><italic>M</italic>(<italic>N</italic>1)<sub><italic>transfer</italic></sub></td>
<td valign="top" align="left">&#x0002A;0.301(0.097)</td>
<td valign="top" align="left">&#x0002A;0.287(0.051)</td>
<td valign="top" align="left">0.334(0.029)</td>
<td valign="top" align="left">0.350(0.029)</td>
<td valign="top" align="left">&#x0002A;0.447(0.037)</td>
</tr>
<tr>
<td valign="top" align="left"><italic>M</italic>(<sub><italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub>)<italic>transfer</italic></sub></td>
<td valign="top" align="left">0.292(0.091)</td>
<td valign="top" align="left">0.272(0.054)</td>
<td valign="top" align="left">&#x0002A;0.337(0.030)</td>
<td valign="top" align="left">0.350(0.028)</td>
<td valign="top" align="left">0.434(0.035)</td>
</tr>
<tr>
<td valign="top" align="left"><italic>M</italic>(<sub><italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub><italic>N</italic><sub>3</sub>)<italic>transfer</italic></sub></td>
<td valign="top" align="left">0.230(0.087)</td>
<td valign="top" align="left">0.217(0.057)</td>
<td valign="top" align="left">0.300(0.032)</td>
<td valign="top" align="left">0.310(0.030)</td>
<td valign="top" align="left">0.376(0.040)</td>
</tr>
<tr>
<td valign="top" align="left"><italic>M</italic>(<sub><italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub><italic>N</italic><sub>3</sub><italic>N</italic><sub>4</sub>)<italic>transfer</italic></sub></td>
<td valign="top" align="left">0.174(0.010)</td>
<td valign="top" align="left">0.172(0.051)</td>
<td valign="top" align="left">0.254(0.029)</td>
<td valign="top" align="left">0.273(0.030)</td>
<td valign="top" align="left">0.310(0.050)</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left"><italic>M</italic>(<italic>u</italic>)<sub><italic>n</italic></sub></td>
<td valign="top" align="left">0.272(0.067)</td>
<td valign="top" align="left">0.274(0.040)</td>
<td valign="top" align="left">0.331(0.032)</td>
<td valign="top" align="left">&#x0002A;0.351(0.024)</td>
<td valign="top" align="left">0.436(0.036)</td>
</tr> <tr>
<td valign="top" align="left">% improvement</td>
<td valign="top" align="left">10.7%</td>
<td valign="top" align="left">4.7%</td>
<td valign="top" align="left">1.8%</td>
<td valign="top" align="left">0.0%</td>
<td valign="top" align="left">2.5%</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>4.4.2. Results</title>
<p>The experimental findings are reported in <xref ref-type="table" rid="T2">Table 2</xref>, in which we applied Progressive Shifted Contributions (<italic>PSC</italic>), moving from the source domain to the target domain data. All of these experiments were carried out using a 10-fold cross-validation procedure that was repeated five times. In other words, each experiment was performed 50 times, and the mean of the Gini scores and the standard deviation (denoted s.d. in tables) were recorded.</p>
<p>For the experiment over source/target <italic>CD</italic>1/<italic>SB</italic>1, we see that <italic>M</italic>(<sub><italic>N</italic><sub>1</sub>)<italic>transfer</italic></sub> had the greatest Gini of 0.301, an improvement of 10.7% over <italic>M</italic>(<italic>u</italic>)<sub><italic>n</italic></sub>. As the <italic>CD</italic>2/<italic>SB</italic>2 target data covers a wider time range and has more examples, <italic>M</italic>(<sub><italic>N</italic><sub>1</sub>)<italic>transfer</italic></sub> retained the highest Gini of 0.287, but the improvement was just 4.7 percent. As <italic>CD</italic>3/<italic>SB</italic>3 target domain time range increased further along with further increases in dataset size, the contribution moved toward the target domain; model <italic>M</italic>(<sub><italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub>)<italic>transfer</italic></sub> had the greatest Gini of 0.337, however, with just a minor improvement of 1.8% above <italic>M</italic>(<italic>u</italic>)<sub><italic>n</italic></sub>. Finally, the contribution was totally moved toward the target domain in <italic>CD</italic>4/<italic>SB</italic>4, resulting in the highest Gini coefficient of 0.351 for <italic>M</italic>(<italic>u</italic>)<sub><italic>n</italic></sub>.</p>
<p>From the experiments conducted on the datasets <italic>CCD</italic> and <italic>CAR</italic>, <italic>M</italic>(<sub><italic>N</italic><sub>1</sub>)<italic>transfer</italic></sub> produced the highest Gini score of 0.447, representing a slight improvement (2.5%) over the previous experiment, <italic>M</italic>(<italic>u</italic>)<sub><italic>n</italic></sub>. This may be due to a greater similarity between source and target domains for <italic>CCD</italic>/<italic>CAR</italic> than that for <italic>CD</italic>/<italic>SB</italic>, which could be investigated as part of further study. However, it also demonstrates that the progression from the source to target data finding the best balance improves the quality of transfer learning.</p>
<p>From the experimental results, it is evident that the effect of progression from source to target has more impact as data is selectedfrom a wider time range in both the source and target domains. The contribution was dependent on the number of re-trainable layers.</p>
</sec>
<sec>
<title>4.4.3. Additional Experiments</title>
<p>We explored the possibility that the Gini improvement was attributable to the complexity of the network&#x00027;s structure. We ran experiments in accordance with Equations 25 and 26. On source domain data, the model with network configuration <italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub><italic>N</italic><sub>3</sub><italic>N</italic><sub>4</sub> was trained and retrained. This model&#x00027;s performance was 0.39 (0.01), which was lower than the base model Gini of 0.43 (0.01). This demonstrates that the added complexity of <italic>M</italic>(<sub><italic>N</italic><sub>1</sub><italic>N</italic><sub>2</sub><italic>N</italic><sub>3</sub><italic>N</italic><sub>4</sub>)<italic>transfer</italic></sub> has no effect on Gini performance. As a result, the improvement must be attributable to the diversity of the source data, which supplements the target data.</p>
<disp-formula id="E27"><label>(25)</label><mml:math id="M31"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="-tex-caligraphic">T</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">N</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D513;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D531;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D522;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E28"><label>(26)</label><mml:math id="M32"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>&#x1D510;</mml:mi><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="-tex-caligraphic">C</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D51B;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D510;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x1D509;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
</sec>
<sec>
<title>4.4.4. Transfer Learning: Summary</title>
<p>We have presented an algorithm for gradually shifting the contribution from the source domain to the target domain. We can assess incremental complements of target domain data with source domain data using the <italic>PSC</italic> algorithm. While certain tasks were done manually, the overarching purpose was to create a framework that can automatically search for the optimum balance of source and target domain data, resulting in the highest Gini score for that combination. Models ranging from Model <italic>v</italic> (using just source domain data) to Model <italic>u</italic> (using only target domain data) were created. These empirical results have shown that the method of transfer learning developed in this article can be applied in the area of model-based credit risk assessment. Furthermore, the results demonstrate that our method is capable of optimizing over a structured series of network architectures to find the best balance between the contribution of source and target domains.</p>
</sec>
</sec>
</sec>
<sec id="s5">
<title>5. Domain Adaptation and Explainability</title>
<p>In this section, we address the two important questions of domain adaptation and explainability.</p>
<sec>
<title>5.1. Explainability</title>
<p>We use SHapley Additive exPlanations (SHAP)<xref ref-type="fn" rid="fn0003"><sup>3</sup></xref> to calculate the feature contributions for source models, target models, and transferred models. Source models are trained and tested on source domain data, using the same neural network configuration as the target models. The sole purpose of the source models is to investigate the feature contributions of the source domain to compare with transferred models.</p>
<p>We start by constructing a neural network model using training data, then feed this model and the training data into SHAP to create a SHAP-explainer model. We then run test data through the SHAP-explainer model to generate a SHAP contribution value for each input feature. We run this experiment with a 10-fold cross-validation setup and calculate the average SHAP contribution value. The average contribution value of each feature, for all features and models, was recorded (due to space limitations, only one summary of these is shown here, as a stacked bar chart in <xref ref-type="fig" rid="F6">Figure 6</xref>).</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Feature contribution comparison for transferred, target, and source models using SHAP-transferring from CD to MD.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-05-868232-g0006.tif"/>
</fig>
</sec>
<sec>
<title>5.2. Domain Differences</title>
<p>To understand the differences between source and target domains, we use KS tests to quantify the difference for each feature. The KS test can be used to compare two samples without making an assumption about the distribution of data. The null hypothesis is that the two samples, source and target data, come from the same distribution. The KS test produces a KS-statistic and <italic>p</italic>-value. The KS-statistic represents the maximum distance between the source data and the target data distributions. The <italic>p</italic>-value represents the significance level, e.g., less than 0.05. We used the maximum distance between the source data and the target data distribution curves (KS-statistic) to provide insights into the differences in features between these two domains.</p>
</sec>
<sec>
<title>5.3. Domain Adaptation</title>
<p>Domain adaptation aims to transform the source data distribution to be similar to the target data distribution. The intention is to use latent features constructed using source data to complement the target data. We propose the following approach to adapt the feature distribution of the source data to mimic the feature distribution of target data. For each feature, the adaptation steps are:</p>
<list list-type="order">
<list-item><p>Group the source data and the target data in <italic>N</italic> quantiles, where <italic>N</italic> should be selected to ensure that we have sufficient data for each quantile, e.g., more than 50 samples. In this study, we selected <italic>N</italic> &#x0003D; 10, after experimenting with various <italic>N</italic> values.</p></list-item>
<list-item><p>For each corresponding source and target quantiles, we calculate <italic>scale</italic> and then adapt/adjust the source feature values:
<disp-formula id="E29"><label>(27)</label><mml:math id="M33"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mtext mathvariant="italic">scale</mml:mtext><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="italic">max</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="italic">target_value</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mtext mathvariant="italic">min</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="italic">target_value</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="italic">max</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="italic">source_value</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mtext mathvariant="italic">min</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="italic">source_value</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E30"><label>(28)</label><mml:math id="M34"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mstyle mathvariant="italic"><mml:mtext mathvariant="italic">offset</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="italic">source_value</mml:mtext><mml:mo>-</mml:mo><mml:mtext mathvariant="italic">min</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="italic">source_value</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002A;</mml:mo><mml:mtext mathvariant="italic">scale</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E31"><label>(29)</label><mml:math id="M35"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mtext mathvariant="italic">adapted_source_value</mml:mtext><mml:mo>=</mml:mo><mml:mtext mathvariant="italic">min</mml:mtext><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="italic">target_value</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mstyle mathvariant="italic"><mml:mtext mathvariant="italic">offset</mml:mtext></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p></list-item>
</list>
<p>The adapted source feature values are used to initially train the neural network before the last layers are retrained using the target features.</p>
<p>Based on observations of explainer models and feature differences, we adapted different sets of features before training, and then trained and tested using the method described in Section 4 on transfer learning. We then compared the performance of models (using AUC) with different adaptation sets, and transferred models without adaptation. We found that adapting all features significantly reduces accuracy, so we tried different combinations of features to adapt to find the most accurate adapted models.</p>
</sec>
<sec>
<title>5.4. Experiments and Results</title>
<sec>
<title>5.4.1. Transfer Learning and Explainability</title>
<p><xref ref-type="table" rid="T3">Table 3</xref> shows the AUC comparison for the target and the transferred models. The accuracy of the transferred models was better than the target models; AUC improved by 0.042 or 7% for the MD domain, and 0.0224 or 3.6% for the SB domain, respectively. This is in line with the results of Suryanto et al. (<xref ref-type="bibr" rid="B21">2019</xref>).</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Target model vs. transferred model.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Domain</bold></th>
<th valign="top" align="center"><bold>Experiment</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold>Improvement</bold></th>
<th valign="top" align="center"><bold><italic>p</italic>-value</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">CD to MD</td>
<td valign="top" align="left">Training using Target only</td>
<td valign="top" align="center">0.5971(0.0823)</td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">CD to MD</td>
<td valign="top" align="left">Training using Source then retraining the last layer using Target</td>
<td valign="top" align="left">0.6391(0.0856)</td>
<td valign="top" align="center">0.0420 (7.0%)</td>
<td valign="top" align="center">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">CD to SB</td>
<td valign="top" align="left">Training using Target only</td>
<td valign="top" align="center">0.6194(0.0456)</td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">CD to SB</td>
<td valign="top" align="left">Training using Source then retraining the last layer using Target</td>
<td valign="top" align="center">0.6419(0.0509)</td>
<td valign="top" align="center">0.0224 (3.6%)</td>
<td valign="top" align="center">&#x0003C;0.01</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>More interestingly, <xref ref-type="fig" rid="F6">Figure 6</xref> shows a comparison of feature contributions from source, target, and transferred models. For the MD (and SB, not shown) domains, &#x0201C;cover&#x0201D; becomes the top contributing feature of the more accurate transferred models. However, it was the least contributing feature for both the target and source models for CD to MD transfer and a weak contributor in the SB target model. All other features contributed less in the transferred model than in the target model, no matter how much they contributed in the source model for CD to MD transfer. It was similar for CD to SB transfer except for interest rate, which was one of the top contributors in the source model (CD domain). These explainer models indicate that transfer learning improves accuracy by boosting the contribution of weak features in the target domain.</p>
<p>To understand the contribution of &#x0201C;cover,&#x0201D; we investigated the difference in the value distribution for &#x0201C;cover&#x0201D; between source and target domains. Comparing the difference in value distribution between source and target with KS statistics we observed that source CD compared to target MD has a larger (and significant) difference than between source CD and target SB. This difference is larger for small values of cover (&#x0003C; 20) in both distribution comparisons but is more obvious for the CD to MD source&#x02013;target difference where the source distribution has over twice as many loans as the target for the smallest value of cover. This difference is much less for the CD to SB source&#x02013;target difference, particularly for larger values of cover. Further results are presented in the following section.</p>
</sec>
<sec>
<title>5.4.2. Domain Difference</title>
<p><xref ref-type="table" rid="T4">Table 4</xref> lists the KS statistics for CD vs. MD as well as CD vs. SB. It shows that some features were very different between source and target domains but some were similar. It also shows that different pairings of source and target domains had different patterns in feature differences. For example, &#x0201C;cover&#x0201D; was very different between CD and MD with a KS-statistic of 0.2729, but similar between CD and SB with a KS-statistic of 0.0502.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Kolmogorov-Smirnov (KS) of input features.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th/>
<th/>
<th valign="top" align="center" colspan="2"><bold>CD vs. MD</bold></th>
<th valign="top" align="center" colspan="2"><bold>CD vs. SB</bold></th>
</tr>
</thead>
<tbody>
<tr style="border-bottom: thin solid #000000;">
<td valign="top" align="left"><bold>No</bold></td>
<td valign="top" align="left"><bold>Short name</bold></td>
<td valign="top" align="left"><bold>KS stats</bold></td>
<td valign="top" align="left"><bold>KS</bold> <italic><bold>p</bold></italic><bold>-value</bold></td>
<td valign="top" align="left"><bold>KS stats</bold></td>
<td valign="top" align="left"><bold>KS</bold> <italic><bold>p</bold></italic><bold>-value</bold></td>
</tr> <tr>
<td valign="top" align="left">1</td>
<td valign="top" align="left">term_36m</td>
<td valign="top" align="left">0.0407</td>
<td valign="top" align="left">&#x0003C;0.24</td>
<td valign="top" align="left">0.0357</td>
<td valign="top" align="left">&#x0003C;0.20</td>
</tr>
<tr>
<td valign="top" align="left">2</td>
<td valign="top" align="left">term_60m</td>
<td valign="top" align="left">0.0407</td>
<td valign="top" align="left">&#x0003C;0.24</td>
<td valign="top" align="left">0.0357</td>
<td valign="top" align="left">&#x0003C;0.20</td>
</tr>
<tr>
<td valign="top" align="left">3</td>
<td valign="top" align="left">grade_n</td>
<td valign="top" align="left">0.0792</td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left">0.0984</td>
<td valign="top" align="left">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">4</td>
<td valign="top" align="left">sub_grade_n</td>
<td valign="top" align="left">0.0884</td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left">0.1069</td>
<td valign="top" align="left">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">5</td>
<td valign="top" align="left">int_rate_n</td>
<td valign="top" align="left">0.0941</td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left">0.1033</td>
<td valign="top" align="left">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">6</td>
<td valign="top" align="left">revol_util_n</td>
<td valign="top" align="left"><bold>0.2248</bold></td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left"><bold>0.2292</bold></td>
<td valign="top" align="left">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">7</td>
<td valign="top" align="left">emp_length_n</td>
<td valign="top" align="left">0.0242</td>
<td valign="top" align="left">&#x0003C;0.85</td>
<td valign="top" align="left">0.0749</td>
<td valign="top" align="left">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">8</td>
<td valign="top" align="left">dti_n</td>
<td valign="top" align="left"><bold>0.1502</bold></td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left"><bold>0.2295</bold></td>
<td valign="top" align="left">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">9</td>
<td valign="top" align="left">installment_n</td>
<td valign="top" align="left"><bold>0.3005</bold></td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left">0.0671</td>
<td valign="top" align="left">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">10</td>
<td valign="top" align="left">annual_inc_n</td>
<td valign="top" align="left">0.0670</td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left">0.0906</td>
<td valign="top" align="left">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">11</td>
<td valign="top" align="left">loan_amnt_n</td>
<td valign="top" align="left"><bold>0.2899</bold></td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left">0.0813</td>
<td valign="top" align="left">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">12</td>
<td valign="top" align="left">cover</td>
<td valign="top" align="left"><bold>0.2736</bold></td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left">0.0585</td>
<td valign="top" align="left">&#x0003C;0.01</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Values in bold represent large differences that are statistically significant</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>5.4.3. Domain Adaptation</title>
<p>To further understand the contribution of cover, we applied our proposed domain adaptation function. Before we trained the transfer model on the source data, we adapted the cover on the source data to make it similar to the target data, and then applied the transfer learning technique to produce an &#x0201C;adapted&#x0201D; and transferred model. The AUC tests for these adapted and transferred models are listed in <xref ref-type="table" rid="T5">Table 5</xref> where they are compared to the transferred model without adaptation. We ran paired <italic>t</italic>-tests to test the &#x0201C;improvements&#x0201D; (AUC increase or decrease) shown in <xref ref-type="table" rid="T5">Tables 5</xref>&#x02013;<bold>7</bold>; all improvements were statistically significant with <italic>p</italic> &#x0003C; 0.01. Since this data was normally distributed <italic>t</italic>-tests were appropriate. Adapting cover works for CD to MD transfer with an AUC 0.01 (1.6%) higher than the transfer-only model, but AUC decreases for a CD to SB transfer.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Adapted model vs. transferred model.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Domain</bold></th>
<th valign="top" align="center"><bold>Experiment</bold></th>
<th valign="top" align="center"><bold>AUC (s.d.)</bold></th>
<th valign="top" align="center"><bold>Improvement</bold></th>
<th valign="top" align="center"><bold><italic>p</italic>-value</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">CD to MD</td>
<td valign="top" align="left">Transfer only</td>
<td valign="top" align="center">0.6391(0.0856)</td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">CD to MD</td>
<td valign="top" align="left">Transfer with cover adapted</td>
<td valign="top" align="center">0.6491(0.0824)</td>
<td valign="top" align="center">0.0100 (1.6%)</td>
<td valign="top" align="center">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">CD to SB</td>
<td valign="top" align="left">Transfer only</td>
<td valign="top" align="center">0.6419(0.0509)</td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">CD to SB</td>
<td valign="top" align="left">Transfer with cover adapted</td>
<td valign="top" align="center">0.6361(0.0502)</td>
<td valign="top" align="center">&#x02013;0.0058 (&#x02013;0.9%)</td>
<td valign="top" align="center">&#x0003C;0.01</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>We also examined SHAP changes with the improved model. After adapting cover, it became the fourth highest contributing feature, and the top three contributing features (int_rate_n, term_36m, and annual_inc_n) were the top contributing features of the target and source models.</p>
<p>We tested various permutations of features to adapt to find the most accurate model for the CD to MD transfer and to establish an optimal strategy for seeking the most accurate adapted model. The experiments on the CD to MD transfer are listed in <xref ref-type="table" rid="T6">Table 6</xref>. When adapting all features, or adapting credit grade and related features, model accuracy was significantly reduced, with AUC 0.1771 (27.7%) or 0.1756 (27.5%) lower than the transfer-only model. Adapting only features with a high KS-statistic (over 0.15), i.e., revolving utility, debt to income ratio, installment, loan amount, and cover, improved accuracy with AUC 0.0172 (2.7%) higher than the transfer-only model. Adding related features, i.e., annual income (annual_inc_n)&#x02014;which is used to derive cover (a high KS feature), further improved accuracy, with AUC 0.0209 (3.3%) higher than the transfer-only model. Removing credit history features that are intrinsic to the borrower, i.e., revolving utility and debt to income ratio, produced an even more accurate model, with AUC 0.0257 (4.0%) higher than the transfer-only model.</p>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>Adapted model vs. transferred model in CD to MD transfer.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="center"><bold>Experiment</bold></th>
<th valign="top" align="center"><bold>AUC (s.d.)</bold></th>
<th valign="top" align="center"><bold>Improvement</bold></th>
<th valign="top" align="center"><bold><italic>p</italic>-value</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Transfer only</td>
<td valign="top" align="center">0.6391(0.0856)</td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">Adapt all features</td>
<td valign="top" align="center">0.4620(0.3048)</td>
<td valign="top" align="center">&#x02013;0.1771 (&#x02013;27.7%)</td>
<td valign="top" align="center">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">Adapt credit grade and related features, i.e., grade, sub-grade, interest rate</td>
<td valign="top" align="center">0.4635(0.3052)</td>
<td valign="top" align="center">&#x02013;0.1756 (&#x02013;27.5%)</td>
<td valign="top" align="center">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">Adapt features with high KS, i.e., revolving utility, debt to income ratio, installment, loan amount and cover</td>
<td valign="top" align="center">0.6563(0.0806)</td>
<td valign="top" align="center">0.0172 (2.7%)</td>
<td valign="top" align="center">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">Adapt features with high KS and related features, i.e., revolving utility, debt to income ratio, installment, loan amount, cover and annual income</td>
<td valign="top" align="center">0.6600(0.07417)</td>
<td valign="top" align="center">0.0209 (3.3%)</td>
<td valign="top" align="center">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">Adapt features with high KS and related features less credit history features, i.e., installment, loan amount, cover and annual income</td>
<td valign="top" align="center">0.6649(0.0731)</td>
<td valign="top" align="center">0.0257 (4.0%)</td>
<td valign="top" align="center">&#x0003C;0.01</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Grade, sub-grade, revolving utility (revol_util_n), and debt to income ratio (dti_n) are features derived from credit history, which are intrinsic to the borrower and are usually highly correlated with the lending outcome, i.e., default or not. The interest rate in the lending club data is derived directly from grade and sub-grade, so we consider it as a credit history feature in our experiment.</p>
<p>We use a similar strategy for the CD to SB transfer. The AUC comparison with the transfer-only model is shown in <xref ref-type="table" rid="T7">Table 7</xref>. Adapting all features, or credit grade related features, significantly reduced model accuracy, with AUC 0.123 (19.3%) or 0.1106 (17.2%) lower than the transfer-only model, respectively. We tested adaptation of the features that we adapted for the most accurate model for the CD to MD transfer, which has a low KS-statistic from CD and SB comparisons. This adapted model was slightly less accurate, with an AUC 0.0015 (0.2%) lower than the transfer-only model. Adapting features with a high KS-statistic, i.e., revolving utility and debt to income ratio, improved model accuracy slightly, with AUC 0.0018 (0.3%) higher than the transfer-only model. These two features do not have related features, and both were credit history features, so we could not improve accuracy further as we did with the CD to MD transfer.</p>
<table-wrap position="float" id="T7">
<label>Table 7</label>
<caption><p>Adapted model vs. transferred model in CD to SB transfer.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="center"><bold>Experiment</bold></th>
<th valign="top" align="center"><bold>AUC (s.d.)</bold></th>
<th valign="top" align="center"><bold>Improvement</bold></th>
<th valign="top" align="center"><bold><italic>p</italic>-value</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Transfer only</td>
<td valign="top" align="center">0.6419(0.0509)</td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">Adapt all features</td>
<td valign="top" align="center">0.5189(0.1666)</td>
<td valign="top" align="center">&#x02013;0.123 (&#x02013;19.2%)</td>
<td valign="top" align="center">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">Adapt credit grade and related features, i.e., grade, sub-grade, interest rate</td>
<td valign="top" align="center">0.5313(0.1624)</td>
<td valign="top" align="center">&#x02013;0.1106 (&#x02013;17.2%)</td>
<td valign="top" align="center">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">Adapt features used in CD to MD transfer, i.e., installment, loan amount, cover, annual income</td>
<td valign="top" align="center">0.6404(0.0475)</td>
<td valign="top" align="center">&#x02013;0.0015 (&#x02013;0.2%)</td>
<td valign="top" align="center">&#x0003C;0.01</td>
</tr>
<tr>
<td valign="top" align="left">Adapt features with high KS, i.e., revolving utility and debt to income ratio</td>
<td valign="top" align="center">0.6437(0.0495)</td>
<td valign="top" align="center">0.0018 (0.3%)</td>
<td valign="top" align="center">&#x0003C;0.01</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Additionally, we investigated the explainability of the most accurate models using SHAP to assess the feature contributions of the most accurate adapted models compared to the source and target models. Through domain adaptation, the contribution of weak features increased in the most accurate adapted models. For the CD to MD transfer, the contribution of annual income, cover, installment, and loan amount increased. For the CD to SB transfer, the contribution of annual income, term 36 months or 60 months, cover, employment length, installment, and loan amount increased.</p>
<p>To evaluate the effectiveness of our adaptation approach, we compared KS values before and after adaptation for the most accurate models, as shown in <xref ref-type="table" rid="T8">Table 8</xref>. The reduction in KS-statistics was between 44.8 and 90.3%, and for features, with high KS-statistics (over 0.15) the reductions were all above 67.4%. Our adaptation approach successfully reduced the differences between the distribution of the source data and the target data.</p>
<table-wrap position="float" id="T8">
<label>Table 8</label>
<caption><p>Kolmogorov-Smirnov test to compare source data and target data, before and after the source data is adapted, ACD is the abbreviation for Adapted Credit card and Debt consolidation data.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="middle" align="left" rowspan="2"><bold>Feature</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>Without adaptation</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>With adaptation</bold></th>
<th valign="middle" align="left" rowspan="2"><bold>Reduction</bold></th>
</tr>
<tr>
<th valign="top" align="center"><bold>Domain</bold></th>
<th valign="top" align="center"><bold>KS-stats</bold></th>
<th valign="top" align="center"><bold><italic>p</italic>-value</bold></th>
<th valign="top" align="center"><bold>Domain</bold></th>
<th valign="top" align="center"><bold>KS-stats</bold></th>
<th valign="top" align="center"><bold><italic>p</italic>-value</bold></th>
</tr>
</thead>
<tbody> <tr>
<td valign="top" align="left">installment</td>
<td valign="top" align="left">CD to MD</td>
<td valign="top" align="left">0.3005</td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left">ACD to MD</td>
<td valign="top" align="left">0.0293</td>
<td valign="top" align="left">&#x0003C;0.64</td>
<td valign="top" align="left">90.3%</td>
</tr>
<tr>
<td valign="top" align="left">annual_inc</td>
<td valign="top" align="left">CD to MD</td>
<td valign="top" align="left">0.0670</td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left">ACD to MD</td>
<td valign="top" align="left">0.0369</td>
<td valign="top" align="left">&#x0003C;0.34</td>
<td valign="top" align="left">44.8%</td>
</tr>
<tr>
<td valign="top" align="left">loan_amnt</td>
<td valign="top" align="left">CD to MD</td>
<td valign="top" align="left">0.2899</td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left">ACD to MD</td>
<td valign="top" align="left">0.0681</td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left">76.5%</td>
</tr>
<tr>
<td valign="top" align="left">cover</td>
<td valign="top" align="left">CD to MD</td>
<td valign="top" align="left">0.2736</td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left">ACD to MD</td>
<td valign="top" align="left">0.0892</td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left">67.4%</td>
</tr>
<tr>
<td valign="top" align="left">revol_util</td>
<td valign="top" align="left">CD to SB</td>
<td valign="top" align="left">0.2292</td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left">ACD to SB</td>
<td valign="top" align="left">0.0536</td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left">76.6%</td>
</tr>
<tr>
<td valign="top" align="left">dti</td>
<td valign="top" align="left">CD to SB</td>
<td valign="top" align="left">0.2295</td>
<td valign="top" align="left">&#x0003C;0.01</td>
<td valign="top" align="left">ACD to SB</td>
<td valign="top" align="left">0.0301</td>
<td valign="top" align="left">&#x0003C;0.39</td>
<td valign="top" align="left">86.9%</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
</sec>
<sec sec-type="discussion" id="s6">
<title>6. Discussion</title>
<p>Transfer learning improves model accuracy by generating intermediate features from the source domain to be selected for retraining on the target domain. This intermediate-feature-generation concept is similar to &#x0201C;self taught learning&#x0201D; proposed by Raina et al. (<xref ref-type="bibr" rid="B15">2007</xref>), which constructed higher-level features using unlabeled data, except that in this article, we used labeled data from the source domain.</p>
<p>The contribution of a weak feature from the target domain increased because it was complemented by new intermediate features from the source domain. We tested an adaptation approach taking the outcome label into consideration. But this did not improve model accuracy. The reason was that the population of positive (outcome=1) cases was too small in the already small target dataset.</p>
<p>Adapting strong credit history features, such as grade and sub-grade, significantly reduced model accuracy, while removing features related to credit history from the adaptation list improved model accuracy. Adapting credit history related features <italic>without consideration of the outcome label</italic> generates unrealistic instances, e.g., changing a borrower&#x00027;s credit grade from high to low without adjusting the outcome from not default to default. These unrealistic instances can negatively impact model accuracy.</p>
</sec>
<sec id="s7">
<title>7. Conclusion and Future Study</title>
<p>In this article, we have proposed and evaluated a method of gradually shifting the contribution from the source to the target domain during transfer learning. The <italic>PSC</italic> method in a structured way varies network architecture and hyperparameters to find the best balance to learn from source and target domains. While in this article, certain design choices were made manually, the overarching aim was to create a framework within which such changes could be optimized automatically, in our setting to find the best Gini score.</p>
<p>In terms of interpretability of models for credit assessment, we found that SHAP is an effective tool in explaining why transfer learning improves the accuracy of credit scoring models. In our experiments transfer learning lifted the contribution of weak features, thereby improving overall prediction accuracy.</p>
<p>Domain adaptation with the right set of features further improved the accuracy of transfer learning models. However, adapting all features normally reduces model accuracy significantly. Reasons to select features to adapt include: differences in feature distribution between source and target domain, quantified by KS-statistics; relationships to already selected features; and domain specific knowledge, e.g., the credit history features intrinsic to the borrowers.</p>
<p>Through domain adaptation, the contribution of weaker features increased in the most accurate adapted models. An adaptation approach that significantly reduces KS-statistics has been critical in producing a successful domain adaptation algorithm.</p>
<p>If a model constructed by a machine learning approach may affect capital requirements, then such models need to be reviewed by a regulatory institution. However, business cases where our approach is applicable are in many domains that do not fall in these categories.</p>
<p>For future study, the proposed strategy to select features for domain adaptation produces more accurate credit scoring models, but the execution of the strategy requires human intervention in observing and applying domain knowledge. We will further explore methods to automate this selection strategy, so it can be a pre-processing step for fully automated transfer learning. Although in this article we have focused on transfer learning based on deep networks, other model architectures could be used. For example, Goussies et al. (<xref ref-type="bibr" rid="B7">2014</xref>) and Segev et al. (<xref ref-type="bibr" rid="B18">2016</xref>) have used ensembles of tree learners as the basis for transfer learning. Comparison of deep networks with such approaches could be investigated as part of further study.</p>
<p>The use of alternatives to KS statistics to estimate the distance between distributions, such as KL-divergence, should be investigated. SHAP is an indirect method to understand the impact of latent intermediate features. A further study exploring and explaining latent intermediate features could improve our understanding of transfer learning and domain adaptation, and better meet transparency and compliance requirements.</p>
<p>Finally, we note that although the significant improvements in accuracy demonstrated are small in terms of percentage improvements, such improvements in real world lending could be of substantial economic importance in reducing lenders&#x00027; losses due to loan defaults.</p>
</sec>
<sec sec-type="data-availability" id="s8">
<title>Data Availability Statement</title>
<p>The software and instructions for downloading and pre-processing the data are provided at the following link to help Reproducible Research URL: <ext-link ext-link-type="uri" xlink:href="https://gitlab.com/research-study/ecmlpkdd2020">https://gitlab.com/research-study/ecmlpkdd2020</ext-link>.</p>
</sec>
<sec id="s9">
<title>Author Contributions</title>
<p>HS, CG, and AG contributed to conception and design of the study, wrote the first draft of the manuscript. HS contributed to implementations and statistical analysis. HS, AM, and MB contributed to research and analysis and wrote sections of the manuscript. All authors contributed to manuscript revision, read, and approved the submitted version.</p>
</sec>
<sec sec-type="funding-information" id="s10">
<title>Funding</title>
<p>This study received funding from Rich Data Corporation, Sydney, Australia. The funder was not involved in the study design, collection, analysis, interpretation of data, the writing of this article or the decision to submit it for publication.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>HS, AM, CG, and AG are employed by Rich Data Corporation, Sydney, Australia. The remaining author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s11">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ack><p>The content of this manuscript has been presented in part at The European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (2019), Suryanto et al. (<xref ref-type="bibr" rid="B21">2019</xref>).</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="web"><person-group person-group-type="author"><collab>BankRate</collab></person-group> (<year>2019</year>). <article-title>58% of millennials have been denied at least one financial product because of their credit score</article-title>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.bankrate.com/credit-cards/credit-denial-survey/">https://www.bankrate.com/credit-cards/credit-denial-survey/</ext-link></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Breiman</surname> <given-names>L.</given-names></name></person-group> (<year>1996</year>). <article-title>Some properties of splitting criteria</article-title>. <source>Mach. Learn</source>. <volume>24</volume>, <fpage>41</fpage>&#x02013;<lpage>47</lpage>. <pub-id pub-id-type="doi">10.1007/BF00117831</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chattopadhyay</surname> <given-names>R.</given-names></name> <name><surname>Ye</surname> <given-names>J.</given-names></name> <name><surname>Panchanathan</surname> <given-names>S.</given-names></name> <name><surname>Fan</surname> <given-names>W.</given-names></name> <name><surname>Davidson</surname> <given-names>I.</given-names></name></person-group> (<year>2011</year>). <article-title>Multi-source domain adaptation and its application to early detection of fatigue</article-title>, in <source>KDD</source> (<publisher-loc>San Diego, CA</publisher-loc>).</citation>
</ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Flach</surname> <given-names>P. A.</given-names></name></person-group> (<year>2003</year>). <article-title>The geometry of roc space: understanding machine learning metrics through roc isometrics</article-title>, in <source>Proceedings of the 20th International Conference on Machine Learning (ICML-03)</source> (<publisher-loc>Washington, DC</publisher-loc>), <fpage>194</fpage>&#x02013;<lpage>201</lpage>.</citation>
</ref>
<ref id="B5">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Forum</surname> <given-names>S. F.</given-names></name></person-group> (<year>2019</year>). <source>Msme finance gap</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.smefinanceforum.org/data-sites/msme-finance-gap/">https://www.smefinanceforum.org/data-sites/msme-finance-gap/</ext-link></citation>
</ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>J.</given-names></name> <name><surname>Fan</surname> <given-names>W.</given-names></name> <name><surname>Jiang</surname> <given-names>J.</given-names></name> <name><surname>Han</surname> <given-names>J.</given-names></name></person-group> (<year>2008</year>). <article-title>Knowledge transfer via multiple model local structure mapping</article-title>, in <source>KDD</source>.</citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Goussies</surname> <given-names>N. A.</given-names></name> <name><surname>Ubalde</surname> <given-names>S.</given-names></name> <name><surname>Mejail</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>Transfer learning decision forests for gesture recognition</article-title>. <source>J. Mach. Learn. Res</source>. <volume>15</volume>, <fpage>3667</fpage>&#x02013;<lpage>3690</lpage>. <pub-id pub-id-type="doi">10.5555/2627435.2750362</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Han</surname> <given-names>X.</given-names></name> <name><surname>Huang</surname> <given-names>Z.</given-names></name> <name><surname>An</surname> <given-names>B.</given-names></name> <name><surname>Bai</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>Adaptive transfer learning on graph neural networks</article-title>, in <source>Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD &#x00027;21</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>565</fpage>&#x02013;<lpage>574</lpage>.</citation>
</ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>J.</given-names></name> <name><surname>Chen</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>Domain adaptation approach for credit risk analysis</article-title>, in <source>Proceedings of the 2018 International Conference on Software Engineering and Information Management, ICSIM2018</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>104</fpage>&#x02013;<lpage>107</lpage>.</citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kouw</surname> <given-names>W. M.</given-names></name> <name><surname>Loog</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <article-title>A review of domain adaptation without target labels</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. <volume>43</volume>, <fpage>766</fpage>&#x02013;<lpage>785</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2019.2945942</pub-id><pub-id pub-id-type="pmid">31603771</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lundberg</surname> <given-names>S.</given-names></name> <name><surname>Lee</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>A unified approach to interpreting model predictions</article-title>, in <source>Advances in Neural Information Processing Systems</source> (<publisher-loc>Long Beach, CA</publisher-loc>), <fpage>4765</fpage>&#x02013;<lpage>4774</lpage>.</citation>
</ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Neyshabur</surname> <given-names>B.</given-names></name> <name><surname>Sedghi</surname> <given-names>H.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name></person-group> (<year>2020</year>). <article-title>What is being transferred in transfer learning?</article-title>, in <source>NeurIPS</source> (<publisher-loc>Vancouver, BC</publisher-loc>).</citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pan</surname> <given-names>S. J.</given-names></name> <name><surname>Tsang</surname> <given-names>I. W.</given-names></name> <name><surname>Kwok</surname> <given-names>J. T.</given-names></name> <name><surname>Yang</surname> <given-names>Q.</given-names></name></person-group> (<year>2011</year>). <article-title>Domain adaptation via transfer component analysis</article-title>. <source>IEEE Trans. Neural Netw</source>. <volume>22</volume>, <fpage>199</fpage>&#x02013;<lpage>210</lpage>. <pub-id pub-id-type="doi">10.1109/TNN.2010.2091281</pub-id><pub-id pub-id-type="pmid">21095864</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pan</surname> <given-names>S. J.</given-names></name> <name><surname>Yang</surname> <given-names>Q.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>A survey on transfer learning</article-title>. <source>IEEE Trans, Knowl, Data Eng</source>. <volume>22</volume>, <fpage>1345</fpage>&#x02013;<lpage>1359</lpage>. <pub-id pub-id-type="doi">10.1109/TKDE.2009.191</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Raina</surname> <given-names>R.</given-names></name> <name><surname>Battle</surname> <given-names>A.</given-names></name> <name><surname>Lee</surname> <given-names>H.</given-names></name> <name><surname>Packer</surname> <given-names>B.</given-names></name> <name><surname>Ng</surname> <given-names>A. Y.</given-names></name></person-group> (<year>2007</year>). <article-title>Self-taught learning: transfer learning from unlabeled data</article-title>, in <source>Proceedings of the 24th International Conference on Machine Learning, ICML 9207</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>759</fpage>&#x02013;<lpage>766</lpage>.</citation>
</ref>
<ref id="B16">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>&#x00158;ez&#x000E1;&#x0010D;</surname> <given-names>M.</given-names></name> <name><surname>&#x00158;ez&#x000E1;&#x0010D;</surname> <given-names>F.</given-names></name></person-group> (<year>2011</year>). <article-title>How to measure the quality of credit scoring models</article-title>. <source>Finance a &#x000FA;v&#x0011B;r: Czech J. Econ. Fin.</source> <volume>61</volume>, <fpage>486</fpage>&#x02013;<lpage>507</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://ideas.repec.org/a/fau/fauart/v61y2011i5p486-507.html">https://ideas.repec.org/a/fau/fauart/v61y2011i5p486-507.html</ext-link></citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rusu</surname> <given-names>A. A.</given-names></name> <name><surname>Rabinowitz</surname> <given-names>N. C.</given-names></name> <name><surname>Desjardins</surname> <given-names>G.</given-names></name> <name><surname>Soyer</surname> <given-names>H.</given-names></name> <name><surname>Kirkpatrick</surname> <given-names>J.</given-names></name> <name><surname>Kavukcuoglu</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Progressive neural networks</article-title>. <source>ArXiv, abs/1606.04671</source>.</citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Segev</surname> <given-names>N.</given-names></name> <name><surname>Harel</surname> <given-names>M.</given-names></name> <name><surname>Mannor</surname> <given-names>S.</given-names></name> <name><surname>Crammer</surname> <given-names>K.</given-names></name> <name><surname>El-Yaniv</surname> <given-names>R.</given-names></name></person-group> (<year>2016</year>). <article-title>Learn on source, refine on target: a model transfer learning framework with random forests</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. <volume>39</volume>, <fpage>1811</fpage>&#x02013;<lpage>1824</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2016.2618118</pub-id><pub-id pub-id-type="pmid">28113392</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Selbst</surname> <given-names>A.</given-names></name> <name><surname>Powles</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>Meaningful information and the right to explanation</article-title>. <source>Int. Data Privacy Law</source> <volume>7</volume>, <fpage>233</fpage>&#x02013;<lpage>242</lpage>. <pub-id pub-id-type="doi">10.1093/idpl/ipx022</pub-id><pub-id pub-id-type="pmid">32106285</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>S.</given-names></name> <name><surname>Xu</surname> <given-names>Z.</given-names></name> <name><surname>Yang</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>Transfer learning with part-based ensembles</article-title>, in <source>MCS</source> (<publisher-loc>Berlin</publisher-loc>).</citation>
</ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Suryanto</surname> <given-names>H.</given-names></name> <name><surname>Guan</surname> <given-names>C.</given-names></name> <name><surname>Voumard</surname> <given-names>A.</given-names></name> <name><surname>Beydoun</surname> <given-names>G.</given-names></name></person-group> (<year>2019</year>). <article-title>Transfer learning in credit risk</article-title>, in <source>The European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases</source> (<publisher-loc>W&#x000FC;rzburg</publisher-loc>: <publisher-name>Springer</publisher-name>).</citation>
</ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Hao</surname> <given-names>S.</given-names></name> <name><surname>Feng</surname> <given-names>W.</given-names></name> <name><surname>Shen</surname> <given-names>Z.</given-names></name></person-group> (<year>2017</year>). <article-title>Balanced distribution adaptation for transfer learning</article-title>, in <source>2017 IEEE International Conference on Data Mining (ICDM)</source> (<publisher-loc>New Orleans, LA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1129</fpage>&#x02013;<lpage>1134</lpage>.</citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Weiss</surname> <given-names>K. R.</given-names></name> <name><surname>Khoshgoftaar</surname> <given-names>T. M.</given-names></name> <name><surname>Wang</surname> <given-names>D.</given-names></name></person-group> (<year>2016</year>). <article-title>A survey of transfer learning</article-title>. <source>J. Big Data</source> <volume>3</volume>, <fpage>1</fpage>&#x02013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1186/s40537-016-0043-6</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="web"><person-group person-group-type="author"><collab>World Bank</collab></person-group> (<year>2017</year>). <source>Financial inclusion and inclusive growth</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://documents.worldbank.org/curated/en/403611493134249446/pdf/WPS8040.pdf">http://documents.worldbank.org/curated/en/403611493134249446/pdf/WPS8040.pdf</ext-link></citation>
</ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xiao</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>R.</given-names></name> <name><surname>Teng</surname> <given-names>G.-E.</given-names></name> <name><surname>Hu</surname> <given-names>Y.</given-names></name></person-group> (<year>2014</year>). <article-title>A transfer learning based classifier ensemble model for customer credit scoring</article-title>, in <source>2014 Seventh International Joint Conference on Computational Sciences and Optimization</source> (<publisher-loc>Beijing</publisher-loc>), <fpage>64</fpage>&#x02013;<lpage>68</lpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Yan</surname> <given-names>R.</given-names></name> <name><surname>Hauptmann</surname> <given-names>A. G.</given-names></name></person-group> (<year>2007</year>). <article-title>Cross-domain video concept detection using adaptive svms</article-title>, in <source>Proceedings of the 15th ACM International Conference on Multimedia</source> (<publisher-loc>Augsburg</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>188</fpage>&#x02013;<lpage>197</lpage>.</citation>
</ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Ogunbona</surname> <given-names>P.</given-names></name></person-group> (<year>2018</year>). <article-title>Unsupervised domain adaptation: a multi-task learning-based method</article-title>. <source>Knowl. Based Syst</source>. <volume>186</volume>, <fpage>104975</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2019.104975</pub-id></citation>
</ref>
</ref-list>
<fn-group>
<fn id="fn0001"><p><sup>1</sup>Available from <ext-link ext-link-type="uri" xlink:href="https://www.lendingclub.com/info/download-data.action">https://www.lendingclub.com/info/download-data.action</ext-link></p></fn>
<fn id="fn0002"><p><sup>2</sup>The hyper parameters optimization has been done before this step.</p></fn>
<fn id="fn0003"><p><sup>3</sup>See <ext-link ext-link-type="uri" xlink:href="https://github.com/slundberg/shap">https://github.com/slundberg/shap</ext-link></p></fn>
</fn-group>
</back>
</article>