<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">761925</article-id>
<article-id pub-id-type="doi">10.3389/frai.2021.761925</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Self-Organising Map Based Framework for Investigating Accounts Suspected of Money Laundering</article-title>
<alt-title alt-title-type="left-running-head">Alshantti and Rasheed</alt-title>
<alt-title alt-title-type="right-running-head">SOM for Money Laundering Investigation</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Alshantti</surname>
<given-names>Abdallah</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1435934/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Rasheed</surname>
<given-names>Adil</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/985039/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Department of Engineering Cybernetics, Norwegian University of Science and Technology</institution>, <addr-line>Trondheim</addr-line>, <country>Norway</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Department of Mathematics and Cybernetics, SINTEF Digital</institution>, <addr-line>Trondheim</addr-line>, <country>Norway</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/598769/overview">Aparna Gupta</ext-link>, Rensselaer Polytechnic Institute, United&#x20;States</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/596324/overview">Simone Righi</ext-link>, University College London, United&#x20;Kingdom</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/596489/overview">Bertrand Kian Hassani</ext-link>, University College London, United&#x20;Kingdom</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Abdallah Alshantti, <email>abdallah.a.s.alshantti@ntnu.no</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Artificial Intelligence in Finance, a section of the journal Frontiers in Artificial Intelligence</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>14</day>
<month>12</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>4</volume>
<elocation-id>761925</elocation-id>
<history>
<date date-type="received">
<day>20</day>
<month>08</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>04</day>
<month>11</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Alshantti and Rasheed.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Alshantti and Rasheed</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>There has been an emerging interest by financial institutions to develop advanced systems that can help enhance their anti-money laundering (AML) programmes. In this study, we present a self-organising map (SOM) based approach to predict which bank accounts are possibly involved in money laundering cases, given their financial transaction histories. Our method takes advantage of the competitive and adaptive properties of SOM to represent the accounts in a lower-dimensional space. Subsequently, categorising the SOM and the accounts into money laundering risk levels and proposing investigative strategies enables us to measure the classification performance. Our results indicate that our framework is well capable of identifying suspicious accounts already investigated by our partner bank, using both proposed investigation strategies. We further validate our model by analysing the performance when modifying different parameters in our dataset.</p>
</abstract>
<kwd-group>
<kwd>self-organising map (SOM)</kwd>
<kwd>money laundering</kwd>
<kwd>suspicious accounts</kwd>
<kwd>risk levels</kwd>
<kwd>investigation strategies</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Money laundering represents a major challenge for governments and financial institutions alike, as the flow of dirty money can hinder a state&#x2019;s development, damage the reputation of the financial system and motivate the generation of further crime (<xref ref-type="bibr" rid="B26">Kumar, 2012</xref>). Since money laundering is an underground market, its true magnitude can never be precisely determined. However, the most recent estimate of the laundered money worldwide considers the sum of illicit transactions worldwide to account for 2&#x2013;5% of the global GDP (<xref ref-type="bibr" rid="B29">Levi and Reuter, 2006</xref>). In addition, financial institutions that fail to deploy satisfactory measures for investigating and reporting suspicious cases, thus not sufficiently filing suspicious activity reports, are subject to heavy fines and penalties by the authorities. Therefore, there has been an increasing interest over the past decades to develop tools that aid in the combat against money laundering.</p>
<p>The most common methods used by financial institutions for detecting and reporting suspicious cases are rule-based systems. A rule-based system is a set of predefined rules gathered from a knowledge base maintained by domain experts. Transactions that match the predefined conditions trigger alerts, which prompt further investigations by bank compliance teams. A major drawback of the rule-based systems is the generation of a significant volume of false positive alerts that are costly in terms of time and resources needed to track down flagged cases (<xref ref-type="bibr" rid="B12">Gao, 2009</xref>). Those false positive alarms are estimated to constitute more than 90% of the total alerts generated by the traditional rule-based systems commonly adopted by banks (<xref ref-type="bibr" rid="B4">Breslow et&#x20;al., 2017</xref>). Consequently, there has been an increasing desire to develop more advanced tools for more precise detection of money laundering transactions.</p>
<p>In several studies, new frameworks and systems were developed for identifying money laundering transactions or accounts (<xref ref-type="bibr" rid="B30">Liu et&#x20;al., 2008</xref>; <xref ref-type="bibr" rid="B28">Le-Khac et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B31">Lopez-Rojas and Axelsson, 2012</xref>; <xref ref-type="bibr" rid="B19">Jun, 2006</xref>; <xref ref-type="bibr" rid="B41">Soltani et&#x20;al., 2016</xref>). While many of these works claim a good detection performance, they often use synthetic data due to the limited availability of financial transactional data. Therefore, the reliability of such methods was not tested on real financial data, which might indicate that there might be no practical use for such frameworks. In addition, money laundering datasets are heavily characterised by a significant class imbalance, where the number of normal transactions tremendously exceeds the number of illicit transactions. The class imbalance represents a major challenge in the detection of suspicious transactions, and as a result, the data is often either undersampled or oversampled to improve the classification performance (<xref ref-type="bibr" rid="B43">Sudjianto et&#x20;al., 2010</xref>). Another significant challenge is the quality of the data and the labels used. It is estimated that sum of money from the money laundering cases investigated annually make up no more than 10% of the total amount laundered (<xref ref-type="bibr" rid="B1">Argentiero et&#x20;al., 2008</xref>). Therefore, while financial institutions may have labelled datasets that indicate which transactions were investigated as suspicious, it is highly likely that these labels represent only a fraction of the total number of money laundering transactions that went undetected.</p>
<p>Overall, the most prevalent limitation in AML literature is that many of the proposed models are developed with the aim of replacing human intervention (<xref ref-type="bibr" rid="B6">Chen et&#x20;al., 2014</xref>). In essence, these studies aim to detect money laundering transactions in a binary classification setting. While machine learning based automation is increasingly deemed as a valuable mean in AML for tasks such as transaction monitoring (<xref ref-type="bibr" rid="B47">Turki et&#x20;al., 2020</xref>) and processing client information (<xref ref-type="bibr" rid="B46">Tiwari et&#x20;al., 2020</xref>), a sensitive field such as AML cannot solely rely on machine learning algorithms to replace human expertise for reporting illegal transactions. Instead, compliance teams have an obligation of providing their reasoning when filing suspicious activity reports (<xref ref-type="bibr" rid="B39">Singh and Lin, 2020</xref>). The requirement to include the explanation when submitting suspected money laundering cases to the authorities entails that the adoption of technologies in the finance compliance sector ought to aid the decision making of case investigators, rather than fully automating the money laundering detection process.</p>
<p>In this work we aim to address these underlying gaps in literature by proposing a SOM-based approach to assist in the decision making process of compliance teams when investigating bank accounts. A SOM is an unsupervised neural network that maps a highly dimensional dataset into a lower-dimensional representation (<xref ref-type="bibr" rid="B24">Kohonen, 2012</xref>). SOMs are commonly used in clustering (<xref ref-type="bibr" rid="B16">Isa et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B14">Haga et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B32">Nilashi et&#x20;al., 2020</xref>), classification (<xref ref-type="bibr" rid="B53">Yorek et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B17">Jain et&#x20;al., 2018</xref>) and oversampling (<xref ref-type="bibr" rid="B7">Douzas and Bacao, 2017</xref>) applications. For its notable success in these applications, we propose a SOM based method for clustering and classifying bank accounts into risk levels that can be further investigated by the bank. Consequently, our implementation of SOM is a hybrid approach that combines the strengths of unsupervised and supervised learning. In contrast to traditional clustering techniques where the parameter selection significantly impacts the grouping of observations, the number of clusters in SOMs is dependent on the number of training observations. The main parameters we introduce contribute to the categorisation of the formed clusters, such that these parameters can be tuned according to how strict or lenient the decision makers preferences are. Moreover, unlike K-means clustering, the dimension reduction in SOM helps in visualising how the bank accounts are distributed across the low-dimensional space, which can be used for further inferring how the accounts are grouped together.</p>
<p>To this end, we develop a novel framework based on SOMs to categorise the bank accounts in our dataset into different risk level groups. First, the data is pre-processed and aggregated to group the transactions by account. This is an important step since grouping the data by accounts enables the model to capture the historic financial patterns exhibited by the accounts. The most important features in the dataset are selected and the dataset is split into a training and test sets in a k-fold cross-validation manner. We then select the SOM and its hyperparameters and develop a categorisation technique for dividing the SOM neurons and the accounts assigned to them into three risk level groups based on the neurons&#x2019; properties. The model&#x2019;s performance is evaluated by proposing two investigation strategies using different utilisations of the risk levels. Finally, we further evaluate the classification performance of our model by varying the number of selected features and the ratio of suspicious accounts in the&#x20;data.</p>
<p>Taking the limitations of existing AML research and our motivation for this work into consideration, we highlight our contributions as follows:<list list-type="simple">
<list-item>
<p>1) We develop a framework that facilitates the investigation of potential money laundering accounts despite the major class imbalance between the normal and suspicious accounts in the dataset. The SOM model maps the observations into a 2-dimensional grid space, and in doing so, the technique attempts to group the unlabelled observations with similar transaction patterns together.</p>
</list-item>
<list-item>
<p>2) We contribute to the state of the art by introducing a system that can well handle poorly labelled financial transaction data. Since the training of our SOM grid does not rely on the whether a given sample is suspicious or not, our method takes into account the challenges of assigning the transactions with their true labels by financial compliance teams. This is due to the fact that money laundering transactions are often carried out in a way that makes them difficult to distinguish from normal transactions (<xref ref-type="bibr" rid="B45">Teichmann, 2020</xref>).</p>
</list-item>
</list>
</p>
<p>The remainder of this paper is divided as follows. In <xref ref-type="sec" rid="s3">Section 3</xref>, we review the existing literature on anti-money laundering and self-organising maps. We describe our dataset and the preprocessing method in <xref ref-type="sec" rid="s4">Section 4</xref>. <xref ref-type="sec" rid="s5">Section 5</xref> discusses the methodology which consists of feature selection, SOM architecture and risk level categorisation. In <xref ref-type="sec" rid="s6">Section 6</xref>, we evaluate the performance of investigation strategies and present the results of the training and test sets. The paper is concluded and motivations for future work are provided in <xref ref-type="sec" rid="s7">Section&#x20;7</xref>.</p>
</sec>
<sec id="s2">
<title>2 Literature Review</title>
<sec id="s2-1">
<title>2.1&#x20;Anti-Money Laundering</title>
<p>Several existing studies on money laundering detection have been focusing on data mining and supervised machine learning techniques when the data provided contains the labels of the previously identified transactions. (<xref ref-type="bibr" rid="B30">Liu et&#x20;al., 2008</xref>) developed a sequence matching algorithm to identify suspicious sequences in an account&#x2019;s history by comparing the similarities and differences with other sequences in the data. Similarly, (<xref ref-type="bibr" rid="B44">Tang and Yin, 2005</xref>) proposed a one-class SVM algorithm to identify suspicious and unusual patterns while highlighting the speed and efficiency of their method. (<xref ref-type="bibr" rid="B18">Jullum et&#x20;al., 2020</xref>) trained an XGBoost supervised model using financial transactional data for predicting whether certain financial transactions should be reported, and demonstrated that their method outperforms the existing approach used by financial institutions. The detection of suspicious transactions linked to terrorism financing was implemented by (<xref ref-type="bibr" rid="B37">Rocha-Salazar et&#x20;al., 2021</xref>) using real datasets provided by a financial institution in Mexico, which was found to reduce the number of false positives in comparison to rule-based system flagging.</p>
<p>Meanwhile, unsupervised methods are employed when the labels are unavailable or when using synthetic datasets. (<xref ref-type="bibr" rid="B51">Wang and Dong, 2009</xref>) proposed a minimum spanning tree clustering method to detect money laundering transactions using different tree levels. In a study by (<xref ref-type="bibr" rid="B6">Chen et&#x20;al., 2014</xref>), clustering of transactions to identify the suspicious transactions grouped together was carried out using an expectation maximisation algorithm. A distance based clustering technique was combined with a local outlier detection method to identify suspicious transactions in a synthetic dataset (<xref ref-type="bibr" rid="B12">Gao, 2009</xref>). (<xref ref-type="bibr" rid="B34">Paula et&#x20;al., 2016</xref>) developed an unsupervised deep learning algorithm based on an autoencoder to detect anomalous transactions in Brazilian exports financial&#x20;data.</p>
<p>In addition, several works follow AML approaches that focus on identifying accounts and customers rather than transactions. Social network analysis was used for identifying the roles and responsibilities of criminals in money laundering networks (<xref ref-type="bibr" rid="B8">Dre&#x17c;ewski et&#x20;al., 2015</xref>), and for analysing overall group properties of criminal networks (<xref ref-type="bibr" rid="B36">Savage et&#x20;al., 2016</xref>). Moreover, (<xref ref-type="bibr" rid="B56">Zhou et&#x20;al., 2017</xref>) developed a statistical classifier for the detection of suspicious accounts involved in illicit virtual currency trade, while achieving a low rate of false positives in their approach. While (<xref ref-type="bibr" rid="B50">Wang and Yang, 2007</xref>) implemented decision trees based approach to determine the money laundering risk levels of customers of a commercial bank, only the customers&#x2019; profiles were used to fit the model without taking into account their transactional&#x20;data.</p>
<p>In this work, we combine the benefits of supervised and unsupervised learning to cluster and categorise our bank&#x2019;s clients into risk level groups. It is noteworthy to recognise that in the works mentioned above, the low proportion of money laundering transaction was mainly resolved by oversampling or undersampling prior to the implementation of the methods. In our work, we instead use datasets with several class ratios to demonstrate that our approach is reasonably effective on both heavily imbalanced datasets and balanced datasets. Additionally, we present an adaptive approach that considers the level of suspicion at an account level instead of analysing every transaction individually. This is necessary, since a significant proportion of money laundering transactions have very similar characteristics to ordinary transactions. The consideration of an account&#x2019;s history in an aggregated manner allows the inclusion of historic patterns that could be linked to money laundering behaviour. Moreover, the aforementioned studies presented models that were aimed to replace human expertise by attempting to replicate their decisions, which is essentially problematic due to the challenges with manually identifying money laundering transactions. Instead, we propose a method that ranks the suspicious level of accounts in order to assist compliance teams at financial institutions with prioritising the clients to further investigate.</p>
</sec>
<sec id="s2-2">
<title>Self-Organising Maps</title>
<p>A self-organising map is an unsupervised neural network that maps a multi-dimensional dataset&#x20;along a lower-dimensional grid (<xref ref-type="bibr" rid="B25">Kohonen, 1990</xref>). Due to their structural properties, SOMs have been widely implemented for different use-cases in various industries. (<xref ref-type="bibr" rid="B53">Yorek et&#x20;al., 2016</xref>) combined SOM with ward clustering to classify living organisms into three distinguished clusters. SOM for clustering was also used for classifying natural language written texts into their respective document types (<xref ref-type="bibr" rid="B33">Pacella et&#x20;al., 2016</xref>). Furthermore, (<xref ref-type="bibr" rid="B20">Kiang et&#x20;al., 2006</xref>) applied SOM on telecommunication questionnaire responses in order to cluster the respondents into several segments by using K-means clustering. (<xref ref-type="bibr" rid="B27">Lacerda and Mello, 2013</xref>) presented a framework that segments handwritten digits using SOM for accurate digit recognition.</p>
<p>In addition to clustering, self-organising maps have been increasingly employed for oversampling the minority class&#x2019;s observations when a class imbalance exists (<xref ref-type="bibr" rid="B21">Kim, 2007</xref>; <xref ref-type="bibr" rid="B5">Cai et&#x20;al., 2014</xref>; <xref ref-type="bibr" rid="B7">Douzas and Bacao, 2017</xref>). Whereas, (<xref ref-type="bibr" rid="B55">Zadeh Shirazi et&#x20;al., 2018</xref>) developed a framework using SOM and complex-valued neural network for the diagnosis and detection of breast cancer using a dataset consisting of five features for hundreds of patients. Meanwhile, (<xref ref-type="bibr" rid="B42">Stephanakis et&#x20;al., 2019</xref>) proposed an extensive SOM-based hybrid method that was used for anomaly detection in cloud computing structures. (<xref ref-type="bibr" rid="B2">Barreto, 2007</xref>) established that unlike other artificial neural network models, SOM-based models can be used for time-series prediction whilst benefiting from SOM&#x2019;s interpretable results.</p>
<p>Furthermore, SOMs have also been explored in the financial domain. In a study by (<xref ref-type="bibr" rid="B9">Du Jardin and S&#xe9;verin, 2011</xref>), SOM was used to forecast the financial models of companies over extended periods of time and demonstrated a better prediction accuracy than traditional methods. SOMs were also employed for predicting bankruptcy by analysing financial records of several companies (<xref ref-type="bibr" rid="B22">Kiviluoto, 1998</xref>). Meanwhile, (<xref ref-type="bibr" rid="B15">Hsu, 2011</xref>) presented a hybrid approach using self-organising maps and genetic programming for predicting stock closing prices. In our work, we draw inspiration from such studies to demonstrate the capability of a SOM-based model in identifying potential money launderers, in contrast to the traditional supervised learning methods which rely heavily on the assigned labels for classification. The employment of SOM in our study stems from the need for grouping of accounts with similar transactional attributes together, in order for the investigators to able to make sense of these clusters. The clustering is followed by the categorisation of the lower dimensional grid to assign a risk level for every client in our dataset which are subsequently used for measuring the performance of the proposed investigation strategies.</p>
</sec>
</sec>
<sec id="s3">
<title>3 Data</title>
<p>In this work, we use the financial transactional data provided to us by our partner bank, DNB, the largest financial group in Norway. The data made available to us for this study represents a fraction of the total financial transactions handled by the bank between January 2014 and December 2016, with the bank clients as the main party and either bank clients or external accounts as the second party. The data has already been labelled by the bank, such that a class label exists for each transaction as to whether the transaction is normal or suspicious. In this context, suspicious transactions are not the alerts generated by the rule-based system, but are the more serious cases which were carefully investigated by the DNB&#x2019;s compliance team as potential money laundering cases - most of which were reported to the financial authorities. Subsequently, the labels in our dataset do not reveal whether the suspicious cases were indeed money laundering cases, since these decisions are made separately by the financial authorities and their outcome is not made available to us in the provided dataset. To this end, we treat every transaction that was thoroughly investigated by the compliance team at the bank as a suspicious transaction.</p>
<sec id="s3-1">
<title>3.1 Aggregation</title>
<p>In practice, it is almost impossible to make a decision purely based on transaction features of an individual transaction when investigating a particular case. Investigators often look at history of the party involved to observe if there are any underlying patterns or changes in a client&#x2019;s financial activity (<xref ref-type="bibr" rid="B40">Singh and Best, 2019</xref>). In this work, we relatively follow the investigators&#x2019; approach by aggregating the financial transactions data on the accounts. The original transactional dataset consists of a combination of categorical, numerical and mixed type features. We refine the dataset to eliminate the redundant and duplicate features. The outline of the refined dataset is presented in <xref ref-type="table" rid="T1">Table&#x20;1</xref>.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Dataset attributes prior to aggregation.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Attribute type</th>
<th align="center">Number of attributes</th>
<th align="center">Attribute names</th>
<th align="center">Unique categories</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Identifier</td>
<td align="center">1</td>
<td align="left">Account number</td>
<td align="left"/>
</tr>
<tr>
<td align="left">Date</td>
<td align="center">1</td>
<td align="left">Transaction date</td>
<td align="left"/>
</tr>
<tr>
<td align="left">Numerical</td>
<td align="center">1</td>
<td align="left">Transaction amount</td>
<td align="left"/>
</tr>
<tr>
<td rowspan="9" align="left">Categorical</td>
<td rowspan="9" align="center">9</td>
<td align="left">CreditDebitCode</td>
<td align="center">2</td>
</tr>
<tr>
<td align="left">TransTypeID</td>
<td align="center">29</td>
</tr>
<tr>
<td align="left">TransMethodID</td>
<td align="center">21</td>
</tr>
<tr>
<td align="left">TransactionChannelID</td>
<td align="center">28</td>
</tr>
<tr>
<td align="left">TextCode</td>
<td align="center">84</td>
</tr>
<tr>
<td align="left">ProductCode</td>
<td align="center">292</td>
</tr>
<tr>
<td align="left">CurrencyCodeOrig</td>
<td align="center">112</td>
</tr>
<tr>
<td align="left">SourceSystemOrg</td>
<td align="center">110</td>
</tr>
<tr>
<td align="left">SourceSystemFetch</td>
<td align="center">3</td>
</tr>
<tr>
<td align="left">Binary</td>
<td align="center">1</td>
<td align="left">Class label</td>
<td align="left"/>
</tr>
</tbody>
</table>
</table-wrap>
<p>The only identifier variable we select is the account number of the first party for every transaction in the dataset. To aggregate the data by account, we first apply one-hot encoding to the categorical variables. This increases the number of categorical variables from 9 variables to 681 variables. During the process of aggregation, these categorical variables are converted to the frequency of occurrence of every one-hot encoded feature for every account&#x2019;s transaction history between 2014 and 2016. The final feature is the binary label feature that indicates whether an account had a transaction that was investigated as a potential money laundering transaction. For accounts without any suspicious cases, all the transactional data was used in the aggregation process. Accounts that were involved in suspicious transaction had their data aggregated from their first transaction until the last suspicious transaction. This was done for two reasons. First, in some cases banks allow their customers to carry on with their financial activity while an investigation is pending. Second, for suspicious customers we are mainly interested in the transactions that led to the investigations, therefore, we are not concerned with the transaction activity that took place afterwards.</p>
</sec>
<sec id="s3-2">
<title>3.2 Dataset Size</title>
<p>When aggregating by accounts, we decide to eliminate bank accounts with more than 50,000 transactions over a 3&#xa0;year period from the data. This is supported by the fact that accounts used by large corporations are quite distinct from the majority of personal and corporate accounts. As such, outlying behaviour was removed to maintain the focus on the classification of bank accounts of customers with average account&#x20;usage.</p>
<p>The ratio of money laundering transactions is tremendously low compared to the normal transactions. As our partner bank, DNB, would not like to disclose the true ratio of suspicious clients in the dataset, we set the ratio of suspicious accounts in the dataset at % 10 of the total accounts in our baseline model. A subset of the suspicious accounts from our data is chosen such that 1,141 suspicious accounts are represented in the dataset. The remaining 10,269 accounts in the baseline model are ordinary accounts that were not investigated for money laundering by the&#x20;bank.</p>
</sec>
<sec id="s3-3">
<title>3.3 Preprocessing</title>
<p>Additional features are generated in order to embed a combination of the most significant attributes for every account. For every observation we engineer a total of 17 features, which include attributes such as the average number of daily credited transactions, average amount per transaction and the cumulative sum of debited transaction. The engineered features are first normalised by taking the natural logarithm of their values &#x2b;1. We then normalise all the engineered features in order for the values to fall in the [0,1] scale. Therefore, after dropping the date and the identifier attributes, all the features in the preprocessed dataset become in the [0,1] range, as the one-hot encoded ratios features are already within the same scale. We then drop all the single-valued features for all the 11,410 observations. Subsequently, we end up with 522 features after dropping all the non-unique features.</p>
</sec>
<sec id="s3-4">
<title>3.4 Training and Test Sets</title>
<p>We divide our baseline dataset which consists of 11,410 accounts and 522 features into training and test sets. In the split, we use a 80/20 training/test set ratio in a 5-fold cross-validation. Additionally, we ensure the ratio of suspicious customers is exactly 10% in both the training and test sets. The labels are initially removed from both the training and test sets and are reattached afterwards for evaluation. Later in our work, we measure the performance of our method using the mean value of the classification metrics of the 5-fold cross-validations, and a confusion matrix obtained by adding the predicted and actual labels of the test samples from all the cross-validation&#x20;runs.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Methodology</title>
<p>In our work, we implement a 4-stage approach in order to predict whether every bank client in our dataset might be involved in a money laundering case or not. We use the preprocessed dataset described in <xref ref-type="sec" rid="s4">Section 4</xref> for fitting and evaluating the presented model. <xref ref-type="fig" rid="F1">Figure&#x20;1</xref> depicts an overview of the stages of our proposed method.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Architecture of the SOM money laundering classification framework.</p>
</caption>
<graphic xlink:href="frai-04-761925-g001.tif"/>
</fig>
<sec id="s4-1">
<title>4.1 Feature Selection</title>
<p>We use an ensemble of feature selection algorithms in order to select the most significant features of the training set contributing to the label of the observations. This is important, as the dimensionality of the dataset can be reduced by discarding the noisy and redundant variables (<xref ref-type="bibr" rid="B13">Guyon et&#x20;al., 2008</xref>). Five feature selection algorithms are implemented on the dataset which rank the features according to their importance. Since the feature selection algorithms rank the features differently, we use an ensemble of the feature selectors. The ensemble computes the average ranking of every feature from the five feature selection methods, and subsequently, the features are ranked according to their mean rankings.</p>
<sec id="s4-1-1">
<title>4.1.1 L2 Regularisation</title>
<p>Regularisation is a technique used in machine learning to increase the training error in order to improve the generalisation on unseen data by preventing the overfitting of data. In L2 regularisation, also known as ridge regularisation, a loss function is computed using the weight coefficients of features and a bias term. The weights are updated using gradient descent optimisation to minimise the loss function. Since the labels are binary, the error function is also the log loss function used in logistic regression to predict the labels. When the regularisation converges, the features with greater weights are considered as the more important ones in the dataset.</p>
</sec>
<sec id="s4-1-2">
<title>4.1.2 Gini Impurity</title>
<p>The Gini impurity for feature importance is computed using the random forest classifier (<xref ref-type="bibr" rid="B3">Breiman, 2001</xref>). A random forest is an ensemble of decision trees, where each tree is a set of nodes and leaves. In classification problems, a given node in a decision tree is characterised by a feature with a certain threshold to decide how to split the dataset into two sets based on how the observations over and below the feature threshold are attributed to the different classes. The Gini impurity for each feature calculates the importance as the sum over the number of splits among the trees comprising the feature, proportionally to the number of samples the feature splits.</p>
</sec>
<sec id="s4-1-3">
<title>4.1.3 ANOVA F-Score</title>
<p>Analysis of Variance, ANOVA, is a univariate F-test that compares the mean of a feature between two or more label classes. The univariance property implies that the means and variances for a single feature against the target class are computed at a given time. The null hypothesis is that a feature&#x2019;s population means is the same for all the different labels. The method is commonly used as a statistical feature selection to measure how well individual features in a dataset can contribute to the target class labels.</p>
</sec>
<sec id="s4-1-4">
<title>4.1.4 Fisher Score</title>
<p>The Fisher score is a similarity-based feature selection method that computes the feature importance, and is typically used in binary classification applications (<xref ref-type="bibr" rid="B10">Hart et&#x20;al., 2001</xref>). The algorithm works by calculating the Laplacian matrix for a given feature using the local affinity and the global affinity matrices for all the observations and their labels. The Fisher score is then calculated using the inverse of the Laplacian score generated from the Laplacian matrix. Higher scores are awarded to features that well separate instances of different classes from each other, while drawing the observations of the same class closer together.</p>
</sec>
<sec id="s4-1-5">
<title>4.1.5 FCBF</title>
<p>Fast Correlation-Based Filter Solution, FCBF, is a mutlivariate feature selection technique developed by (<xref ref-type="bibr" rid="B54">Yu and Liu, 2003</xref>) that calculates the importance of features by estimating the interdependencies between them. The dependencies between a set of features can be calculated using a Symmetrical Uncertainty (SU) value based on information gain and entropy measures. In this fast feature selection approach, features which are given a higher score, implying more importance, are those which are highly correlated to the class labels, but are less correlated to other features.</p>
</sec>
</sec>
<sec id="s4-2">
<title>4.2&#x20;Self-Organising Map</title>
<p>A SOM is an unsupervised neural network algorithm developed by (<xref ref-type="bibr" rid="B25">Kohonen, 1990</xref>), where a high dimension dataset is typically mapped into a two dimensional representation arranged in either a rectangular or a hexagonal topology. Extensions to the traditional SOM include the time-adaptive self-organising map (<xref ref-type="bibr" rid="B38">Shah-Hosseini and Safabakhsh, 2003</xref>), which is a dynamic implementation of SOM that updates the learning parameters as more datapoints are added or modified over a period of time. However, due to the scarcity of observations and the class imbalance in our dataset, we implement a framework based on the traditional SOM instead of the time-adaptive self-organising maps. After selecting only the most relevant features in the dataset as highlighted in <xref ref-type="sec" rid="s5-1">Section 5.1</xref>, the training labels are detached from the dataset, such that the training of the SOM model is carried out without the class labels.</p>
<sec id="s4-2-1">
<title>4.2.1 SOM Description</title>
<p>The SOM consists of two layers: the prototype input layer and the output layer. The number of neurons in the output layer are determined by selecting the respective dimension parameters when creating the SOM. Each neuron is represented by a high-dimensional vector in the input layer, where the dimension size of each prototype vector is equivalent to the number of features in the data. The training of the SOM algorithm is implemented by the steps described below:</p>
<sec id="s4-2-1-1">
<title>1 Initialisation</title>
<p>The weights of the input prototype vector is initialised. This is done by either assigning the weights randomly, sampling observations from the data as the weights or by using linear methods such as the first two principal components of the principal component analysis for assigning the initial weights to the prototype vector.</p>
</sec>
<sec id="s4-2-1-2">
<title>2 Choosing a Random Sample</title>
<p>An observation from the dataset is chosen at random for training the weights of the SOM layers.</p>
</sec>
<sec id="s4-2-1-3">
<title>3 Matching</title>
<p>The best matching unit (BMU) is found by computing the Euclidean distance between the observation and the prototype vectors corresponding to the neurons in the output space. The prototype vector that is the closest to the observation, denoted by the minimum distance, will assign its neuron in the outer layer as the BMU. This is described by:<disp-formula id="e1">
<mml:math id="m1">
<mml:mo>&#x2225;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2225;</mml:mo>
<mml:mspace width="0.22em"/>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>min</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2225;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2225;</mml:mo>
</mml:math>
<label>(1)</label>
</disp-formula>where <italic>x</italic>
<sub>
<italic>j</italic>
</sub> is the observation vector, <italic>c</italic> is the index of the BMU, <italic>m</italic>
<sub>
<italic>c</italic>
</sub> is the prototype vector of the BMU, <italic>i</italic> is the number of neurons, and <italic>m</italic>
<sub>
<italic>i</italic>
</sub> is the prototype vector corresponding to neuron&#x20;<italic>i</italic>.</p>
</sec>
<sec id="s4-2-1-4">
<title>4 Neighbourhood Calculation</title>
<p>The neighbouring neurons of the BMU are determined. In the first stages of the training, the neurons are relatively close to each other. The distance between neurons increases over time, thus the number of neighbours decreases.</p>
</sec>
<sec id="s4-2-1-5">
<title>5 Weight Updating</title>
<p>In this stage, the BMU is rewarded by closely matching the observation vector. Neighbouring neurons are also matched to the observation sample, but to a lesser extent. The SOM update rule for the prototype vector <italic>i</italic> is:<disp-formula id="e2">
<mml:math id="m2">
<mml:msub>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mspace width="0.22em"/>
<mml:msub>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mspace width="0.22em"/>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>where <italic>t</italic> is the time step, <italic>&#x3b1;</italic> is the learning rate at time <italic>t</italic>, and <italic>h</italic>
<sub>
<italic>ci</italic>
</sub> is the neighbourhood function. The learning rate parameter is selected when choosing the model hyperparameters and fall within the [0,1] range. The neighbourhood function weights the neighbourhood kernel around the BMU in the output map and is usually in the form of a Gaussian function. The neighbourhood function at a given time step can be calculated as:<disp-formula id="e3">
<mml:math id="m3">
<mml:msub>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="italic">exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mo>&#x2225;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msup>
<mml:mrow>
<mml:mo>&#x2225;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(3)</label>
</disp-formula>where <italic>r</italic>
<sub>
<italic>c</italic>
</sub> is the position vector of the BMU neuron, <italic>r</italic>
<sub>
<italic>i</italic>
</sub> is the position vector of the other neurons in the SOM output map, and <italic>&#x3c3;</italic>(<italic>t</italic>) is the neighbourhood spread radius function at time&#x20;<italic>t</italic>.</p>
</sec>
<sec id="s4-2-1-6">
<title>6 Iteration</title>
<p>Steps 2&#x2013;5 are repeated based on the number of iterations specified before training the SOM algorithm.</p>
</sec>
</sec>
<sec id="s4-2-2">
<title>4.2.2 SOM Hyperparameter Selection</title>
<p>The dimensions of the SOM grids are selected according to <inline-formula id="inf1">
<mml:math id="m4">
<mml:mn>5</mml:mn>
<mml:msqrt>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msqrt>
</mml:math>
</inline-formula> rule proposed by (<xref ref-type="bibr" rid="B48">Vesanto and Alhoniemi, 2000</xref>), where <italic>N</italic>
<sub>
<italic>tr</italic>
</sub> is the number of training observation samples used for SOM training. The dataset used in the baseline model consists of 11,410 observations, 80% of which are used in each cross-validation training, hence 9,128 samples. This gives us <inline-formula id="inf2">
<mml:math id="m5">
<mml:mn>5</mml:mn>
<mml:msqrt>
<mml:mrow>
<mml:mn>9128</mml:mn>
</mml:mrow>
</mml:msqrt>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>477.70</mml:mn>
</mml:math>
</inline-formula> neurons. Applying the square root and rounding to closest integer gives us 22 neurons in each dimension, thus, a 22&#x20;&#xd7; 22 SOM grid for the baseline&#x20;model.</p>
<p>We choose a hexagonal topology for the SOM grid implementation since each non-edge neuron is surrounded by six neurons, rather than four neighbouring neurons commonly computed in the rectangular topology. In addition, we choose a Gaussian function for updating the neurons&#x2019; prototype weights since the Gaussian function allows all the neurons to be updated, not just the ones in close proximity to the BMU. For the learning rate and the neighbourhood spread radius function, we tune these hyperparameters on the training set prior to training the SOMs by calculating the quantisation error of the grid over a range of values for the two parameters. Since we train our model using 5-fold cross validations, we train five different SOMs, each with its unique learning rate and neighbourhood spread parameters based on the quantisation error generated from the training data in every cross-validation&#x20;run.</p>
</sec>
</sec>
<sec id="s4-3">
<title>4.3 Categorisation</title>
<p>Following the model training, the neurons in the SOM grid are categorised into three risk levels based on their distance from neurons in the neighbourhood, which we refer to as the inter-neural distance, and their suspicious ratio composition. For a given neuron, the inter-neural distance is the normalised sum of Euclidean distances between the neuron&#x2019;s weight vector, <italic>m</italic>
<sub>
<italic>i</italic>
</sub> from <xref ref-type="disp-formula" rid="e2">Eq. 2</xref>, and its neighbouring weight vectors when the model converges. Neurons with larger inter-neural distances are well separated from their neighbours, and therefore have more distinct properties. On the other hand, neurons with lower inter-neural distances have similar properties to their neighbouring neurons. Since the inter-neural distances of all SOM neurons fall in the [0,1] range, we divide this scale into five equally sized segments, <italic>P</italic>
<sub>1</sub>, &#x2026;, <italic>P</italic>
<sub>5</sub> for categorising the neuron&#x2019;s risk levels:<disp-formula id="e4">
<mml:math id="m6">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>0,0.2</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mspace width="0.28em"/>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mspace width="0.28em"/>
<mml:mn>0</mml:mn>
<mml:mo>&#x2264;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3c;</mml:mo>
<mml:mn>0.2</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>0.2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>0.4</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mspace width="0.28em"/>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mspace width="0.28em"/>
<mml:mn>0.2</mml:mn>
<mml:mo>&#x2264;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3c;</mml:mo>
<mml:mn>0.4</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>0.4</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>0.6</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mspace width="0.28em"/>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mspace width="0.28em"/>
<mml:mn>0.4</mml:mn>
<mml:mo>&#x2264;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3c;</mml:mo>
<mml:mn>0.6</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>0.6</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>0.8</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mspace width="0.28em"/>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mspace width="0.28em"/>
<mml:mn>0.6</mml:mn>
<mml:mo>&#x2264;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3c;</mml:mo>
<mml:mn>0.8</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>5</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>0.8</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>5</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mspace width="0.28em"/>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mspace width="0.28em"/>
<mml:mn>0.8</mml:mn>
<mml:mo>&#x2264;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2264;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(4)</label>
</disp-formula>where <italic>d</italic>
<sub>
<italic>i</italic>
</sub> is the inter-neural distance of neuron&#x20;<italic>i</italic>.</p>
<p>As discussed earlier, each training observation is assigned to its BMU. The suspicious composition of a SOM neuron is the proportion of suspicious observations from the total number of observation assigned to them following the reattachment of labels after training. This indicates that neurons with a high suspicious composition tend to mainly comprise suspicious observations, hence can be seen as highly suspicious neurons themselves. Meanwhile, neurons with a low suspicious composition are treated as nodes that mainly encapsulate normal accounts. In contrast to the inter-neural distance, we divide the suspicious composition [0,1] range into three segments, <italic>Q</italic>
<sub>1</sub>, <italic>Q</italic>
<sub>2</sub>, <italic>Q</italic>
<sub>3</sub>, that vary in size. The segments are defined as follows:<disp-formula id="e5">
<mml:math id="m7">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>Q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>Q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mspace width="0.28em"/>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mspace width="0.28em"/>
<mml:mn>0</mml:mn>
<mml:mo>&#x2264;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>Q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>z</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>Q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mspace width="0.28em"/>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mspace width="0.28em"/>
<mml:mi>z</mml:mi>
<mml:mo>&#x2264;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>z</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>Q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b2;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>Q</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mspace width="0.28em"/>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mspace width="0.28em"/>
<mml:mi>z</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b2;</mml:mi>
<mml:mo>&#x2264;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2264;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(5)</label>
</disp-formula>where <italic>z</italic> is the ratio of suspicious accounts in our dataset, <italic>k</italic>
<sub>
<italic>i</italic>
</sub> is the suspicious composition of neuron <italic>i</italic>, and <italic>&#x3b2;</italic> is the boundary threshold. In principle, the boundary threshold, <italic>&#x3b2;</italic>, can take the form of any value within [0,1]. However, we set <italic>&#x3b2;</italic> &#x3d; 0.25 in our SOM implementation to ensure the neurons in <italic>Q</italic>
<sub>2</sub> have a comparable representation, similar to neurons in segments <italic>Q</italic>
<sub>1</sub> and&#x20;<italic>Q</italic>
<sub>3</sub>.</p>
<p>The inter-neural distance segments and suspicious composition segments are used for constructing the neuron categorisation matrix, such that every neuron belongs to a risk level. We establish that neurons with large inter-neural distances and large suspicious compositions are very likely attributed to observations that pose a high risk for money laundering. Meanwhile, neurons that are in close proximity to their neighbours and generally have a smaller proportion of suspicious accounts are unlikely to encapsulate suspicious observations. Nodes with intermediate inter-neural distance and suspicious composition values are considered as medium-risk nodes. Our risk level categorisation of neurons based on the inter-neural distance and suspicious composition is demonstrated in <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>SOM neurons risk categorisation matrix.</p>
</caption>
<graphic xlink:href="frai-04-761925-g002.tif"/>
</fig>
</sec>
<sec id="s4-4">
<title>4.4 Strategy Selection</title>
<p>Samples are the given the same risk level as their best matching neuron. In this way, we can observe how the samples are distributed along the risk categorisation matrix, just like the SOM neurons. Therefore, a given observation can either belong to the low-risk, medium-risk or high-risk money laundering category. Since we have three risk level categories, and two target class labels: normal or suspicious, we propose two strategies to embed the risk categories into the binary labels of the accounts. These two strategies represent the different approaches that can be adopted by the banks or financial institutions in order to prioritise the investigation of accounts over a period of time based on the desired considerations of the risk groups. This is also instrumental for evaluating the performance of our model and highlighting the differences between the two approaches. The proposed two strategies are:<list list-type="simple">
<list-item>
<p>&#x2022; <italic>Safe Strategy</italic>: Medium-risk and high-risk accounts are classified as suspicious.</p>
</list-item>
<list-item>
<p>&#x2022; <italic>Fast Strategy</italic>: Only high-risk accounts are classified as suspicious.</p>
</list-item>
</list>
</p>
<p>In the safe strategy, the accounts that were categorised in the medium-risk and high-risk groups are considered as potentially suspicious and are to be investigated as such. In this strategy, banks are willing to act safely by investigating the transaction patterns of accounts falling in the medium and high risk groups. In contrast, in the fast strategy banks prioritise only the accounts in the high risk group for investigations. This can decrease the volume of accounts to be investigated and the amount of time spent investigating them. We therefore use both strategies for evaluating the performance of our model, while comparing the metrics of both approaches.</p>
</sec>
</sec>
<sec id="s5">
<title>5 Results</title>
<p>We implement our model in Python using the Minisom library (<xref ref-type="bibr" rid="B49">Vettigli, 2021</xref>) to generate the self-organising maps and using the SKLearn library (<xref ref-type="bibr" rid="B35">Pedregosa et&#x20;al., 2011</xref>) to evaluate the performance of our proposed method. In this section, we present the results produced after the implementation of the baseline model, and we further evaluate the model&#x2019;s performance when experimenting with our dataset&#x2019;s structure.</p>
<sec id="s5-1">
<title>5.1 Baseline Model</title>
<p>In the baseline model, 11,410 accounts are split into a 80/20 training/test split in a 5-fold cross-validation. The ratio of suspicious accounts in the dataset is 10<italic>%</italic>, and the same ratio is maintained in the training and test cross-validation sets. In every cross-validation run we use the ensemble of feature selectors on the training set to select the top 25 ranked features generated from the five feature selectors in our dataset and discard the remaining features. The most important features from the training set are also selected for test data in the cross-validations prior to evaluation. This implies that a different set of features is generated for the training and test sets during every&#x20;run.</p>
<sec id="s5-1-1">
<title>5.1.1 Training Set Analysis</title>
<p>To provide an insight on the performance of our method on the baseline training set, we present the training results of our fifth cross-validation run. In this manner, the SOM plot and the neurons distribution along the risk categorisation matrix from one of the iterations can be visualised. <xref ref-type="fig" rid="F3">Figure&#x20;3</xref> demonstrates the unsupervised SOM generated from the training set with 25 selected features.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>The self-organising map plot after training. Dark grey cells in the plot represent neurons that are further away from their neighbouring neurons, while the light grey or white cells are those that are in close proximity to their neighbours. Darker intensity markers of either the red crosses or green circles indicate the attribution of many observations of the same class to a given neuron.</p>
</caption>
<graphic xlink:href="frai-04-761925-g003.tif"/>
</fig>
<p>It can be observed from the plot in <xref ref-type="fig" rid="F3">Figure&#x20;3</xref> that suspicious accounts, marked with red crosses in the plot, appear to accumulate towards the upper right side of the SOM. This is not exclusive to suspicious accounts, as other normal accounts can be seen attributed to the same neurons as the suspicious accounts. We can also observe how few suspicious observations are scattered around the grid. A distinct boundary that consists of neurons with large distances from each other can also be visualised from the self-organising map. These neurons have a large inter-neural distance, and they appear to divide the map into several regions. It can also be observed how a large region under the boundary on the lower left side is entirely composed of normal accounts. This can give an intuition that the some of the normal accounts have features that significantly separates them from suspicious accounts and other normal accounts. It is worth noting that the positioning of accounts, thus markers, vary between the training cross-validation runs since the training observations and the initialised weight vectors are changed. Nevertheless, the main properties of the SOM such as the dense suspicious regions and the partitions created by high inter-neural distance nodes are maintained throughout the cross-validations.</p>
<p>The SOM plot of the fifth cross-validation is used for categorising the neurons along the neurons risk categorisation matrix depicted in <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>. Since the baseline dataset comprises a 10% suspicious observations ratio, we use <italic>z</italic>&#x20;&#x3d; 0.10 for <xref ref-type="disp-formula" rid="e5">Eq. 5</xref> to calculate the boundaries of the suspicious composition segments: <italic>Q</italic>
<sub>1</sub> &#x3d; [0, 0.1), <italic>Q</italic>
<sub>2</sub> &#x3d; [0.1, 0.35) and <italic>Q</italic>
<sub>3</sub> &#x3d; [0.35, 1]. Subsequently, the neurons distribution along the risk categorisation matrix is provided in <xref ref-type="table" rid="T2">Table&#x20;2</xref>.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Fifth-fold cross-validation SOM neurons inter-neural distances against their suspicious accounts composition.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Inter-neural distance</th>
<th colspan="3" align="center">Suspicious composition</th>
<th rowspan="2" align="center">Total</th>
</tr>
<tr>
<th align="center">0&#x2013;9.99%</th>
<th align="center">10&#x2013;34.99%</th>
<th align="center">35&#x2b;%</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">
<inline-formula id="inf3">
<mml:math id="m8">
<mml:mo>&#x3c;</mml:mo>
<mml:mn>0.20</mml:mn>
</mml:math>
</inline-formula>
</td>
<td align="char" char=".">98</td>
<td align="char" char=".">7</td>
<td align="char" char=".">7</td>
<td align="char" char=".">112</td>
</tr>
<tr>
<td align="left">0.20&#x2013;0.39</td>
<td align="char" char=".">154</td>
<td align="char" char=".">40</td>
<td align="char" char=".">40</td>
<td align="char" char=".">234</td>
</tr>
<tr>
<td align="left">0.40&#x2013;0.59</td>
<td align="char" char=".">80</td>
<td align="char" char=".">6</td>
<td align="char" char=".">18</td>
<td align="char" char=".">104</td>
</tr>
<tr>
<td align="left">0.60&#x2013;0.79</td>
<td align="char" char=".">26</td>
<td align="char" char=".">2</td>
<td align="char" char=".">1</td>
<td align="char" char=".">29</td>
</tr>
<tr>
<td align="left">
<inline-formula id="inf4">
<mml:math id="m9">
<mml:mo>&#x2265;</mml:mo>
<mml:mn>0.80</mml:mn>
</mml:math>
</inline-formula>
</td>
<td align="char" char=".">5</td>
<td align="char" char=".">0</td>
<td align="char" char=".">0</td>
<td align="char" char=".">5</td>
</tr>
<tr>
<td align="left">Total</td>
<td align="char" char=".">363</td>
<td align="char" char=".">55</td>
<td align="char" char=".">66</td>
<td align="char" char=".">484</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>From <xref ref-type="table" rid="T2">Table&#x20;2</xref>, it can be noted how the vast majority of the neurons have an inter-neural distance of less than 0.6. This implies that most neurons are in close vicinity to their neighbours, the only exception being the dark grey nodes in the SOM grid in <xref ref-type="fig" rid="F3">Figure&#x20;3</xref>. Moreover, most neurons tend to have a suspicious composition of less than 10<italic>%</italic>. This is reasonable, given that the ratio of suspicious accounts in our dataset is 10<italic>%</italic>. Categorising the risk level of neurons according to the matrix in <xref ref-type="fig" rid="F2">Figure&#x20;2</xref> reveals that 365 neurons are low-risk neurons, 60 neurons are medium-risk and 59 neurons are high-risk. Given that observations are matched with the neurons in the SOM grid, we further present the distribution of training accounts in the categorisation matrix in <xref ref-type="table" rid="T3">Table&#x20;3</xref>.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>The distribution of the fifth-fold cross-validation training accounts along the risk categorisation matrix. Notation of the numbers is: all accounts (suspicious accounts).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Inter-neural distance</th>
<th colspan="3" align="center">Suspicious composition</th>
<th rowspan="2" align="center">Total</th>
</tr>
<tr>
<th align="center">0&#x2013;9.99%</th>
<th align="center">10&#x2013;34.99%</th>
<th align="center">35&#x2b;%</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">
<inline-formula id="inf5">
<mml:math id="m10">
<mml:mo>&#x3c;</mml:mo>
<mml:mn>0.20</mml:mn>
</mml:math>
</inline-formula>
</td>
<td align="char" char="( .">3122 (27)</td>
<td align="char" char="( .">119 (21)</td>
<td align="char" char="( .">158 (111)</td>
<td align="char" char="( .">3399 (159)</td>
</tr>
<tr>
<td align="left">0.20&#x2013;0.39</td>
<td align="char" char="( .">2973 (29)</td>
<td align="char" char="( .">602 (109)</td>
<td align="char" char="( .">647 (418)</td>
<td align="char" char="( .">4222 (556)</td>
</tr>
<tr>
<td align="left">0.40&#x2013;0.59</td>
<td align="char" char="( .">855 (9)</td>
<td align="char" char="( .">65 (11)</td>
<td align="char" char="( .">233 (164)</td>
<td align="char" char="( .">1,153 (184)</td>
</tr>
<tr>
<td align="left">0.60&#x2013;0.79</td>
<td align="char" char="( .">265 (4)</td>
<td align="char" char="( .">52 (7)</td>
<td align="char" char="( .">4 (2)</td>
<td align="char" char="( .">321 (13)</td>
</tr>
<tr>
<td align="left">
<inline-formula id="inf6">
<mml:math id="m11">
<mml:mo>&#x2265;</mml:mo>
<mml:mn>0.80</mml:mn>
</mml:math>
</inline-formula>
</td>
<td align="char" char="( .">33 (0)</td>
<td align="char" char="( .">0 (0)</td>
<td align="char" char="( .">0 (0)</td>
<td align="char" char="( .">33 (0)</td>
</tr>
<tr>
<td align="left">Total</td>
<td align="char" char="( .">7248 (69)</td>
<td align="char" char="( .">838 (148)</td>
<td align="char" char="( .">1,042 (695)</td>
<td align="char" char="( .">9,128 (912)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Similar to the neurons distributions, it can be observed in <xref ref-type="table" rid="T3">Table&#x20;3</xref> how the majority accounts were attributed to neurons with low inter-neural distances and low suspicious compositions. We can see, however, that the majority of suspicious accounts were assigned to neurons with a large suspicious compositions. In other words, suspicious observations tend to be grouped together by the same neurons. Using this information, the risk grouping of training accounts is demonstrated in <xref ref-type="table" rid="T4">Table&#x20;4</xref>.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Risk categorisation of the fifth-fold cross-validation training observations.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Risk category</th>
<th align="center">All accounts</th>
<th align="center">Suspicious accounts</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Low</td>
<td align="center">7334</td>
<td align="center">90</td>
</tr>
<tr>
<td align="left">Medium</td>
<td align="center">910</td>
<td align="center">238</td>
</tr>
<tr>
<td align="left">High</td>
<td align="center">884</td>
<td align="center">584</td>
</tr>
<tr>
<td align="left">Total</td>
<td align="center">9,128</td>
<td align="center">912</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>It can be evident from <xref ref-type="table" rid="T3">Table&#x20;3</xref> how the training samples in the cross-validation run were categorised into the low, medium and high risk levels. We can see that while most of the non-suspicious accounts belonged to the low-risk group, 90 suspicious observations were misclassified as low-risk. In addition, the number of accounts in the low-risk category is significantly greater than total number of accounts in the medium and high-risk categories, which are comparable in size. The suspicious accounts, however, tend to be more concentrated in the high-risk category.</p>
</sec>
<sec id="s5-1-2">
<title>5.1.2 Test Set Analysis</title>
<p>To evaluate the performance of our model on the test data, we combine the results from all five cross-validations, such that every sample in our dataset is used only once for testing. As such, we can represent the test data distribution on the risk categorisation matrix and the risk level grouping in <xref ref-type="table" rid="T5">Tables 5</xref>, <xref ref-type="table" rid="T6">6</xref> respectively.</p>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>The distribution of test accounts along the risk categorisation matrix. Notation of the numbers is: all accounts (suspicious accounts).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Inter-neural distance</th>
<th colspan="3" align="center">Suspicious composition</th>
<th rowspan="2" align="center">Total</th>
</tr>
<tr>
<th align="center">0&#x2013;9.99%</th>
<th align="center">10&#x2013;34.99%</th>
<th align="center">35&#x2b;%</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">
<inline-formula id="inf7">
<mml:math id="m12">
<mml:mo>&#x3c;</mml:mo>
<mml:mn>0.20</mml:mn>
</mml:math>
</inline-formula>
</td>
<td align="center">2379 (33)</td>
<td align="center">94 (9)</td>
<td align="center">72 (43)</td>
<td align="center">2545 (85)</td>
</tr>
<tr>
<td align="left">0.20&#x2013;0.39</td>
<td align="center">2809 (46)</td>
<td align="center">513 (78)</td>
<td align="center">407 (247)</td>
<td align="center">3729 (371)</td>
</tr>
<tr>
<td align="left">0.40&#x2013;0.59</td>
<td align="center">2548 (64)</td>
<td align="center">513 (83)</td>
<td align="center">534 (324)</td>
<td align="center">3595 (471)</td>
</tr>
<tr>
<td align="left">0.60&#x2013;0.79</td>
<td align="center">940 (12)</td>
<td align="center">174 (37)</td>
<td align="center">238 (128)</td>
<td align="center">1,352 (177)</td>
</tr>
<tr>
<td align="left">
<inline-formula id="inf8">
<mml:math id="m13">
<mml:mo>&#x2265;</mml:mo>
<mml:mn>0.80</mml:mn>
</mml:math>
</inline-formula>
</td>
<td align="center">123 (5)</td>
<td align="center">11 (2)</td>
<td align="center">55 (30)</td>
<td align="center">189 (37)</td>
</tr>
<tr>
<td align="left">Total</td>
<td align="center">8799 (160)</td>
<td align="center">1,305 (209)</td>
<td align="center">1,306 (772)</td>
<td align="center">11,410 (1,141)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T6" position="float">
<label>TABLE 6</label>
<caption>
<p>Risk categorisation of test observations.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Risk category</th>
<th align="center">All accounts</th>
<th align="center">Suspicious accounts</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Low</td>
<td align="center">8770</td>
<td align="center">164</td>
</tr>
<tr>
<td align="left">Medium</td>
<td align="center">1,395</td>
<td align="center">246</td>
</tr>
<tr>
<td align="left">High</td>
<td align="center">1,245</td>
<td align="center">731</td>
</tr>
<tr>
<td align="left">Total</td>
<td align="center">11,410</td>
<td align="center">1,141</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>From <xref ref-type="table" rid="T5">Table&#x20;5</xref> we can notice how the testing accounts are spread across the risk categorisation matrix. Similar to the training samples distribution, most of the accounts were assigned to BMUs with a low suspicious composition. An interesting observation, however, is that many accounts in the cross-validation test sets were assigned to neurons with inter-neural distances between 0.4 and 0.59, compared to the total training samples under the same inter-neural distance in <xref ref-type="table" rid="T3">Table&#x20;3</xref>. This entails that the SOM model performs a slightly different matching between samples and neurons when the model is learning from the input, compared to when the test data is applied to the trained model. The risk levels of the test accounts after all five cross-validations in <xref ref-type="table" rid="T6">Table&#x20;6</xref> appear similar to the risk grouping of training accounts from <xref ref-type="table" rid="T4">Table&#x20;4</xref>. Most of the suspicious accounts were categorised as high-risk, while evidently smaller populations were assigned with low and medium-risk categories.</p>
<p>Our evaluation of the model relies on the investigation strategies proposed in Section 5.4, since we formulated three risk levels for the binary class labels in our dataset. We recall that in the safe strategy, the medium and high-risk accounts are classified as suspicious, whereas, the fast strategy only considers the high-risk accounts as suspicious. On that basis, we use the risk level categorisation of the test observations to construct a confusion matrix for the safe strategy and a confusion matrix for the fast strategy represented in <xref ref-type="table" rid="T7">Tables 7</xref>, <xref ref-type="table" rid="T8">8</xref>, respectively.</p>
<table-wrap id="T7" position="float">
<label>TABLE 7</label>
<caption>
<p>Safe strategy confusion matrix.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left"/>
<th rowspan="2" align="left"/>
<th colspan="2" align="center">Predicted</th>
</tr>
<tr>
<th align="center">Normal</th>
<th align="center">Suspicious</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="2" align="left">
<bold>Actual</bold>
</td>
<td align="left">
<bold>Normal</bold>
</td>
<td align="center">8606</td>
<td align="center">1,663</td>
</tr>
<tr>
<td align="left">
<bold>Suspicious</bold>
</td>
<td align="center">164</td>
<td align="center">977</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T8" position="float">
<label>TABLE 8</label>
<caption>
<p>Fast strategy confusion matrix.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left"/>
<th rowspan="2" align="left"/>
<th colspan="2" align="center">Predicted</th>
</tr>
<tr>
<th align="center">Normal</th>
<th align="center">Suspicious</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">
<bold>Actual</bold>
</td>
<td align="left">
<bold>Normal</bold>
</td>
<td align="center">9,637</td>
<td align="center">632</td>
</tr>
<tr>
<td align="left"/>
<td align="left">
<bold>Suspicious</bold>
</td>
<td align="center">405</td>
<td align="center">736</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As it can be seen from the confusion matrices, the majority of the accounts that were classified as low-risk are in fact normal accounts. The difference between both strategies is more evident when looking at the number of false positives and the number of false negatives. The safe strategy is more conservative and detects more suspicious accounts. The drawback, however, is that 1,663 normal accounts were incorrectly classified as suspicious. In the fast strategy only 632 accounts were misclassified as suspicious, which is a drastic improvement in terms of reducing the number of false positives. However, this comes at the cost of obtaining more than twice the number of false negatives of the safe strategy. <xref ref-type="table" rid="T9">Table&#x20;9</xref> presents the classification metrics for both strategies using the output of the confusion matrices.</p>
<table-wrap id="T9" position="float">
<label>TABLE 9</label>
<caption>
<p>Classification performance metrics for the safe and fast strategies.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Strategy</th>
<th align="center">Accuracy</th>
<th align="center">Precision</th>
<th align="center">Recall</th>
<th align="center">F1-score</th>
<th align="center">AUC</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Safe</td>
<td align="center">0.8399</td>
<td align="center">0.3728</td>
<td align="center">0.8562</td>
<td align="center">0.5188</td>
<td align="center">0.8472</td>
</tr>
<tr>
<td align="left">Fast</td>
<td align="center">0.9091</td>
<td align="center">0.5480</td>
<td align="center">0.6451</td>
<td align="center">0.5897</td>
<td align="center">0.7918</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>It can observed from <xref ref-type="table" rid="T9">Table&#x20;9</xref> how the performance metrics differ for both strategies. The accuracy score is higher for the fast strategy, which is reflected by the large number of correctly classified normal accounts, or true negatives. The elimination of the medium-risk category in the fast strategy also improved the precision score for the fast strategy since the medium-risk category contained a substantial amount of normal observations. The precision scores for both strategies is nevertheless still low as several normal accounts have transactional features which are similar to suspicious accounts. As expected, the safe strategy has a better recall score than the fast risk investigation strategy. Classifying the accounts in the medium-risk category as suspicious enabled the detection of more suspicious accounts, thus the higher recall score for the safe strategy in comparison with the fast strategy. The calculated F1-score asserted equal weights to the recall and precision scores. Therefore, a greater F1-score is obtained by the fast strategy. Meanwhile, the receiver operating characteristic area under curve (AUC) scores are comparable for both approaches, but slightly higher for the safe strategy as a result of the increased rate of the true positive classification on various thresholds.</p>
<p>From the classification metrics, the precision-recall trade-off can be very well observed. A more conservative classification based on risk categories enables the detection of more suspicious accounts. However, this comes at the expense of investigating a larger number of normal accounts. The fast strategy is concerned with reducing the number of normal accounts to be investigated. However, this leads to missing out on some of the suspicious accounts that were categorised earlier in the medium-risk group. The F1-score can be tuned according to preference, if any, over the rate of false positives and the rate of false negatives.</p>
</sec>
</sec>
<sec id="s5-2">
<title>5.2 Experimental Results</title>
<p>In principle, financial transaction datasets are heavily characterised by class imbalance where abnormal activity represents a small fraction of all transactions. As such, further experiments were carried out to analyse the model&#x2019;s performance when changing the dataset&#x2019;s structure. More precisely, it is interesting to study how the model behaves when tuning the ratio of suspicious accounts in our dataset. We run models on datasets with a suspicious class size of 5, 10, 20, and 50% of the total accounts in the datasets. To implement this, we select 500 suspicious observations to be used in all the experiments and modify the number of normal accounts according to the selected class ratios. Moreover, the number of selected features are tuned when training the SOMs, to investigate the impact of the number of features on the model&#x2019;s ability to correctly classify the observations.</p>
<p>Given the two proposed investigation strategies, we identify the most important metrics for evaluating the performance as the recall score and the precision score. Although the F1-score is a useful metric for combining the recall and the precision performances, we instead use the recall and precision scores to assess to what extent the investigation strategies are capable of achieving their objective. Similar to the baseline model, every experiment was run in 5-fold cross-validations and the mean values of the recall and precision scores were calculated. These are presented in <xref ref-type="table" rid="T10">Tables 10</xref>,&#x20;<xref ref-type="table" rid="T11">11</xref>.</p>
<table-wrap id="T10" position="float">
<label>TABLE 10</label>
<caption>
<p>Classification rate recall scores.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="3" align="left">Suspicious ratio (%)</th>
<th colspan="5" align="center">Safe strategy</th>
<th colspan="5" align="center">Fast strategy</th>
</tr>
<tr>
<th colspan="5" align="center">Number of Features</th>
<th colspan="5" align="center">Number of Features</th>
</tr>
<tr>
<th align="center">10</th>
<th align="center">25</th>
<th align="center">50</th>
<th align="center">100</th>
<th align="center">All</th>
<th align="center">10</th>
<th align="center">25</th>
<th align="center">50</th>
<th align="center">100</th>
<th align="center">All</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">5</td>
<td align="char" char=".">0.790</td>
<td align="char" char=".">0.802</td>
<td align="char" char=".">0.778</td>
<td align="char" char=".">0.824</td>
<td align="char" char=".">0.826</td>
<td align="char" char=".">0.368</td>
<td align="char" char=".">0.594</td>
<td align="char" char=".">0.516</td>
<td align="char" char=".">0.464</td>
<td align="char" char=".">0.432</td>
</tr>
<tr>
<td align="left">10</td>
<td align="char" char=".">0.836</td>
<td align="char" char=".">0.840</td>
<td align="char" char=".">0.850</td>
<td align="char" char=".">0.818</td>
<td align="char" char=".">0.828</td>
<td align="char" char=".">0.660</td>
<td align="char" char=".">0.624</td>
<td align="char" char=".">0.610</td>
<td align="char" char=".">0.580</td>
<td align="char" char=".">0.530</td>
</tr>
<tr>
<td align="left">20</td>
<td align="char" char=".">0.832</td>
<td align="char" char=".">0.826</td>
<td align="char" char=".">0.856</td>
<td align="char" char=".">0.838</td>
<td align="char" char=".">0.822</td>
<td align="char" char=".">0.680</td>
<td align="char" char=".">0.714</td>
<td align="char" char=".">0.644</td>
<td align="char" char=".">0.584</td>
<td align="char" char=".">0.596</td>
</tr>
<tr>
<td align="left">50</td>
<td align="char" char=".">0.792</td>
<td align="char" char=".">0.830</td>
<td align="char" char=".">0.842</td>
<td align="char" char=".">0.840</td>
<td align="char" char=".">0.864</td>
<td align="char" char=".">0.606</td>
<td align="char" char=".">0.644</td>
<td align="char" char=".">0.582</td>
<td align="char" char=".">0.692</td>
<td align="char" char=".">0.738</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T11" position="float">
<label>TABLE 11</label>
<caption>
<p>Classification rate precision scores.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="3" align="left">Suspicious ratio (%)</th>
<th colspan="5" align="center">Safe strategy</th>
<th colspan="5" align="center">Fast strategy</th>
</tr>
<tr>
<th colspan="5" align="center">Number of Features</th>
<th colspan="5" align="center">Number of Features</th>
</tr>
<tr>
<th align="center">10</th>
<th align="center">25</th>
<th align="center">50</th>
<th align="center">100</th>
<th align="center">All</th>
<th align="center">10</th>
<th align="center">25</th>
<th align="center">50</th>
<th align="center">100</th>
<th align="center">All</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">5</td>
<td align="char" char=".">0.228</td>
<td align="char" char=".">0.257</td>
<td align="char" char=".">0.238</td>
<td align="char" char=".">0.190</td>
<td align="char" char=".">0.160</td>
<td align="char" char=".">0.402</td>
<td align="char" char=".">0.467</td>
<td align="char" char=".">0.445</td>
<td align="char" char=".">0.297</td>
<td align="char" char=".">0.262</td>
</tr>
<tr>
<td align="left">10</td>
<td align="char" char=".">0.391</td>
<td align="char" char=".">0.392</td>
<td align="char" char=".">0.355</td>
<td align="char" char=".">0.323</td>
<td align="char" char=".">0.314</td>
<td align="char" char=".">0.564</td>
<td align="char" char=".">0.553</td>
<td align="char" char=".">0.544</td>
<td align="char" char=".">0.437</td>
<td align="char" char=".">0.436</td>
</tr>
<tr>
<td align="left">20</td>
<td align="char" char=".">0.530</td>
<td align="char" char=".">0.557</td>
<td align="char" char=".">0.560</td>
<td align="char" char=".">0.520</td>
<td align="char" char=".">0.528</td>
<td align="char" char=".">0.670</td>
<td align="char" char=".">0.661</td>
<td align="char" char=".">0.669</td>
<td align="char" char=".">0.608</td>
<td align="char" char=".">0.630</td>
</tr>
<tr>
<td align="left">50</td>
<td align="char" char=".">0.793</td>
<td align="char" char=".">0.814</td>
<td align="char" char=".">0.819</td>
<td align="char" char=".">0.819</td>
<td align="char" char=".">0.804</td>
<td align="char" char=".">0.849</td>
<td align="char" char=".">0.860</td>
<td align="char" char=".">0.837</td>
<td align="char" char=".">0.831</td>
<td align="char" char=".">0.845</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>
<xref ref-type="table" rid="T10">Table&#x20;10</xref> shows the recall scores for the safe and the fast strategies when varying the number of features and the ratio of suspicious accounts in the dataset. The highest classification recall scores for both strategies are usually achieved when selecting the most important 25 or 50 variables. The ratio of suspicious accounts had a low impact on the recall score of the safe strategy and a moderate impact on the fast strategy. Similar to the baseline model, the safe strategy was better at reducing the number of false negatives, hence the higher recall scores of the strategy.</p>
<p>In <xref ref-type="table" rid="T11">Table&#x20;11</xref>, the change of the precision scores as the model parameters are varied can be observed. Selecting the best 25 features generally improved the strategies&#x2019; ability to reduce the number of false positives. Unlike the recall scores, however, reducing the class imbalance appears to significantly enhance the precision scores for both strategies. This is expected since we maintained the same number of suspicious accounts while increasing the number of normal accounts when reducing the ratio of suspicious accounts, due to the scarcity of suspicious accounts in the dataset provided to us. As such, an increasing number of normal observations was attributed to SOM neurons in the high risk category when a low suspicious accounts ratio was used in the dataset.</p>
<p>Furthermore, we used the Welch&#x2019;s <italic>t</italic>-test (<xref ref-type="bibr" rid="B52">Welch, 1947</xref>) to investigate to what extent are the classification scores of the same strategy statistically significant when modifying the class ratios. To easily interpret the results of the significance test, we only used the scores of the experiments involving 25 selected features. In addition, we computed the statistical significance of a given strategy at once, since it is already expected that the strategies behave differently, thus, are already statistically significant. The <italic>p</italic>-values of the significance tests for the recall and precision scores are shown in <xref ref-type="table" rid="T12">Tables 12</xref>,&#x20;<xref ref-type="table" rid="T13">13</xref>.</p>
<table-wrap id="T12" position="float">
<label>TABLE 12</label>
<caption>
<p>Recall scores&#x2019; <italic>p</italic>-values between different suspicious ratios.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="3" align="left">Suspicious ratio (%)</th>
<th colspan="3" align="center">Safe strategy</th>
<th colspan="3" align="center">Fast strategy</th>
</tr>
<tr>
<th colspan="3" align="center">Suspicious Ratio</th>
<th colspan="3" align="center">Suspicious Ratio</th>
</tr>
<tr>
<th align="center">10%</th>
<th align="center">20%</th>
<th align="center">50%</th>
<th align="center">10%</th>
<th align="center">20%</th>
<th align="center">50%</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">5</td>
<td align="char" char=".">0.299</td>
<td align="char" char=".">0.549</td>
<td align="char" char=".">0.457</td>
<td align="char" char=".">0.587</td>
<td align="char" char=".">0.059</td>
<td align="char" char=".">0.378</td>
</tr>
<tr>
<td align="left">10</td>
<td align="left"/>
<td align="char" char=".">0.564</td>
<td align="char" char=".">0.609</td>
<td align="left"/>
<td align="char" char=".">0.049</td>
<td align="char" char=".">0.609</td>
</tr>
<tr>
<td align="left">20</td>
<td align="left"/>
<td align="left"/>
<td align="char" char=".">0.879</td>
<td align="left"/>
<td align="left"/>
<td align="char" char=".">0.114</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T13" position="float">
<label>TABLE 13</label>
<caption>
<p>Precision scores&#x2019; <italic>p</italic>-values between different suspicious ratios.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="3" align="left">Suspicious ratio (%)</th>
<th colspan="3" align="center">Safe strategy</th>
<th colspan="3" align="center">Fast strategy</th>
</tr>
<tr>
<th colspan="3" align="center">Suspicious Ratio</th>
<th colspan="3" align="center">Suspicious Ratio</th>
</tr>
<tr>
<th align="center">10%</th>
<th align="center">20%</th>
<th align="center">50%</th>
<th align="center">10%</th>
<th align="center">20%</th>
<th align="center">50%</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">5</td>
<td align="char" char=".">0.000</td>
<td align="char" char=".">0.000</td>
<td align="char" char=".">0.000</td>
<td align="char" char=".">0.024</td>
<td align="char" char=".">0.000</td>
<td align="char" char=".">0.000</td>
</tr>
<tr>
<td align="left">10</td>
<td align="left"/>
<td align="char" char=".">0.000</td>
<td align="char" char=".">0.000</td>
<td align="left"/>
<td align="char" char=".">0.007</td>
<td align="char" char=".">0.000</td>
</tr>
<tr>
<td align="left">20</td>
<td align="left"/>
<td align="left"/>
<td align="char" char=".">0.000</td>
<td align="left"/>
<td align="left"/>
<td align="char" char=".">0.000</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>It is evident from <xref ref-type="table" rid="T12">Table&#x20;12</xref> that the recall scores of the safe strategy are not statistically significant as they are all greater than 0.05: the conventional <italic>p</italic>-value threshold of 5<italic>%</italic>. In the fast strategy, the recall scores are also statistically insignificant, with the exception being a suspicious ratio of 10% against 20%. Therefore, it can be inferred that varying the class ratio has a minimal impact on the ratio of false negatives of both strategies. In contrast, <xref ref-type="table" rid="T13">Table&#x20;13</xref> reveals how the class imbalance greatly affects the change in the ratio of false positive classifications. Observing the outstandingly low <italic>p</italic>-values when comparing the class ratios demonstrate that the classification of normal accounts tend to improve when the dataset is characterised by a lower class imbalance. Hence, the significance test results indicate that our model produces a consistent classification of the suspicious accounts, while the increasing misclassification of normal accounts when reducing the proportion of suspicious accounts is inevitable due to the increased number of observations in our dataset.</p>
</sec>
</sec>
<sec id="s6">
<title>6 Conclusion</title>
<p>In this work, we presented a method using self-organising maps to identify and detect suspicious accounts. In contrast to the other studies presented in the literature, we developed a model that detects money laundering activity in an imbalanced dataset while demonstrating robustness against the inadequately labelled data. The poor labelling is a general problem in money laundering datasets which stems from the fact that a remarkable proportion of money laundering transactions are undetected by conventional rule-based alert systems. Therefore, developing models based on supervised methods is not practical, as the class labels in transactional datasets do not necessarily capture the ground truths regarding whether a transaction was carried for funneling dirty money or not. To this end, we developed a framework based on SOMs to categorise the bank accounts in our dataset into three risk level groups. Our approach is adaptive, as we used the risk levels to propose two investigations strategies which can be employed by financial compliance teams in order to prioritise the investigation of suspicious accounts.</p>
<p>Evaluating the model demonstrated that self-organising maps tend to group the suspicious accounts together on the SOM grid, despite not reading the labels. The two proposed strategies allowed us to more evidently observe the false-negatives, false-positives dilemma that indeed exists in anti-money laundering practices. We highlight that our model presents a novel contribution to the literature by demonstrating its capability in detecting the majority of suspicious cases when choosing a safe investigation strategy, and reasonably reducing the number of false alerts when employing a less conservative strategy. In addition, unlike the classical black-box machine learning methods, SOMs enable the visualisation of samples along a low dimensional grid such that it is possible to observe the clusters of the grid and potentially categorise the accounts into more groups.</p>
<p>We note that while our framework presented a good detection of suspicious bank clients, our study is not without its limitations. First, the efficiency of the model drops when the there is a significant class imbalance in the dataset. The risk categorisation of our SOM is a semi-supervised task since we introduce a threshold based on the suspicious accounts ratio in the data. Consequently, while the model succeeds in clustering the suspicious observations together, many normal accounts are also attributed to these suspicious clusters, hence more time and resources spent on false alerts. Contrarily, some of the suspicious accounts are also classified as normal, which carries a much higher cost for financial institutions for failing to report suspicious activity. Secondly, due to the scarcity of suspected money laundering accounts, we used the binary labels to combine all suspicious accounts together. In practice, suspected money launderers are investigated differently based on a range of factors: corporate or individual accounts, daily or savings accounts, local or international transactions.</p>
<p>Future extensions to this work can include embedding more features such as account type, previous bankruptcies and account creation date, which might contribute to more distinction between normal and money laundering activity. In addition, we aim to obtain more data in order to use SOM for generating more clusters that can strengthen the understanding of the various underlying financial criminal behaviours. As a potential extension to this work, we plan to explore the impact of combining alternative unsupervised approaches such as growing neural gas (<xref ref-type="bibr" rid="B11">Fritzke et&#x20;al., 1995</xref>) with supervised learning models for investigating money laundering accounts. Another promising direction for future works is incorporating tools commonly used in information retrieval systems such as the HITS algorithm (<xref ref-type="bibr" rid="B23">Kleinberg et&#x20;al., 2011</xref>) for ranking the bank accounts based on suspicion level, which can help in prioritising the order by which clients are investigated.</p>
</sec>
</body>
<back>
<sec id="s7">
<title>Data Availability Statement</title>
<p>The data analyzed in this study is subject to the following licenses/restrictions: The data used for the purpose of this study was provided to us by our partner bank, DNB, and is subject to privacy and confidentiality agreements that prohibits us from publicly sharing it. Requests to access these datasets should be directed to Karl Aksel Fest&#xf8;, <email>karl.aksel.festo@dnb.no</email>.</p>
</sec>
<sec id="s8">
<title>Author Contributions</title>
<p>AA: Conceptualization, Methodology, Software, Writing (Original Draft), and Visualization. AR: Conceptualization, Methodology, Supervision, Writing (Review), and Editing.</p>
</sec>
<sec id="s9">
<title>Funding</title>
<p>This work was supported by DNB through the funding of this project and providing the data for the purpose of this research.</p>
</sec>
<sec sec-type="COI-statement" id="s10">
<title>Conflict of Interest</title>
<p>Author AR was employed by company SINTEF Digital. The authors declare that this study received funding from DNB. The funder had the following involvement in the study: data sharing, supervision, revision of manuscript and approval to submit this study for publication.</p>
<p>The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s11">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>We express our gratitude to Aria Rahmati for his input and valuable contribution. We would also like to thank Frank Westad and Damiano Varagnolo for their coordination and useful remarks.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Argentiero</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bagella</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Busato</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Money Laundering in a Two-Sector Model: Using Theory for Measurement</article-title>. <source>Eur. J.&#x20;L. Econ</source> <volume>26</volume>, <fpage>341</fpage>&#x2013;<lpage>359</lpage>. <pub-id pub-id-type="doi">10.1007/s10657-008-9074-6</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Barreto</surname>
<given-names>G. A.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Time Series Prediction with the Self-Organizing Map: A Review</article-title>. <source>Perspectives Of Neural-Symbolic Integration</source>, <fpage>135</fpage>&#x2013;<lpage>158</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-540-73954-8_6</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Breiman</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>Random Forests</article-title>. <source>Machine Learn.</source> <volume>45</volume>, <fpage>5</fpage>&#x2013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1023/a:1010933404324</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Breslow</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Hagstroem</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Mikkelsen</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Robu</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2017</year>). <source>The New Frontier in Anti&#x2013;money Laundering</source>. <publisher-loc>New York, NY</publisher-loc>. </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cai</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Man</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Imbalanced Evolving Self-Organizing Learning</article-title>. <source>Neurocomputing</source> <volume>133</volume>, <fpage>258</fpage>&#x2013;<lpage>270</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2013.11.010</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Dinh Van Khoa</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Nazir</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Teoh</surname>
<given-names>E. N.</given-names>
</name>
<name>
<surname>Karupiah</surname>
<given-names>E. K.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). &#x201c;<article-title>Exploration of the Effectiveness of Expectation Maximization Algorithm for Suspicious Transaction Detection in Anti-money Laundering</article-title>,&#x201d; in <source>2014 IEEE Conference on Open Systems (ICOS)</source>, <fpage>145</fpage>&#x2013;<lpage>149</lpage>. <pub-id pub-id-type="doi">10.1109/icos.2014.7042645</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Douzas</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Bacao</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Self-organizing Map Oversampling (Somo) for Imbalanced Data Set Learning</article-title>. <source>Expert Syst. Appl.</source> <volume>82</volume>, <fpage>40</fpage>&#x2013;<lpage>52</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2017.03.073</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dre&#x17c;ewski</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Sepielak</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Filipkowski</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>The Application of Social Network Analysis Algorithms in a System Supporting Money Laundering Detection</article-title>. <source>Inf. Sci.</source> <volume>295</volume>, <fpage>18</fpage>&#x2013;<lpage>32</lpage>. </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Du Jardin</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>S&#xe9;verin</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Predicting Corporate Bankruptcy Using a Self-Organizing Map: An Empirical Study to Improve the Forecasting Horizon of a Financial Failure Model</article-title>. <source>Decis. Support Syst.</source> <volume>51</volume>, <fpage>701</fpage>&#x2013;<lpage>711</lpage>. <pub-id pub-id-type="doi">10.1016/j.dss.2011.04.001</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fritzke</surname>
<given-names>B.</given-names>
</name>
<etal/>
</person-group> (<year>1995</year>). <article-title>A Growing Neural Gas Network Learns Topologies</article-title>. <source>Adv. Neural Inf. Process. Syst.</source>. <publisher-name>Citeseer</publisher-name> <volume>7</volume>, <fpage>625</fpage>&#x2013;<lpage>632</lpage>. </citation>
</ref>
<ref id="B12">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Gao</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2009</year>). &#x201c;<article-title>Application of Cluster-Based Local Outlier Factor Algorithm in Anti-money Laundering</article-title>,&#x201d; in <conf-name>2009 International Conference on Management and Service Science</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>4</lpage>. <pub-id pub-id-type="doi">10.1109/icmss.2009.5302396</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Guyon</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Gunn</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Nikravesh</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zadeh</surname>
<given-names>L. A.</given-names>
</name>
</person-group> (<year>2008</year>). <source>Feature Extraction: Foundations and Applications</source>, <volume>207</volume>. <publisher-name>Springer</publisher-name>. </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Haga</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Siekkinen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Sundvik</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Initial Stage Clustering when Estimating Accounting Quality Measures with Self-Organizing Maps</article-title>. <source>Expert Syst. Appl.</source> <volume>42</volume>, <fpage>8327</fpage>&#x2013;<lpage>8336</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2015.06.049</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Hart</surname>
<given-names>P. E.</given-names>
</name>
<name>
<surname>David</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Duda</surname>
<given-names>R. O.</given-names>
</name>
</person-group> (<year>2001</year>). <source>Pattern Classification</source>. <publisher-name>Wiley Hoboken</publisher-name>. </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hsu</surname>
<given-names>C.-M.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>A Hybrid Procedure for Stock price Prediction by Integrating Self-Organizing Map and Genetic Programming</article-title>. <source>Expert Syst. Appl.</source> <volume>38</volume>, <fpage>14026</fpage>&#x2013;<lpage>14036</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2011.04.210</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Isa</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Kallimani</surname>
<given-names>V. P.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>L. H.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Using the Self Organizing Map for Clustering of Text Documents</article-title>. <source>Expert Syst. Appl.</source> <volume>36</volume>, <fpage>9584</fpage>&#x2013;<lpage>9591</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2008.07.082</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jain</surname>
<given-names>D. K.</given-names>
</name>
<name>
<surname>Dubey</surname>
<given-names>S. B.</given-names>
</name>
<name>
<surname>Choubey</surname>
<given-names>R. K.</given-names>
</name>
<name>
<surname>Sinhal</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Arjaria</surname>
<given-names>S. K.</given-names>
</name>
<name>
<surname>Jain</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>An Approach for Hyperspectral Image Classification by Optimizing Svm Using Self Organizing Map</article-title>. <source>J.&#x20;Comput. Sci.</source> <volume>25</volume>, <fpage>252</fpage>&#x2013;<lpage>259</lpage>. <pub-id pub-id-type="doi">10.1016/j.jocs.2017.07.016</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jullum</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>L&#xf8;land</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Huseby</surname>
<given-names>R. B.</given-names>
</name>
<name>
<surname>&#xc5;nonsen</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Lorentzen</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Detecting Money Laundering Transactions with Machine Learning</article-title>. <source>J.&#x20;Money Laundering Control.</source> <volume>23</volume> (<issue>1</issue>), <fpage>173</fpage>&#x2013;<lpage>186</lpage>. <pub-id pub-id-type="doi">10.1108/jmlc-07-2019-0055</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Jun</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2006</year>). &#x201c;<article-title>A Peer Dataset Comparison Outlier Detection Model Applied to Financial Surveillance</article-title>,&#x201d; in <conf-name>18th International Conference on Pattern Recognition (ICPR&#x2019;06)</conf-name> (<publisher-name>IEEE</publisher-name>), <volume>4</volume>, <fpage>900</fpage>&#x2013;<lpage>903</lpage>. <pub-id pub-id-type="doi">10.1109/icpr.2006.150</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kiang</surname>
<given-names>M. Y.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>M. Y.</given-names>
</name>
<name>
<surname>Fisher</surname>
<given-names>D. M.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>An Extended Self-Organizing Map Network for Market Segmentation-A Telecommunication Example</article-title>. <source>Decis. Support Syst.</source> <volume>42</volume>, <fpage>36</fpage>&#x2013;<lpage>47</lpage>. <pub-id pub-id-type="doi">10.1016/j.dss.2004.09.012</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Kim</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2007</year>). &#x201c;<article-title>An Effective Under-sampling Method for Class Imbalance Data Problem</article-title>,&#x201d; in <conf-name>Proceedings of the 8th Symposium on Advanced Intelligent Systems</conf-name>, <fpage>825</fpage>&#x2013;<lpage>829</lpage>. </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kiviluoto</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>1998</year>). <article-title>Predicting Bankruptcies with the Self-Organizing Map</article-title>. <source>Neurocomputing</source> <volume>21</volume>, <fpage>191</fpage>&#x2013;<lpage>201</lpage>. <pub-id pub-id-type="doi">10.1016/s0925-2312(98)00038-1</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kleinberg</surname>
<given-names>J.&#x20;M.</given-names>
</name>
<name>
<surname>Newman</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Barab&#xE1;si</surname>
<given-names>A. L.</given-names>
</name>
<name>
<surname>Watts</surname>
<given-names>D. J.</given-names>
</name>
</person-group> (<year>2011</year>). <source>Authoritative Sources in a Hyperlinked Environment</source>. <publisher-name>Princeton University Press</publisher-name>.</citation>
</ref>
<ref id="B24">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kohonen</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2012</year>). <source>Self-organizing Maps</source>, <volume>30</volume>. <publisher-name>Springer Science &#x26; Business Media</publisher-name>. </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kohonen</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>1990</year>). <article-title>The Self-Organizing Map</article-title>. <source>Proc. IEEE</source> <volume>78</volume>, <fpage>1464</fpage>&#x2013;<lpage>1480</lpage>. <pub-id pub-id-type="doi">10.1109/5.58325</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kumar</surname>
<given-names>V. A.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Money Laundering: Concept, Significance and its Impact</article-title>. <source>Eur. J.&#x20;Business Manage.</source> <volume>4</volume>. </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lacerda</surname>
<given-names>E. B.</given-names>
</name>
<name>
<surname>Mello</surname>
<given-names>C. A. B.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Segmentation of Connected Handwritten Digits Using Self-Organizing Maps</article-title>. <source>Expert Syst. Appl.</source> <volume>40</volume>, <fpage>5867</fpage>&#x2013;<lpage>5877</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2013.05.006</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Le-Khac</surname>
<given-names>N.-A.</given-names>
</name>
<name>
<surname>Markos</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kechadi</surname>
<given-names>M.-T.</given-names>
</name>
</person-group> (<year>2009</year>). &#x201c;<article-title>Towards a New Data Mining-Based Approach for Anti-money Laundering in an International Investment Bank</article-title>,&#x201d; in <conf-name>International Conference on Digital Forensics and Cyber Crime</conf-name> (<publisher-name>Springer</publisher-name>), <fpage>77</fpage>&#x2013;<lpage>84</lpage>. </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Levi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Reuter</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Money Laundering</article-title>. <source>Crime and Justice</source> <volume>34</volume>, <fpage>289</fpage>&#x2013;<lpage>375</lpage>. <pub-id pub-id-type="doi">10.1086/501508</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2008</year>). &#x201c;<article-title>Sequence Matching for Suspicious Activity Detection in Anti-money Laundering</article-title>,&#x201d; in <conf-name>International Conference on Intelligence and Security Informatics</conf-name> (<publisher-name>Springer</publisher-name>), <fpage>50</fpage>&#x2013;<lpage>61</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-540-69304-8_6</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Lopez-Rojas</surname>
<given-names>E. A.</given-names>
</name>
<name>
<surname>Axelsson</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Multi Agent Based Simulation (Mabs) of Financial Transactions for Anti Money Laundering (Aml)</article-title>,&#x201d; in <conf-name>Nordic Conference on Secure IT Systems</conf-name> (<publisher-loc>Karlskrona, Sweden</publisher-loc>: <publisher-name>Blekinge Institute of Technology</publisher-name>). </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nilashi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ahmadi</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Sheikhtaheri</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Naemi</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Alotaibi</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Abdulsalam Alarood</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Remote Tracking of Parkinson&#x27;s Disease Progression Using Ensembles of Deep Belief Network and Self-Organizing Map</article-title>. <source>Expert Syst. Appl.</source> <volume>159</volume>, <fpage>113562</fpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2020.113562</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Pacella</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Grieco</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Blaco</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>20162016</year>). &#x201c;<article-title>On the Use of Self-Organizing Map for Text Clustering in Engineering Change Process Analysis: a Case Study</article-title>,&#x201d; in <source>Computational Intelligence and Neuroscience</source>. <pub-id pub-id-type="doi">10.1155/2016/5139574</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Paula</surname>
<given-names>E. L.</given-names>
</name>
<name>
<surname>Ladeira</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Carvalho</surname>
<given-names>R. N.</given-names>
</name>
<name>
<surname>Marzagao</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Deep Learning Anomaly Detection as Support Fraud Investigation in Brazilian Exports and Anti-money Laundering</article-title>,&#x201d; in <conf-name>2016 15th IEEE International Conference on Machine Learning and Applications (ICMLA)</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>954</fpage>&#x2013;<lpage>960</lpage>. <pub-id pub-id-type="doi">10.1109/icmla.2016.0172</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pedregosa</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Varoquaux</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Gramfort</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Michel</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Thirion</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Grisel</surname>
<given-names>O.</given-names>
</name>
<etal/>
</person-group> (<year>2011</year>). <article-title>Scikit-learn: Machine Learning in Python</article-title>. <source>J.&#x20;Machine Learn. Res.</source> <volume>12</volume>, <fpage>2825</fpage>&#x2013;<lpage>2830</lpage>. </citation>
</ref>
<ref id="B36">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Savage</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Chou</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2016</year>). <source>Detection of Money Laundering Groups Using Supervised Learning in Networks</source>. <comment>arXiv preprint arXiv:1608.00708</comment>. </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rocha-Salazar</surname>
<given-names>J.-d.-J</given-names>
</name>
<name>
<surname>Segovia-Vargas</surname>
<given-names>M.-J.</given-names>
</name>
<name>
<surname>Camacho-Mi&#x00F1;anos</surname>
<given-names>M.-d.-M</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Money Laundering and Terrorism Financing Detection Using Neural Networks and an Abnormality Indicator</article-title>. <source>Expert Syst. Appl.</source> <volume>169</volume>, <fpage>114470</fpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2020.114470</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shah-Hosseini</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Safabakhsh</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>Tasom: a New Time Adaptive Self-Organizing Map</article-title>. <source>IEEE Trans. Syst. Man. Cybern. B</source> <volume>33</volume>, <fpage>271</fpage>&#x2013;<lpage>282</lpage>. <pub-id pub-id-type="doi">10.1109/tsmcb.2003.810442</pub-id> </citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Singh</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Can Artificial Intelligence, Regtech and Charitytech Provide Effective Solutions for Anti-money Laundering and Counter-terror Financing Initiatives in Charitable Fundraising</article-title>. <source>J.&#x20;Money Laundering Control.</source> <volume>24</volume> (<issue>3</issue>), <fpage>464</fpage>&#x2013;<lpage>482</lpage>. <pub-id pub-id-type="doi">10.1108/jmlc-09-2020-0100</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Singh</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Best</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Anti-money Laundering: Using Data Visualization to Identify Suspicious Activity</article-title>. <source>Int. J.&#x20;Account. Inf. Syst.</source> <volume>34</volume>, <fpage>100418</fpage>. <pub-id pub-id-type="doi">10.1016/j.accinf.2019.06.001</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Soltani</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Nguyen</surname>
<given-names>U. T.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Faghani</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Yagoub</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>An</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>A New Algorithm for Money Laundering Detection Based on Structural Similarity</article-title>,&#x201d; in <conf-name>2016 IEEE 7th Annual Ubiquitous Computing, Electronics &#x26; Mobile Communication Conference (UEMCON)</conf-name> (<publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1109/uemcon.2016.7777919</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Stephanakis</surname>
<given-names>I. M.</given-names>
</name>
<name>
<surname>Chochliouros</surname>
<given-names>I. P.</given-names>
</name>
<name>
<surname>Sfakianakis</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Shirazi</surname>
<given-names>S. N.</given-names>
</name>
<name>
<surname>Hutchison</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Hybrid Self-Organizing Feature Map (Som) for Anomaly Detection in Cloud Infrastructures Using Granular Clustering Based upon Value-Difference Metrics</article-title>. <source>Inf. Sci.</source> <volume>494</volume>, <fpage>247</fpage>&#x2013;<lpage>277</lpage>. <pub-id pub-id-type="doi">10.1016/j.ins.2019.03.069</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sudjianto</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Nair</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Yuan</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Kern</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Cela-D&#xed;az</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Statistical Methods for Fighting Financial Crimes</article-title>. <source>Technometrics</source> <volume>52</volume>, <fpage>5</fpage>&#x2013;<lpage>19</lpage>. <pub-id pub-id-type="doi">10.1198/tech.2010.07032</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Yin</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2005</year>). &#x201c;<article-title>Developing an Intelligent Data Discriminating System of Anti-money Laundering Based on Svm</article-title>,&#x201d; in <conf-name>2005 International conference on machine learning and cybernetics</conf-name> (<publisher-name>IEEE</publisher-name>), <volume>6</volume>, <fpage>3453</fpage>&#x2013;<lpage>3457</lpage>. <pub-id pub-id-type="doi">10.1109/icmlc.2005.1527539</pub-id> </citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Teichmann</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Recent Trends in Money Laundering</article-title>. <source>Crime L. Soc Change</source> <volume>73</volume>, <fpage>237</fpage>&#x2013;<lpage>247</lpage>. <pub-id pub-id-type="doi">10.1007/s10611-019-09859-0</pub-id> </citation>
</ref>
<ref id="B46">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Tiwari</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Gepp</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Kumar</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>A Review of Money Laundering Literature: the State of Research in Key Areas</article-title>,&#x201d; in <source>Pacific Accounting Review</source>. <pub-id pub-id-type="doi">10.1108/par-06-2019-0065</pub-id> </citation>
</ref>
<ref id="B47">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Turki</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hamdan</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Ajmi</surname>
<given-names>J.&#x20;A.</given-names>
</name>
<name>
<surname>Razzaque</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Regulatory Technology (Regtech) and Money Laundering Prevention: Exploratory Study from bahrain</article-title>,&#x201d; in <conf-name>International Conference on Advanced Machine Learning Technologies and Applications</conf-name> (<publisher-name>Springer</publisher-name>), <fpage>349</fpage>&#x2013;<lpage>359</lpage>. <pub-id pub-id-type="doi">10.1007/978-981-15-3383-9_32</pub-id> </citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vesanto</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Alhoniemi</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>Clustering of the Self-Organizing Map</article-title>. <source>IEEE Trans. Neural Netw.</source> <volume>11</volume>, <fpage>586</fpage>&#x2013;<lpage>600</lpage>. <pub-id pub-id-type="doi">10.1109/72.846731</pub-id> </citation>
</ref>
<ref id="B49">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Vettigli</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Minisom: Minimalistic and Numpy-Based Implementation of the Self Organizing Map</article-title>. <comment>GitHub.[Online]</comment>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://github.com/JustGlowing/minisom/">https://github.com/JustGlowing/minisom/</ext-link>
</comment>[<comment>Dataset</comment>]. </citation>
</ref>
<ref id="B50">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>S.-N.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.-G.</given-names>
</name>
</person-group> (<year>2007</year>). &#x201c;<article-title>A Money Laundering Risk Evaluation Method Based on Decision Tree</article-title>,&#x201d; in <conf-name>2007 International Conference on Machine Learning and Cybernetics</conf-name> (<publisher-name>IEEE</publisher-name>), <volume>1</volume>, <fpage>283</fpage>&#x2013;<lpage>286</lpage>. <pub-id pub-id-type="doi">10.1109/icmlc.2007.4370155</pub-id> </citation>
</ref>
<ref id="B51">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Dong</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2009</year>). &#x201c;<article-title>Research on Money Laundering Detection Based on Improved Minimum Spanning Tree Clustering and its Application</article-title>,&#x201d; in <conf-name>2009 Second international symposium on knowledge acquisition and modeling</conf-name> (<publisher-name>IEEE</publisher-name>), <volume>2</volume>, <fpage>62</fpage>&#x2013;<lpage>64</lpage>. <pub-id pub-id-type="doi">10.1109/kam.2009.221</pub-id> </citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Welch</surname>
<given-names>B. L.</given-names>
</name>
</person-group> (<year>1947</year>). <article-title>The Generalization of `Student&#x27;s&#x27; Problem when Several Different Population Variances Are Involved</article-title>. <source>Biometrika</source> <volume>34</volume>, <fpage>28</fpage>&#x2013;<lpage>35</lpage>. <pub-id pub-id-type="doi">10.2307/2332510</pub-id> </citation>
</ref>
<ref id="B53">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yorek</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Ugulu</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Aydin</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Using Self-Organizing Neural Network Map Combined with ward&#x2019;s Clustering Algorithm for Visualization of Students&#x2019; Cognitive Structural Models about Aliveness Concept</article-title>. <source>Comput. Intelligence Neurosci.</source> <volume>2016</volume>, <fpage>1</fpage>&#x2013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1155/2016/2476256</pub-id> </citation>
</ref>
<ref id="B54">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2003</year>). &#x201c;<article-title>Feature Selection for High-Dimensional Data: A Fast Correlation-Based Filter Solution</article-title>,&#x201d; in <conf-name>Proceedings of the 20th international conference on machine learning</conf-name> (<publisher-loc>Washington, DC</publisher-loc>: <publisher-name>ICML-03</publisher-name>), <fpage>856</fpage>&#x2013;<lpage>863</lpage>. </citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zadeh Shirazi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Seyyed Mahdavi Chabok</surname>
<given-names>S. J.</given-names>
</name>
<name>
<surname>Mohammadi</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>A Novel and Reliable Computational Intelligence System for Breast Cancer Detection</article-title>. <source>Med. Biol. Eng. Comput.</source> <volume>56</volume>, <fpage>721</fpage>&#x2013;<lpage>732</lpage>. <pub-id pub-id-type="doi">10.1007/s11517-017-1721-z</pub-id> </citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Jin</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Analyzing and Detecting Money-Laundering Accounts in Online Social Networks</article-title>. <source>IEEE Netw.</source> <volume>32</volume>, <fpage>115</fpage>&#x2013;<lpage>121</lpage>. </citation>
</ref>
</ref-list>
</back>
</article>