<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurorobot.</journal-id>
<journal-title>Frontiers in Neurorobotics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurorobot.</abbrev-journal-title>
<issn pub-type="epub">1662-5218</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnbot.2021.669534</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Learning Then, Learning Now, and Every Second in Between: Lifelong Learning With a Simulated Humanoid Robot</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Logacjov</surname> <given-names>Aleksej</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1219752/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Kerzel</surname> <given-names>Matthias</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/680229/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Wermter</surname> <given-names>Stefan</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/21776/overview"/>
</contrib>
</contrib-group>
<aff><institution>Department of Informatics, Research Group Knowledge Technology, Universit&#x000E4;t Hamburg</institution>, <addr-line>Hamburg</addr-line>, <country>Germany</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: German I. Parisi, University of Hamburg, Germany</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Luca Genovesi, University of Milano-Bicocca, Italy; Joost Van De Weijer, Universitat Aut&#x000F2;noma de Barcelona, Spain</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Aleksej Logacjov <email>a.logacjov&#x00040;gmail.com</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>01</day>
<month>07</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>15</volume>
<elocation-id>669534</elocation-id>
<history>
<date date-type="received">
<day>19</day>
<month>02</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>18</day>
<month>05</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2021 Logacjov, Kerzel and Wermter.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Logacjov, Kerzel and Wermter</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract><p>Long-term human-robot interaction requires the continuous acquisition of knowledge. This ability is referred to as lifelong learning (LL). LL is a long-standing challenge in machine learning due to catastrophic forgetting, which states that continuously learning from novel experiences leads to a decrease in the performance of previously acquired knowledge. Two recently published LL approaches are the Growing Dual-Memory (GDM) and the Self-organizing Incremental Neural Network&#x0002B; (SOINN&#x0002B;). Both are growing neural networks that create new neurons in response to novel sensory experiences. The latter approach shows state-of-the-art clustering performance on sequentially available data with low memory requirements regarding the number of nodes. However, classification capabilities are not investigated. Two novel contributions are made in our research paper: (I) An extended SOINN&#x0002B; approach, called associative SOINN&#x0002B; (A-SOINN&#x0002B;), is proposed. It adopts two main properties of the GDM model to facilitate classification. (II) A new LL object recognition dataset (v-NICO-World-LL) is presented. It is recorded in a nearly photorealistic virtual environment, where a virtual humanoid robot manipulates 100 different objects belonging to 10 classes. Real-world and artificially created background images, grouped into four different complexity levels, are utilized. The A-SOINN&#x0002B; reaches similar state-of-the-art classification accuracy results as the best GDM architecture of this work and consists of 30 to 350 times fewer neurons, evaluated on two LL object recognition datasets, the novel v-NICO-World-LL and the well-known CORe50. Furthermore, we observe an approximately 268 times lower training time. These reduced numbers result in lower memory and computational requirements, indicating higher suitability for autonomous social robots with low computational resources to facilitate a more efficient LL during long-term human-robot interactions.</p></abstract>
<kwd-group>
<kwd>lifelong learning</kwd>
<kwd>self-organizing incremental neural network</kwd>
<kwd>growing dual-memory</kwd>
<kwd>lifelong learning dataset</kwd>
<kwd>simulated humanoid robot</kwd>
<kwd>long-term human-robot interaction</kwd>
</kwd-group>
<contract-sponsor id="cn001">Universit&#x000E4;t Hamburg<named-content content-type="fundref-id">10.13039/501100005711</named-content></contract-sponsor>
<counts>
<fig-count count="9"/>
<table-count count="6"/>
<equation-count count="14"/>
<ref-count count="48"/>
<page-count count="20"/>
<word-count count="13174"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Social robots that interact with humans in their everyday lives are exposed to a dynamic and challenging environment. This dynamic environment provides continuous data streams (Parisi et al., <xref ref-type="bibr" rid="B32">2019</xref>) that are potentially infinite and non-stationary (Ghesmoune et al., <xref ref-type="bibr" rid="B9">2016</xref>; Wiwatcharakoses and Berrar, <xref ref-type="bibr" rid="B45">2019</xref>). Humans can continuously learn throughout their lifespan (Parisi et al., <xref ref-type="bibr" rid="B32">2019</xref>). Hence, to realize an intelligent behavior for long-term human-robot interaction (long-term HRI) in such a complex environment, robots are required to continually acquire new knowledge and adapt to unpredictable changes over time (Dautenhahn, <xref ref-type="bibr" rid="B6">2004</xref>; Parisi et al., <xref ref-type="bibr" rid="B31">2018</xref>; Lomonaco et al., <xref ref-type="bibr" rid="B23">2019</xref>), e.g., by remembering aspects of past interactions with humans (Leite et al., <xref ref-type="bibr" rid="B17">2013</xref>). Current deep learning (DL) approaches, especially convolutional neural networks (CNNs), succeed in many machine learning applications (Najafabadi et al., <xref ref-type="bibr" rid="B29">2015</xref>; Guo et al., <xref ref-type="bibr" rid="B10">2016</xref>). Nevertheless, CNNs rely on a large dataset of (partially) labeled training samples, assuming that all samples are available during training (LeCun et al., <xref ref-type="bibr" rid="B16">2015</xref>; Guo et al., <xref ref-type="bibr" rid="B10">2016</xref>; Parisi et al., <xref ref-type="bibr" rid="B32">2019</xref>). Information that was never seen before is frequently observed in a dynamic environment, requiring a conventional DL approach to be retrained on new and previous observations (Parisi et al., <xref ref-type="bibr" rid="B31">2018</xref>, <xref ref-type="bibr" rid="B32">2019</xref>). This can lead to infeasible memory requirements since all previously observed training samples need to be explicitly stored. The ability to acquire new knowledge over time while preserving previously learned tasks without retraining an architecture from scratch is referred to as <italic>lifelong learning</italic> (LL) (Parisi et al., <xref ref-type="bibr" rid="B32">2019</xref>). LL is a long-standing challenge in machine learning due to <italic>catastrophic forgetting</italic>, which states that continually learning from non-stationary data distributions generally leads to a decrease in the performance of previously learned tasks (Parisi et al., <xref ref-type="bibr" rid="B32">2019</xref>).</p>
<p>Two recently published LL approaches are the <italic>Growing Dual-Memory</italic> (GDM), proposed by Parisi et al. (<xref ref-type="bibr" rid="B31">2018</xref>), and the <italic>Self-organizing Incremental Neural Network&#x0002B;</italic> (SOINN&#x0002B;), proposed by Wiwatcharakoses and Berrar (<xref ref-type="bibr" rid="B45">2019</xref>). Both are growing neural networks that create new neurons in response to novel sensory experiences. The former model shows state-of-the-art classification and the latter state-of-the-art clustering results on sequentially available data. One of the main differences between these approaches is the utilization of different forgetting strategies. The GDM uses a predefined threshold that defines the maximal age that edges in the network can have before being pruned. If this threshold is too low, catastrophic forgetting can occur since critical units of previously learned tasks can be deleted due to their age (Liew et al., <xref ref-type="bibr" rid="B21">2019</xref>). A too high maximal age can lead to an infeasible amount of neurons. Therefore, an appropriate maximal age must be determined [e.g., by cross-validation (Wiwatcharakoses and Berrar, <xref ref-type="bibr" rid="B45">2019</xref>)], which may not be possible in the real world, with no fixed training data size (Liew et al., <xref ref-type="bibr" rid="B21">2019</xref>). Even though Parisi et al. (<xref ref-type="bibr" rid="B31">2018</xref>) state that the periodic replay of neural activation trajectories could be used to prevent the deletion of previously learned knowledge, there was no investigation made concerning the interplay of the maximum age and replay in terms of forgetting. On the other hand, the SOINN&#x0002B; does not require a predefined maximum age threshold. It calculates this value during the learning process itself. Wiwatcharakoses and Berrar (<xref ref-type="bibr" rid="B45">2019</xref>) showed that this forgetting strategy leads to a higher clustering performance and a lower amount of neurons than other growing neural networks. Nevertheless, the authors neither investigated the classification capabilities nor examined the model&#x00027;s behavior on higher than 10-dimensional data.</p>
<p>The promising results of Wiwatcharakoses and Berrar (<xref ref-type="bibr" rid="B45">2019</xref>) are the motivation for the development of an extended version of the SOINN&#x0002B; approach for classification tasks. We make two main contributions. First, we propose two extensions to the SOINN&#x0002B;, adopted from the GDM architecture. (1) An associative matrix is utilized, which stores how often a neuron observed a particular input label, enabling classification. (2) Two additional constraints are introduced that regulate node creation and node adaptation. A new neuron is only created if the input label is unequal to the network&#x00027;s predicted label. On the other hand, nodes are only updated if the input label is equal to the network prediction. We call this extended version associative SOINN&#x0002B; (A-SOINN&#x0002B;). Second, we present a novel LL object recognition dataset, called v-NICO-World-LL. It exhibits three novelties, to the best of our knowledge, not yet realized in combination by other LL object recognition datasets. (1) The background images are grouped into four different complexity groups, enabling the evaluation of LL models in environments of different levels of complexity. (2) A virtual robot is manipulating objects instead of a human to simulate a long-term HRI scenario where the robot receives different objects over time. (3) It is highly controlled and reproducible, mainly because it is recorded in a nearly photorealistic virtual environment. The dataset consists of 100 different objects, split into 10 categories. Twenty real-world and artificial images are used for the background to simulate different environments. We compare the A-SOINN&#x0002B; to a conceptually similar state-of-the-art LL approach, the GDM on two LL datasets, the v-NICO-World-LL and the CORe50 (Lomonaco and Maltoni, <xref ref-type="bibr" rid="B22">2017</xref>). Furthermore, we investigate whether the A-SOINN&#x0002B; can reach state-of-the-art classification accuracy results compared to the GDM while showing fewer memory requirements in terms of created neurons which would be desirable for robots with low computational resources. The A-SOINN&#x0002B; shows similar high accuracy results as the best GDM model. Simultaneously, a lower amount of units is created, resulting in fewer memory and computational requirements. A further ablation study shows that the new node creation constraint extension leads to the generation of fewer neurons.</p>
<p>Our paper is organized as follows. The following section (Section 2) gives a brief overview of the related work and focuses on the GDM, the SOINN&#x0002B;, and different LL benchmark datasets. The A-SOINN&#x0002B; approach and the new LL dataset are presented in Section 3. Section 4 shows the experimental setup and the results. The discussion of these results is presented in Section 5. A conclusion and future work are given in Section 6.</p>
</sec>
<sec id="s2">
<title>2. Related Work</title>
<sec>
<title>2.1. Lifelong Learning Approaches</title>
<p>Different approaches try to mitigate catastrophic forgetting in different ways. According to Parisi et al. (<xref ref-type="bibr" rid="B32">2019</xref>) and Lesort et al. (<xref ref-type="bibr" rid="B18">2020</xref>), those approaches can conceptually be grouped into different categories. This section describes the four categories presented in the work of Lesort et al. (<xref ref-type="bibr" rid="B18">2020</xref>), together with example approaches.</p>
<sec>
<title>Dynamic Architecture</title>
<p>Approaches belonging to this category modify the model&#x00027;s architecture dynamically, either explicitly or implicitly. Approaches with explicit architectural modifications create either new models as soon as a new task occurs and connect those models [e.g., Progressive Neural Networks (PNN) (Rusu et al., <xref ref-type="bibr" rid="B39">2016</xref>)] or create new neurons inside the model for each new task, like the <italic>Self-organizing Incremental Neural Network&#x0002B;</italic> (SOINN&#x0002B;) (Wiwatcharakoses and Berrar, <xref ref-type="bibr" rid="B45">2019</xref>) (Subsection 2.3). Approaches that make implicit modifications are not directly changing the architecture. Instead, either a part of the model is deactivated (e.g., freezing weights during backpropagation), or the forward pass paths are modified while learning a new task. Freezing weight approaches are, for example, the hard attention to the task (HAT) (Serr&#x000E0; et al., <xref ref-type="bibr" rid="B40">2018</xref>), the PackNet (Mallya and Lazebnik, <xref ref-type="bibr" rid="B26">2017</xref>), or the Piggyback (PB) (Mallya et al., <xref ref-type="bibr" rid="B25">2018</xref>) approach. The adaptation of the forward pass path is realized in PathNet (Fernando et al., <xref ref-type="bibr" rid="B8">2017</xref>).</p>
</sec>
<sec>
<title>Regularization</title>
<p>Lesort et al. (<xref ref-type="bibr" rid="B18">2020</xref>) present two types of regularization approaches, penalty computing and knowledge distillation. Penalty computing approaches regulate how strong weights are updated while learning a task. Approaches like elastic weight consolidation (EWC) (Kirkpatrick et al., <xref ref-type="bibr" rid="B15">2016</xref>), Synaptic Intelligence (SI) (Zenke et al., <xref ref-type="bibr" rid="B48">2017</xref>), and Memory Aware Synapses (MAS) (Aljundi et al., <xref ref-type="bibr" rid="B2">2018</xref>) search for important weights inside the model and penalize substantial changes to them (Parisi et al., <xref ref-type="bibr" rid="B32">2019</xref>). Hence, the weights are protected from being modified if they are crucial for previously learned tasks (Lesort et al., <xref ref-type="bibr" rid="B18">2020</xref>). Knowledge distillation in a LL context is realized by training a model <italic>A</italic> on task <italic>t</italic><sub><italic>i</italic></sub>. After <italic>A</italic> learned to solve the task, a model <italic>B</italic> is trained to solve a new task <italic>t</italic><sub><italic>i</italic>&#x0002B;1</sub> and to generate the same output as <italic>A</italic>. Hence, knowledge is distilled from <italic>A</italic> to <italic>B</italic>. In the end, <italic>B</italic> is required to solve both tasks. The learning without forgetting (LWF) approach of Li and Hoiem (<xref ref-type="bibr" rid="B20">2018</xref>) is an example of a knowledge distillation technique (Parisi et al., <xref ref-type="bibr" rid="B32">2019</xref>).</p>
</sec>
<sec>
<title>Rehearsal</title>
<p>Rehearsal approaches save raw data samples of previous tasks and incorporate them into the new task&#x00027;s training set. These samples can either be randomly or carefully chosen to save representatives of past tasks. The Incremental Classifier and Representation Learning (iCaRL) (Rebuffi et al., <xref ref-type="bibr" rid="B37">2016</xref>) is an example rehearsal approach that keeps the most representative samples of previous tasks for future learning. This strategy allows weight strengthening for already learned memories.</p>
</sec>
<sec>
<title>Generative Replay</title>
<p>Compared to rehearsal approaches, generative replay (or pseudo-rehearsal) algorithms learn to artificially generate data samples for past tasks instead of saving raw data of previously seen tasks. Generative models learn the distribution of data previously encountered to replay past experiences when learning on new data. These models are often generative adversarial networks or auto-encoders (Lesort et al., <xref ref-type="bibr" rid="B18">2020</xref>). An example is the generative replay approach (GR) proposed by Shin et al. (<xref ref-type="bibr" rid="B41">2017</xref>).</p>
</sec>
<sec>
<title>Hybrid</title>
<p>According to Lesort et al. (<xref ref-type="bibr" rid="B18">2020</xref>), most LL approaches rely on more than one of the four mentioned strategies, often leading to better solutions. One example is the previously mentioned iCaRL approach of Rebuffi et al. (<xref ref-type="bibr" rid="B37">2016</xref>). Additionally to the usage of raw data of past tasks, iCaRL uses knowledge distillation. Instead of transferring information between different neural networks, this knowledge is transferred within a single model between different time steps, similarly to the LWF (Li and Hoiem, <xref ref-type="bibr" rid="B20">2018</xref>) approach. The Learning a Unified Classifier Incrementally via Rebalancing (LUCIR) (Hou et al., <xref ref-type="bibr" rid="B12">2019</xref>) is similar to iCaRL. However, they introduce three new components to mitigate catastrophic forgetting caused by the imbalance between new and old data. This imbalance is also tackled by the Bias Correction (BiC) approach (Wu et al., <xref ref-type="bibr" rid="B46">2019</xref>). The authors show evidence that the last fully connected layer of a CNN has a strong bias toward new classes, which they corrected by utilizing a bias correction layer after the last network layer. A further rehearsal and distillation regularization approach is the Pooled Outputs Distillation for Small-Task Incremental Learning (PODNet) (Douillard et al., <xref ref-type="bibr" rid="B7">2020</xref>). The authors introduce a novel distillation-loss to ensure a balance between reducing forgetting and learning new tasks for long-term incremental learning, as well as a multi-mode similarity classifier that is more robust to data distribution shifts. The Dynamic Expandable Network (DEN) of Yoon et al. (<xref ref-type="bibr" rid="B47">2018</xref>) is a regularization and an architectural approach. It identifies neurons in a deep neural network relevant for a new task and selectively retrains them. However, if this selective retraining fails to achieve a desired loss, the network is expanded with additional neurons while unnecessary units are eliminated. Catastrophic forgetting is mitigated by duplicating neurons that drifted too much from their original values. Hou et al. (<xref ref-type="bibr" rid="B11">2018</xref>) proposed the Adaptation by Distillation approach, which combines knowledge distillation from an intermediate <italic>Expert CNN</italic> to learn new tasks with caching of small data subsets of previous tasks to preserve old knowledge during training a CNN. Therefore, it is a regularization and a rehearsal approach. The Dynamic Generative Memory (DGM), a combination of generative replay and dynamic architecture model, is presented by Ostapenko et al. (<xref ref-type="bibr" rid="B30">2019</xref>). A generative model is trained incrementally to learn task distributions over time. Samples of the current task and synthesized samples of all previous tasks are used to train a task solver. Additionally, the generative model can be expanded to ensure constant expressive power and sufficient capacity. The Riemannian Walk (RWalk) approach of Chaudhry et al. (<xref ref-type="bibr" rid="B5">2018</xref>) is a rehearsal and penalty computing regularization technique based on Kullback&#x02013;Leibler-divergence (between output distributions) for parameter importance score computation. Representative samples of previous tasks are stored for replay to improve performance. A combination of regularization and implicit dynamic architecture approach, called Continual Learning with Adaptive Weights (CLAW), is presented by Adel et al. (<xref ref-type="bibr" rid="B1">2020</xref>). The architecture adaptation is data-driven by learning which neurons need to be trained and what is the maximum adaptation applied to these neurons using variational inference. The Incremental Learning With Dual Memory (IL2M) (Belouadah and Popescu, <xref ref-type="bibr" rid="B3">2019</xref>) consists of two parts: (1) a deep learning model that is incrementally trained on new samples and a constant number of previous representative samples, and (2) an additional memory that stores previous task statistics which are periodically used to rectify the network. The Growing Dual Memory (GDM) approach of Parisi et al. (<xref ref-type="bibr" rid="B31">2018</xref>) is an architectural and a generative replay approach. In Subsection 2.2, it is described in more detail.</p>
</sec>
</sec>
<sec>
<title>2.2. Growing Dual-Memory</title>
<p>The Growing Dual-Memory consists of two interconnected recurrent Growing When Required (GWR) (Marsland et al., <xref ref-type="bibr" rid="B27">2002</xref>) models called <italic>growing episodic memory</italic> (G-EM) and <italic>growing semantic memory</italic> (G-SM), where both are extended versions of the <italic>Gamma-GWR</italic> model proposed by Parisi et al. (<xref ref-type="bibr" rid="B33">2017</xref>). Depending on the similarity between the input and the nearest neighbor in the network, i.e., the best-matching unit (BMU), new neurons are created or existing neurons trained. This similarity (or activation) is computed using the exponential function of the negative Euclidean distance. If the similarity is too low, compared to a fixed predefined activation threshold, a new neuron is created. Otherwise, the BMU and its neighbors are trained toward the input. Additional context vectors are used for each neuron. These context vectors are integrated into the similarity computation, allowing the utilization of past network activations to learn the input&#x00027;s temporal structure (Parisi et al., <xref ref-type="bibr" rid="B31">2018</xref>). In both G-EM and G-SM, edges that exceed a predefined maximum age threshold are considered to be &#x0201C;too old&#x0201D; and are, therefore, removed. Isolated nodes are deleted. In the G-EM layer, temporal connections between neurons are introduced where consecutively activated units show a stronger temporal link. Given a neuron <italic>s</italic>(<italic>i</italic> &#x02212; 1), the temporal connections allow the determination of the most probable next neuron <italic>s</italic>(<italic>i</italic>) of a prototype trajectory. These trajectories are replayed to G-EM and G-SM in the absence of external input. The weights of the BMU inside the G-EM and its temporal information are used to train the G-SM layer. In G-SM, Parisi et al. (<xref ref-type="bibr" rid="B31">2018</xref>) equipped the Gamma-GWR network with two additional constraints. It will only create new nodes if it cannot predict the current input correctly, and it will only update nodes if the input label can be adequately predicted. Both constraints are realized by maintaining an associative matrix <italic>H</italic>(<italic>j, l</italic>) for label predictions. It stores a histogram of input labels for each node <italic>j</italic>. The prediction of the GDM network is the label the BMU most frequently observed in the past. Additionally, Parisi et al. (<xref ref-type="bibr" rid="B31">2018</xref>) used a pre-trained and fine-tuned VGG-16 CNN (Simonyan and Zisserman, <xref ref-type="bibr" rid="B42">2014</xref>) as a feature extractor to create 256-dimensional feature vectors for each input image. Transfer learning was performed using the CORe50 dataset (Subsection 2.4). These feature vectors are then used as the input for the G-EM layer.</p>
</sec>
<sec>
<title>2.3. Self-Organizing Incremental Neural Network&#x0002B;</title>
<p>The Self-organizing Incremental Neural Network&#x0002B; approach of Wiwatcharakoses and Berrar (<xref ref-type="bibr" rid="B45">2019</xref>) has similarities to the Gamma-GWR architecture. Nevertheless, it differs fundamentally in how the deletion and creation of nodes or edges take place. The SOINN&#x0002B; approach considers the deletion of edges and nodes as an intrinsic part of the learning itself. It deletes edges and nodes only if they are not relevant enough for the learning task (Wiwatcharakoses and Berrar, <xref ref-type="bibr" rid="B45">2019</xref>). The key aspects of the SOINN&#x0002B; approach are the <italic>lifetime</italic> of an edge as well as the <italic>trustworthiness, idle time</italic>, and <italic>(un-)utility</italic> of a node. <italic>Trustworthiness</italic> is used to determine whether to link two nodes with an edge or not. Particularly, the more trust there is in these nodes, the more likely they are connected. The more often a node was chosen as the BMU, the higher is the trust in it. This technique tries to avoid connecting nodes that represent noise. The <italic>lifetime/age</italic> of already deleted edges and the lifetime of all edges reachable from the BMU are used to calculate an edge pruning threshold. If the lifetime of an edge exceeds this threshold, it is removed. This technique enables the deletion of edges that connect nodes of different clusters to isolate these. The <italic>unutility</italic> is the ratio between the idle time and the winning time of a node. The former represents the number of iterations passed since the node was a BMU, and the latter the number of times it was chosen as the BMU. A node pruning threshold is calculated using the unutilities of deleted and existing nodes. If the unutility of a node exceeds this threshold, it is a candidate for deletion.</p>
</sec>
<sec>
<title>2.4. Benchmark Datasets</title>
<p>Most LL benchmark datasets are adopted from other fields not explicitly created for LL (Lesort et al., <xref ref-type="bibr" rid="B18">2020</xref>). The samples of these datasets are split, artificially modified (e.g., rotation or permutation), or concatenated to create a sequence of tasks (Lesort et al., <xref ref-type="bibr" rid="B18">2020</xref>). Examples are the permuted MNIST used by Kirkpatrick et al. (<xref ref-type="bibr" rid="B15">2016</xref>), the rotated MNIST (Lopez-Paz and Ranzato, <xref ref-type="bibr" rid="B24">2017</xref>), or the incremental CIFAR100 (Rebuffi et al., <xref ref-type="bibr" rid="B37">2016</xref>). Real-world datasets suitable for LL are the iCubWorld28 (Pasquale et al., <xref ref-type="bibr" rid="B34">2015</xref>) or the iCubWorld Transformations (Pasquale et al., <xref ref-type="bibr" rid="B36">2016</xref>; Pasquale et al., <xref ref-type="bibr" rid="B35">2017</xref>). In these datasets, a human operator manipulates different objects in front of the iCub (Metta et al., <xref ref-type="bibr" rid="B28">2008</xref>) robot&#x00027;s cameras. A tracking routine is used to move the robot&#x00027;s gaze toward the object and extract a bounding box around it. Lomonaco and Maltoni (<xref ref-type="bibr" rid="B22">2017</xref>) proposed the CORe50, a dataset specially designed for continuous object recognition. Fifty different objects belonging to 10 object categories are recorded. Each category consists of five instances. In contrast to the iCubWorld datasets, each object is recorded in front of 11 different real-world environments, both outdoor and indoor. In each video (15 s, 20 fps), a human hand moves and rotates the object smoothly in front of a camera. A motion-based tracker is used to create images with a resolution of 128 &#x000D7; 128 pixels by cutting out a bounding box around the object. However, both dataset acquisition processes are not controlled enough for perfect reproducibility. Besides, the ability to change individual aspects in the scenario (e.g., the background) without changing others in any way is not given. Nevertheless, this is desirable to analyze the behavior of LL models on particular environmental changes. The Toys-200 dataset of Stojanov et al. (<xref ref-type="bibr" rid="B43">2019</xref>) is recorded in a 3D virtual environment and consists of 200 toy-like objects. These are translated and scaled in front of a moving camera. The background contains further objects of the dataset distributed over a floor. In contrast to the CORe50 and the iCubWorld datasets, it is highly reproducible. However, it contains no operator (human or robot) who manipulates the objects, which would be the case in a long-term HRI. Such manipulation can lead to occlusion, making the task more difficult. Furthermore, except for the floor scenario, no real-world background images are considered.</p>
</sec>
</sec>
<sec sec-type="methods" id="s3">
<title>3. Methodology</title>
<sec>
<title>3.1. Associative SOINN&#x0002B;</title>
<p>The associative SOINN&#x0002B; (A-SOINN&#x0002B;) extends the SOINN&#x0002B; of Wiwatcharakoses and Berrar (<xref ref-type="bibr" rid="B45">2019</xref>) in two points. Both are adopted from the GDM approach. (1) An associative matrix <italic>H</italic>(<italic>j, l</italic>) is maintained, which stores how often a neuron observed a particular input label. This point gives this extension its name and enables classification, which was previously not possible in the SOINN&#x0002B;. (2) Top-down cues to regulate the network&#x00027;s structural plasticity are introduced in the form of additional constraints for node creation and weight adaptation. Additionally, inspired by Parisi et al. (<xref ref-type="bibr" rid="B31">2018</xref>), a pre-trained and fine-tuned VGG-16 (Simonyan and Zisserman, <xref ref-type="bibr" rid="B42">2014</xref>) CNN architecture is used as a feature extractor to extract the most relevant features out of images and reduce the dimensionality. The A-SOINN&#x0002B; is an architectural LL approach as it adapts its shape in response to novel input. In contrast to GDM, it does not perform memory replay and is, therefore, no hybrid approach. The architecture of the A-SOINN&#x0002B; and the feature extractor are illustrated in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>This figure shows an illustration of the A-SOINN&#x0002B; <bold>(A)</bold> and the feature extractor <bold>(B)</bold>. Over time, different frames become available. These are fed into the feature extractor, illustrated by the thick, black arrow. The feature extractor creates feature vectors representing the image frames. These feature vectors are fed into the A-SOINN&#x0002B; model afterward, illustrated by the thin, yellow arrow. The proposed model stores a histogram of input labels for each node to facilitate classification, shown for one node in the top right corner. This figure is inspired by the GDM illustration of Parisi et al. (<xref ref-type="bibr" rid="B31">2018</xref>). The 3D object is a derivative of <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/3d-models/antique-pocket-watch-fa99fca27aa54adaabb03473616ff117">&#x0201C;Antique Pocket Watch&#x0201D;</ext-link> by <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/michael_grodkowski">&#x0201C;michael_grodkowski&#x0201D;</ext-link> licensed under <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</ext-link>.</p></caption>
<graphic xlink:href="fnbot-15-669534-g0001.tif"/>
</fig>
<p>Furthermore, the (associative) SOINN&#x0002B; algorithm<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref> is shown in algorithms 1&#x02013;6. The adaptations made in this work have the prefix: <bold>(A-SOINN&#x0002B;)</bold>.</p>
<sec>
<title>Main Routine</title>
<p>The network is initialized with three random, unconnected neurons with a winning time of one and an idle time of zero (Algorithm 1, steps 1&#x02013;4). The associative matrix is initially empty (Algorithm 1, Step 5). As soon as a new training sample <italic>x</italic>(<italic>t</italic>) becomes available, the BMU <italic>b</italic> and sBMU <italic>s</italic> are computed using the Euclidean distance (Algorithm 1, step 7). Furthermore, the associative matrix entries of the BMU <italic>b</italic> are updated, such that <italic>H</italic>(<italic>b, l</italic>(<italic>t</italic>)) is increased by a predefined &#x003B4;<sup>&#x0002B;</sup> and <italic>H</italic>(<italic>b, k</italic>) decreased by &#x003B4;<sup>&#x02212;</sup> for all <italic>k</italic> &#x02260; <italic>l</italic>(<italic>t</italic>) (Algorithm 1, steps 8 and 9) (Parisi et al., <xref ref-type="bibr" rid="B33">2017</xref>). The similarity thresholds of <italic>b</italic> and <italic>s</italic> are calculated afterward (Algorithm 2) using:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>&#x003C4;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mstyle displaystyle="true"><mml:msub><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mstyle><mml:mo>&#x02016;</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02016;</mml:mo><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mtext class="textrm" mathvariant="normal">if&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02260;</mml:mo><mml:mi>&#x02205;</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle displaystyle="true"><mml:msub><mml:mrow><mml:mo class="qopname">min</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>A</mml:mi><mml:mo>\</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:mstyle><mml:mo>&#x02016;</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02016;</mml:mo><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mtext class="textrm" mathvariant="normal">otherwise</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow><mml:mtext class="textrm" mathvariant="normal">,</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>for <italic>i</italic> &#x02208; [<italic>b, s</italic>]. The weight vector is denoted by <italic>w</italic><sub><italic>i</italic></sub> and the neighbors of <italic>i</italic> by <italic>N</italic><sub><italic>i</italic></sub>. Hence, the similarity threshold of a node <italic>i</italic> is the maximum distance from <italic>i</italic> to all of its neighbors. If <italic>i</italic> has no neighbors, &#x003C4;(<italic>i</italic>) is the distance between <italic>i</italic> and its closest node. The prediction (Algorithm 1, Step 11) of the network is the label the BMU observed most frequently, computed with:</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x003BE;</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">arg&#x000A0;max</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>L</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mi>H</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>l</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Suppose the Euclidean distance between the input <italic>x</italic>(<italic>t</italic>) and <italic>b</italic> or between <italic>x</italic>(<italic>t</italic>) and <italic>s</italic> is larger than the corresponding similarity threshold, and simultaneously the prediction &#x003BE;<sub><italic>b</italic></sub> of <italic>b</italic> is unequal to the input label <italic>l</italic> (i.e., creation condition is true). In that case, a new node is created (Algorithm 1, Step 15). The input <italic>x</italic>(<italic>t</italic>) is used as the weight vector of the new node. Winning time and idle time are initiated with one and zero, respectively. If, on the other hand, no neuron is created, three subroutines are executed, namely node merging, node linking, and edge deletion (Algorithm 1, steps 16&#x02013;19), but only if the input label is unequal to the prediction (i.e., adaptation condition is true). The node deletion subroutine is performed at the end (Algorithm 1, Step 21).</p>
<table-wrap position="float">
<label>Algorithm 1</label>
<caption><p>(Associative) Self-Organizing Incremental Neural Network &#x0002B;</p></caption>
<table frame="hsides" rules="groups">
<tbody>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;1: <italic>A</italic> &#x02190; Set of neurons with random weight vectors.</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;2: <italic>WT</italic>(<italic>i</italic>) &#x02190; 1, &#x02200;<italic>i</italic> &#x02208; [1, 2, 3] &#x000A0;&#x000A0;&#x000A0;// Set the winning time to one.</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;3: <italic>IT</italic>(<italic>i</italic>) &#x02190; 0, &#x02200;<italic>i</italic> &#x02208; [1, 2, 3] &#x000A0;&#x000A0;&#x000A0;// Set the idle time to zero.</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;4: <italic>E</italic> &#x02190; &#x02205; &#x000A0;&#x000A0;&#x000A0;// Initialize an empty set of connections.</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;5: <bold>(A-SOINN&#x0002B;)</bold>: <italic>H</italic>(<italic>i</italic>, &#x000B7;) &#x02190; &#x02205;, &#x02200;<italic>i</italic> &#x02208; [1, 2, 3] &#x000A0;&#x000A0;&#x000A0;// Initialize an associative matrix with no labels</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;6: <bold>for</bold> each input <italic>x</italic>(<italic>t</italic>) with label <italic>l</italic>(<italic>t</italic>) <bold>do</bold></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;7: &#x000A0;&#x000A0;&#x000A0;<inline-formula><mml:math id="M3"><mml:mi>b</mml:mi><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">arg min</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>A</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mo>&#x02016;</mml:mo><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, &#x000A0;&#x000A0;&#x000A0;// Find the BMU <inline-formula><mml:math id="M4"><mml:mi>s</mml:mi><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mrow><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">arg min</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>A</mml:mi><mml:mo>\</mml:mo><mml:mrow><mml:mo stretchy="false">{</mml:mo><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mo stretchy="false">}</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mo>&#x02016;</mml:mo><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> &#x000A0;&#x000A0;&#x000A0;// Find the (s)BMU.</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;8: &#x000A0;&#x000A0;&#x000A0;<bold>(A-SOINN&#x0002B;)</bold>: &#x000A0;&#x000A0;&#x000A0;&#x00394;<italic>H</italic>(<italic>b, l</italic>(<italic>t</italic>)) &#x0003D; &#x003B4;<sup>&#x0002B;</sup>, &#x00394;<italic>H</italic>(<italic>b, k</italic>) &#x0003D; &#x02212; &#x003B4;<sup>&#x02212;</sup>, &#x02200;<italic>k</italic> &#x02208; <italic>L</italic>\{<italic>l</italic>(<italic>t</italic>)} &#x000A0;&#x000A0;&#x000A0;// Update associative matrix</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;9: &#x000A0;&#x000A0;&#x000A0;&#x000A0;<italic>N</italic><sub><italic>b</italic></sub>, <italic>N</italic><sub><italic>s</italic></sub> &#x02190; neighbors of <italic>b</italic>, <italic>s</italic></td></tr>
<tr><td align="left" valign="top">10: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x003C4;(<italic>b</italic>), &#x003C4;(<italic>s</italic>) &#x02190; Calculation of similarity thresholds (Algorithm 2)</td></tr>
<tr><td align="left" valign="top">11: &#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>(A-SOINN&#x0002B;)</bold>: &#x003BE;<sub><italic>b</italic></sub> &#x02190; arg max<sub><italic>l</italic> &#x02208; <italic>L</italic></sub><italic>H</italic>(<italic>b, l</italic>) &#x000A0;&#x000A0;&#x000A0;// Get prediction (Equation 2).</td></tr>
<tr><td align="left" valign="top">12: &#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>(A-SOINN&#x0002B;)</bold>: creation_condition &#x02190; &#x003BE;<sub><italic>b</italic></sub> &#x02260; <italic>l</italic>(<italic>t</italic>)</td></tr>
<tr><td align="left" valign="top">13: &#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>(A-SOINN&#x0002B;)</bold>: adaptation_condition &#x02190; &#x003BE;<sub><italic>b</italic></sub> &#x0003D; <italic>l</italic>(<italic>t</italic>)</td></tr>
<tr><td align="left" valign="top">14: &#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>if</bold> (||<italic>x</italic>(<italic>t</italic>) &#x02212; <italic>w</italic><sub><italic>b</italic></sub>|| &#x02265; &#x003C4;(<italic>b</italic>) <bold>or</bold> ||<italic>x</italic>(<italic>t</italic>) &#x02212; <italic>w</italic><sub><italic>s</italic></sub>|| &#x02265; &#x003C4;(<italic>s</italic>)) <bold>and</bold> creation_condition <bold>then</bold></td></tr>
<tr><td align="left" valign="top">15: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<italic>w</italic><sub><italic>r</italic></sub> &#x02190; <italic>x</italic>(<italic>t</italic>), <italic>A</italic> &#x02190; <italic>A</italic> &#x0222A; {<italic>r</italic>}, <italic>WT</italic>(<italic>r</italic>) &#x02190; 1, <italic>IT</italic>(<italic>r</italic>) &#x02190; 0 &#x000A0;&#x000A0;&#x000A0;// Create a new node. </td></tr>
<tr><td align="left" valign="top">16: &#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>else if</bold> adaptation_condition <bold>then</bold></td></tr>
<tr><td align="left" valign="top">17: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;Node merging (Algorithm 3)</td></tr>
<tr><td align="left" valign="top">18: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;Node linking (Algorithm 4)</td></tr>
<tr><td align="left" valign="top">19: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;Edge deletion (Algorithm 5) </td></tr>
<tr><td align="left" valign="top">20: &#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>end if</bold></td></tr>
<tr><td align="left" valign="top">21: &#x000A0;&#x000A0;&#x000A0;&#x000A0;Node deletion (Algorithm 6) </td></tr>
<tr><td align="left" valign="top">22: <bold>end for</bold></td></tr> 
</tbody>
</table>
</table-wrap>
<table-wrap position="float">
<label>Algorithm 2</label>
<caption><p>(Associative) SOINN&#x0002B; Calculation of similarity thresholds</p></caption>
<table frame="hsides" rules="groups">
<tbody>
<tr><td align="left" valign="top">1: <bold>for</bold> <italic>i</italic> in [<italic>b, s</italic>] <bold>do</bold></td></tr>
<tr><td align="left" valign="top">2: &#x000A0;&#x000A0;<bold>if</bold> <italic>N</italic><sub><italic>i</italic></sub> &#x02260; &#x02205; <bold>then</bold></td></tr>
<tr><td align="left" valign="top">3: &#x000A0;&#x000A0;&#x000A0;&#x000A0;<inline-formula><mml:math id="M5"><mml:mi>&#x003C4;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02190;</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder><mml:mo>&#x02016;</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02016;</mml:mo></mml:math></inline-formula> &#x000A0;&#x000A0;&#x000A0;// (Equation 1) </td></tr>
<tr><td align="left" valign="top">4: &#x000A0;&#x000A0;<bold>else</bold></td></tr>
<tr><td align="left" valign="top">5: &#x000A0;&#x000A0;&#x000A0;&#x000A0;<inline-formula><mml:math id="M6"><mml:mi>&#x003C4;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02190;</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">min</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>A</mml:mi><mml:mo>\</mml:mo><mml:mrow><mml:mo stretchy="false">{</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">}</mml:mo></mml:mrow></mml:mrow></mml:munder><mml:mo>&#x02016;</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02016;</mml:mo></mml:math></inline-formula> &#x000A0;&#x000A0;&#x000A0;// (Equation 1) </td></tr>
<tr><td align="left" valign="top">6: &#x000A0;&#x000A0;<bold>end if</bold> </td></tr>
<tr><td align="left" valign="top">7: <bold>end for</bold></td></tr> 
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>Node Merging</title>
<p>In the node merging subroutine (Algorithm 3), the weights of <italic>b</italic> and its neighbors <italic>i</italic> &#x02208; <italic>N</italic><sub><italic>b</italic></sub> are adapted toward the input with:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>W</mml:mi><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mtext class="textrm" mathvariant="normal">,</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E4"><label>(4)</label><mml:math id="M8"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mo>&#x00394;</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>W</mml:mi><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mtext class="textrm" mathvariant="normal">.</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>While Wiwatcharakoses and Berrar (<xref ref-type="bibr" rid="B45">2019</xref>) present a pull factor &#x003B7; to regulate the neighboring neurons&#x00027; inertia, we use the learning rate &#x003F5;<sub><italic>n</italic></sub> to resemble the GDM approach better. It can be considered as the inverse pull factor <inline-formula><mml:math id="M9"><mml:msub><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>&#x003B7;</mml:mi></mml:mrow></mml:mfrac></mml:math></inline-formula>. A further learning rate &#x003F5;<sub><italic>b</italic></sub> is introduced to facilitate a more controlled weight adaptation of the BMU. The amount of learning is also modulated by the winning time to avoid strong adaptations of well-trained units.</p>
<table-wrap position="float">
<label>Algorithm 3</label>
<caption><p>(Associative) SOINN&#x0002B; Node Merging</p></caption>
<table frame="hsides" rules="groups">
<tbody>
<tr><td align="left" valign="top">1: <italic>WT</italic>(<italic>b</italic>) &#x02190; <italic>WT</italic>(<italic>b</italic>) &#x0002B; 1</td></tr>
<tr><td align="left" valign="top">2: <bold>(A-SOINN&#x0002B;)</bold>: <inline-formula><mml:math id="M10"><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>W</mml:mi><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></inline-formula> &#x000A0;&#x000A0;&#x000A0;// Update BMU&#x00027;s weight vector (Equation 3).</td></tr>
<tr><td align="left" valign="top">3: <bold>for</bold> <italic>i</italic> in <italic>N</italic><sub><italic>b</italic></sub> <bold>do</bold></td></tr>
<tr><td align="left" valign="top">4: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<inline-formula><mml:math id="M11"><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>W</mml:mi><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></inline-formula> &#x000A0;&#x000A0;&#x000A0;// Update the neighbors of b (Equation 4). </td></tr>
<tr><td align="left" valign="top">5: <bold>end for</bold></td></tr>
<tr><td align="left" valign="top">6: <italic>IT</italic>(<italic>b</italic>) &#x02190; 0 &#x000A0;&#x000A0;&#x000A0;// Set the idle time of the BMU to zero.</td></tr> 
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>Node Linking</title>
<p>The node linking subroutine determines whether the BMU <italic>b</italic> and the sBMU <italic>s</italic> are connected (Algorithm 4). Here, no modifications are made compared to the original approach. A connection between <italic>b</italic> and <italic>s</italic> is created if at least one of three conditions (Algorithm 4, Step 9) is true. The first is that the number of edges in the network is lower than three to ensure edge creation in the initial training steps. The second and third conditions depend on the trustworthiness <italic>T</italic>(<italic>i</italic>) of a node <italic>i</italic>, which is defined as:</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M12"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>W</mml:mi><mml:mi>T</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>with the maximum winning time of all nodes <italic>max</italic>(<italic>WT</italic>). Furthermore, the values <inline-formula><mml:math id="M13"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:math></inline-formula>, <inline-formula><mml:math id="M14"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:math></inline-formula>, &#x003C3;<sub><italic>b</italic></sub>, and &#x003C3;<sub><italic>s</italic></sub> are required. Let &#x003C4;<sub><italic>b</italic></sub> be the set of similarity thresholds of each connected BMU (i.e., having at least one edge) ever encountered. <inline-formula><mml:math id="M15"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:math></inline-formula> is the arithmetic mean of &#x003C4;<sub><italic>b</italic></sub> and &#x003C3;<sub><italic>b</italic></sub> the corresponding standard deviation. <inline-formula><mml:math id="M16"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:math></inline-formula> and &#x003C3;<sub><italic>s</italic></sub> are analogously defined for the sBMUs. The two additional conditions for node linking are defined as:</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M17"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>&#x003C4;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0003C;</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover><mml:mo>&#x0002B;</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mtext class="textrm" mathvariant="normal">,</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E7"><label>(7)</label><mml:math id="M18"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>&#x003C4;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0003C;</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover><mml:mo>&#x0002B;</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mtext class="textrm" mathvariant="normal">.</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Hence, the more trust there is in <italic>b</italic> or <italic>s</italic>, the more likely they are linked. Using this conditional connection technique tries to connect nodes that do not represent noise (Wiwatcharakoses and Berrar, <xref ref-type="bibr" rid="B45">2019</xref>).</p>
<table-wrap position="float">
<label>Algorithm 4</label>
<caption><p>(Associative) SOINN&#x0002B; Node linking</p></caption>
<table frame="hsides" rules="groups">
<tbody>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;1: <bold>for</bold> <italic>i</italic> in <italic>A</italic> <bold>do</bold></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;2: &#x000A0;&#x000A0;&#x000A0;<italic>max</italic>(<italic>WT</italic>) &#x02190; Maximum winning time of all nodes</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;3: &#x000A0;&#x000A0;&#x000A0;<inline-formula><mml:math id="M19"><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02190;</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>W</mml:mi><mml:mi>T</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula> &#x000A0;&#x000A0;&#x000A0;// Update trustworthiness of each node (Equation 5). </td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;4: <bold>end for</bold></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;5: <inline-formula><mml:math id="M20"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:math></inline-formula>, &#x003C3;<sub><italic>b</italic></sub> &#x02190; mean and standard deviation of the similarity thresholds of all BMUs with an edge to sBMU</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;6: <inline-formula><mml:math id="M21"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:math></inline-formula>, &#x003C3;<sub><italic>s</italic></sub> &#x02190; mean and standard deviation of the similarity thresholds of all sBMUs with an edge to BMU</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;7: condition1 <inline-formula><mml:math id="M22"><mml:mo>&#x02190;</mml:mo><mml:mi>&#x003C4;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0003C;</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover><mml:mo>&#x0002B;</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> &#x000A0;&#x000A0;&#x000A0;// (Equation 6)</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;8: condition2 <inline-formula><mml:math id="M23"><mml:mo>&#x02190;</mml:mo><mml:mi>&#x003C4;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0003C;</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover><mml:mo>&#x0002B;</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x000B7;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> &#x000A0;&#x000A0;&#x000A0;// (Equation 7)</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;9: <bold>if</bold> (|<italic>E</italic>| &#x0003C; 3) <bold>or</bold> condition1 <bold>or</bold> condition2 <bold>then</bold></td></tr>
<tr><td align="left" valign="top">10: &#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>if</bold> (<italic>b, s</italic>) &#x02209; <italic>E</italic> <bold>then</bold></td></tr>
<tr><td align="left" valign="top">11: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<italic>E</italic> &#x02190; <italic>E</italic>&#x0222A;{(<italic>b, s</italic>)} &#x000A0;&#x000A0;&#x000A0;// Create an edge between the BMU and the sBMU.</td></tr>
<tr><td align="left" valign="top">12: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;Update <inline-formula><mml:math id="M24"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:math></inline-formula>, <inline-formula><mml:math id="M25"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:math></inline-formula>, &#x003C3;<sub><italic>b</italic></sub>, &#x003C3;<sub><italic>s</italic></sub> </td></tr>
<tr><td align="left" valign="top">13: &#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>end if</bold> </td></tr>
<tr><td align="left" valign="top">14: <bold>end if</bold></td></tr>
<tr><td align="left" valign="top">15: <bold>if</bold> (<italic>b, s</italic>) &#x02208; <italic>E</italic> <bold>then</bold></td></tr>
<tr><td align="left" valign="top">16: &#x000A0;&#x000A0;&#x000A0;&#x000A0;<italic>LT</italic>(<italic>b, s</italic>) &#x02190; 0 &#x000A0;&#x000A0;&#x000A0;// Set the lifetime of the BMU, sBMU edge to zero. </td></tr>
<tr><td align="left" valign="top">17: <bold>end if</bold></td></tr>
<tr><td align="left" valign="top">18: <bold>for</bold> all <italic>i</italic> in <italic>N</italic><sub><italic>b</italic></sub> <bold>do</bold></td></tr>
<tr><td align="left" valign="top">19: &#x000A0;&#x000A0;&#x000A0;&#x000A0;<italic>LT</italic>(<italic>b, i</italic>) &#x02190; <italic>LT</italic>(<italic>b, i</italic>) &#x0002B; 1 &#x000A0;&#x000A0;&#x000A0;// Update the lifetime of all neighbors of the BMU. </td></tr>
<tr><td align="left" valign="top">20: <bold>end for</bold></td></tr> 
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>Edge Deletion</title>
<p>The lifetime (or age) <italic>LT</italic>(<italic>i, j</italic>) of an edge <italic>e</italic> &#x0003D; (<italic>i, j</italic>) plays a crucial role in the edge deletion subroutine (Algorithm 5). If <italic>e</italic> is connected to the BMU <italic>b</italic> and the sBMU <italic>s</italic>, its lifetime is reset to one (Algorithm 4, Step 16). If it is connected to <italic>b</italic> but not to <italic>s</italic>, it is increased by one (Algorithm 4, Step 19). First, the lifetimes of all edges that can be reached from <italic>b</italic> are computed (<inline-formula><mml:math id="M26"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) (Algorithm 5, Step 1). Edges with an exceptionally high lifetime (i.e., outliers) are detected using the threshold:</p>
<disp-formula id="E8"><label>(8)</label><mml:math id="M27"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x003C9;</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>75</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x000B7;</mml:mo><mml:mi>I</mml:mi><mml:mi>Q</mml:mi><mml:mi>R</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p><inline-formula><mml:math id="M28"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>75</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> represents the 75<sup><italic>th</italic></sup> percentile of the lifetimes in <inline-formula><mml:math id="M29"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="M30"><mml:mi>I</mml:mi><mml:mi>Q</mml:mi><mml:mi>R</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>75</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>25</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> the interquartile range (Upton and Cook, <xref ref-type="bibr" rid="B44">1996</xref>). Additionally, the average lifetime of all deleted edges <inline-formula><mml:math id="M31"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, as well as the number of deleted edges <inline-formula><mml:math id="M32"><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:math></inline-formula>, are considered, resulting in the final threshold:</p>
<disp-formula id="E9"><label>(9)</label><mml:math id="M33"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mo>&#x003BB;</mml:mo></mml:mrow><mml:mrow><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003C9;</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Edges connected to <italic>b</italic> with a higher lifetime than &#x003BB;<sub><italic>edge</italic></sub> are deleted. This deletion technique removes edges that connect nodes of different clusters to isolate those (Wiwatcharakoses and Berrar, <xref ref-type="bibr" rid="B45">2019</xref>).</p>
<table-wrap position="float">
<label>Algorithm 5</label>
<caption><p>(Associative) SOINN&#x0002B; Edge deletion</p></caption>
<table frame="hsides" rules="groups">
<tbody>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;1: <inline-formula><mml:math id="M34"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo></mml:math></inline-formula> set of lifetimes of edges through which the BMU can be reached</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;2: <inline-formula><mml:math id="M35"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>75</mml:mn></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:mn>75</mml:mn><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:math></inline-formula> percentile of elements in <inline-formula><mml:math id="M36"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;3: <inline-formula><mml:math id="M37"><mml:msub><mml:mrow><mml:mi>&#x003C9;</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>75</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x000B7;</mml:mo><mml:mi>I</mml:mi><mml:mi>Q</mml:mi><mml:mi>R</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> &#x000A0;&#x000A0;&#x000A0;// (Equation 8)</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;4: <inline-formula><mml:math id="M38"><mml:msub><mml:mrow><mml:mo>&#x003BB;</mml:mo></mml:mrow><mml:mrow><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003C9;</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mi>g</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> &#x000A0;&#x000A0;&#x000A0;// (Equation 9)</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;5: <bold>for</bold> all <italic>i</italic> in <italic>N</italic><sub><italic>b</italic></sub> <bold>do</bold></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;6: &#x000A0;&#x000A0;&#x000A0;<bold>if</bold> <italic>LT</italic>(<italic>b, i</italic>) &#x0003E; &#x003BB;<sub><italic>edge</italic></sub> <bold>then</bold></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;7: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<italic>E</italic> &#x02190; <italic>E</italic>\{(<italic>b, i</italic>)} &#x000A0;&#x000A0;&#x000A0;// Remove the edge.</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;8: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>(A-SOINN&#x0002B;)</bold>: Update <inline-formula><mml:math id="M39"><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:math></inline-formula>, <inline-formula><mml:math id="M40"><mml:msubsup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, and <inline-formula><mml:math id="M41"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> </td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;9: &#x000A0;&#x000A0;&#x000A0;<bold>end if</bold> </td></tr>
<tr><td align="left" valign="top">10: <bold>end for</bold></td></tr> 
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>Node Deletion</title>
<p>The node deletion subroutine is executed independently of whether a node was added or merged (Algorithm 1, Step 21). The unutility of a node <italic>i</italic> is defined as:</p>
<disp-formula id="E10"><label>(10)</label><mml:math id="M42"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>I</mml:mi><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>W</mml:mi><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>If a node was rarely chosen as the BMU, its unutility is high. On the other hand, it is low if the node was relatively often a BMU. Let <inline-formula><mml:math id="M43"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:math></inline-formula> be the set of unutilities of not isolated nodes in the network. The threshold for the outlierness of an unutility is defined as:</p>
<disp-formula id="E11"><label>(11)</label><mml:math id="M44"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x003C9;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x000B7;</mml:mo><mml:mi>s</mml:mi><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>D</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>with the scaled median absolute deviation (sMAD) (Leys et al., <xref ref-type="bibr" rid="B19">2013</xref>) of the elements in <inline-formula><mml:math id="M45"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:math></inline-formula>. Wiwatcharakoses and Berrar (<xref ref-type="bibr" rid="B45">2019</xref>) do not explain why the sMAD is used here and the IQR in the edge deletion subroutine. We assume that they achieved a higher performance with this configuration. Let <inline-formula><mml:math id="M46"><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:math></inline-formula> be the number of deleted nodes and <inline-formula><mml:math id="M47"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> the average unutility of these deleted nodes. Furthermore, let <inline-formula><mml:math id="M48"><mml:msub><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">I</mml:mi></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mi>A</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mfrac></mml:math></inline-formula> be the ratio between the unconnected nodes <inline-formula><mml:math id="M49"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">I</mml:mi></mml:mrow></mml:math></inline-formula> and all nodes <italic>A</italic>. The threshold for node deletion is then defined as:</p>
<disp-formula id="E12"><label>(12)</label><mml:math id="M50"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mo>&#x003BB;</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:mo>|</mml:mo><mml:mi>A</mml:mi><mml:mo>\</mml:mo><mml:mi mathvariant="-tex-caligraphic">I</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003C9;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:mo>|</mml:mo><mml:mi>A</mml:mi><mml:mo>\</mml:mo><mml:mi mathvariant="-tex-caligraphic">I</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Three conditions are required to be fulfilled for a node <italic>i</italic> to be deleted: (1) at least one edge exists in the network, (2) the unutility of <italic>i</italic> is larger than &#x003BB;<sub><italic>node</italic></sub>, and (3) <italic>i</italic> is unconnected.</p>
</sec>
<sec>
<title>Further Adaptations</title>
<p>For efficiency reasons, further adaptations to the original algorithm are performed. Instead of maintaining three sets of deleted parts of the network, i.e., the set of all deleted nodes <italic>A</italic><sub><italic>del</italic></sub>, the set of all deleted unutilities <inline-formula><mml:math id="M51"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, and the set of lifetimes of deleted edges <inline-formula><mml:math id="M52"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, the A-SOINN&#x0002B; approach only maintains the properties of those sets. Notably, by continuously updating the sizes <inline-formula><mml:math id="M53"><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>=</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">A</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:math></inline-formula> and <inline-formula><mml:math id="M54"><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:math></inline-formula>, the sums <inline-formula><mml:math id="M55"><mml:msubsup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M56"><mml:msubsup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, as well as the mean values <inline-formula><mml:math id="M57"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="M58"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (Algorithm 5, Step 8 and Algorithm 6, Step 12). Thus, three sets are replaced by six values, reducing the memory requirements, especially for large numbers of deleted nodes and edges. We utilized context learning for the A-SOINN&#x0002B; in the first place. However, preliminary tests showed no improvements if doing so. Hence, we do not consider context learning in the description of the algorithm.</p>
<table-wrap position="float">
<label>Algorithm 6</label>
<caption><p>(Associative) SOINN&#x0002B; Node deletion</p></caption>
<table frame="hsides" rules="groups">
<tbody>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;1: <bold>for</bold> <italic>i</italic> in <italic>A</italic> <bold>do</bold></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;2: &#x000A0;&#x000A0;&#x000A0;&#x000A0;<inline-formula><mml:math id="M59"><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02190;</mml:mo><mml:mfrac><mml:mrow><mml:mi>I</mml:mi><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>W</mml:mi><mml:mi>T</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></inline-formula> &#x000A0;&#x000A0;&#x000A0;// Update unutilities of all nodes (Equation 10). </td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;3: <bold>end for</bold></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;4: <inline-formula><mml:math id="M60"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow><mml:mo>&#x02190;</mml:mo></mml:math></inline-formula> Set of unutilities of all nodes with edges</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;5: <inline-formula><mml:math id="M61"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">I</mml:mi></mml:mrow><mml:mo>&#x02190;</mml:mo></mml:math></inline-formula> Set of unconnected nodes</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;6: <inline-formula><mml:math id="M62"><mml:msub><mml:mrow><mml:mi>&#x003C9;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:mi>m</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x000B7;</mml:mo><mml:mi>s</mml:mi><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>D</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> &#x000A0;&#x000A0;&#x000A0;// (Equation 11)</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;7: <inline-formula><mml:math id="M63"><mml:msub><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">I</mml:mi></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mi>A</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mfrac></mml:math></inline-formula></td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;8: <inline-formula><mml:math id="M64"><mml:msub><mml:mrow><mml:mo>&#x003BB;</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:mo>|</mml:mo><mml:mi>A</mml:mi><mml:mo>\</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">I</mml:mi></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mfrac><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003C9;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>d</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:mo>|</mml:mo><mml:mi>A</mml:mi><mml:mo>\</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">I</mml:mi></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> &#x000A0;&#x000A0;&#x000A0;// (Equation 12)</td></tr>
<tr><td align="left" valign="top">&#x000A0;&#x000A0;9: <bold>for</bold> <italic>i</italic> in <italic>A</italic> <bold>do</bold></td></tr>
<tr><td align="left" valign="top">10: &#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>if</bold> |<italic>E</italic>| &#x0003E; 1 <bold>and</bold> <italic>U</italic>(<italic>i</italic>) &#x0003E; &#x003BB;<sub><italic>node</italic></sub> <bold>and</bold> <italic>i</italic> has no edges <bold>then</bold></td></tr>
<tr><td align="left" valign="top">11: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<italic>A</italic> &#x02190; <italic>A</italic>\{<italic>i</italic>} &#x000A0;&#x000A0;&#x000A0;// Delete the node.</td></tr>
<tr><td align="left" valign="top">12: &#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>(A-SOINN&#x0002B;)</bold>: Update <inline-formula><mml:math id="M65"><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:math></inline-formula>, <inline-formula><mml:math id="M66"><mml:msubsup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, and <inline-formula><mml:math id="M67"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">U</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x00304;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> </td></tr>
<tr><td align="left" valign="top">13: &#x000A0;&#x000A0;&#x000A0;&#x000A0;<bold>end if</bold> </td></tr>
<tr><td align="left" valign="top">14: <bold>end for</bold></td></tr>
<tr><td align="left" valign="top">15: <bold>for</bold> <italic>i</italic> in <italic>A</italic> <bold>do</bold></td></tr>
<tr><td align="left" valign="top">16: &#x000A0;&#x000A0;&#x000A0;&#x000A0;<italic>IT</italic>(<italic>i</italic>) &#x02190; <italic>IT</italic>(<italic>i</italic>) &#x0002B; 1 &#x000A0;&#x000A0;&#x000A0;// Increment the idle time of each node. </td></tr>
<tr><td align="left" valign="top">17: <bold>end for</bold></td></tr> 
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec>
<title>3.2. v-NICO-World-LL, a New LL Dataset</title>
<p>In long-term HRI, robots are required to actively explore and learn about their environment while interacting with a human. Part of this is grasping items and viewing them from different angles (along with collecting other multi-modal data, like tactile information). We make a step toward this behavior with our proposed LL dataset, the &#x0201C;virtual NICO World for Lifelong Learning&#x0201D;(<italic>v-NICO-World-LL</italic>). It is inspired by the CORe50 of Lomonaco and Maltoni (<xref ref-type="bibr" rid="B22">2017</xref>) but exhibits three novel features currently not considered by other object recognition LL datasets:</p>
<list list-type="order">
<list-item><p>The environments are grouped into different levels of complexity.</p></list-item>
<list-item><p>A robot is manipulating the objects.</p></list-item>
<list-item><p>The dataset is recorded in a nearly photorealistic controlled virtual environment.</p></list-item>
</list>
<p>The proposed dataset is recorded in a controlled virtual environment in Blender (Blender Foundation, <xref ref-type="bibr" rid="B4">2020</xref>) with a 3D model of the humanoid robot NICO (Kerzel et al., <xref ref-type="bibr" rid="B14">2017</xref>; Kerzel et al., <xref ref-type="bibr" rid="B13">2020</xref>) developed by the Knowledge Technology Group at the Universit&#x000E4;t Hamburg. Using a virtual environment facilitates a controlled and reproducible data acquisition process. In the CORe50 and the iCubWorld datasets, a human operator is holding the objects. However, in a long-term HRI scenario, a robot needs to recognize objects in its hand as well. We fill this gap by placing objects in the robot&#x00027;s left hand. This aspect leads to a different perspective on the object. Furthermore, the robot hand differs from a human hand due to its shape and color (e.g., it has only three fingers and is white-colored). This results in different occlusions and contrasts compared to a human hand. The fingers are positioned so that it looks as if the robot is holding the object. <xref ref-type="fig" rid="F2">Figure 2A</xref> shows the general scenario. NICO starts with an outstretched arm in front of a background panel. The black lines demonstrate the field of vision of the main camera placed at the right eye and filling the entire background panel. The same scenario but from the main camera&#x00027;s perspective is shown in <xref ref-type="fig" rid="F2">Figure 2B</xref>. Since only a fraction of the frames contains the object of interest, cropped images focusing on the object are created using a second camera (orange lines). This camera moves in two dimensions (i.e., left/right and up/down) to follow the robot&#x00027;s hand.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>The Blender scenario from two different perspectives. <bold>(A)</bold> shows the robot, the plane, and the two fields of vision as black and orange lines. While the former lines represent the fixed main camera positioned at the right eye, the latter represent the moving camera focused on the robot&#x00027;s hand. <bold>(B)</bold> is showing the perspective of the main camera.</p></caption>
<graphic xlink:href="fnbot-15-669534-g0002.tif"/>
</fig>
<p>The robot&#x00027;s arm is moving smoothly in a 10-s animation by manipulating an elbow, a wrist, and a palm joint, as well as three shoulder joints. RGB images are rendered with 24 fps from the perspective of both cameras. The image resolutions are 512 &#x000D7; 512 pixels and 256 &#x000D7; 256 pixels for the main and moving camera, respectively. The dataset consists of 100 objects belonging to 10 categories: ball, book, bottle, cup, doughnut, glasses, pen, pocket watch, present, and vase. These objects are chosen such that they could appear in a long-term HRI scenario. The 10 instances of a category <italic>C</italic> are denoted with <italic>I</italic><sub>0</sub> to <italic>I</italic><sub>9</sub> throughout this work. All objects are downloaded from the four websites: <ext-link ext-link-type="uri" xlink:href="https://www.turbosquid.com">https://www.turbosquid.com</ext-link>, <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com">https://sketchfab.com</ext-link>, <ext-link ext-link-type="uri" xlink:href="https://free3d.com">https://free3d.com</ext-link>, and <ext-link ext-link-type="uri" xlink:href="https://www.cgtrader.com">https://www.cgtrader.com</ext-link> in the period from April 26, 2020, to June 6, 2020. They are slightly modified in Blender to fit into the scenario. The objects are manipulated in front of 20 different backgrounds, divided into four background complexities, each with five different background instances. These different background complexities allow the controlled evaluation of how LL models respond to different degrees of environmental conditions. <xref ref-type="fig" rid="F3">Figure 3</xref> shows all backgrounds. Each row represents a background complexity, and each column an instance of this complexity. <italic>Background complexity</italic> 1 (B1), shown in the first row of <xref ref-type="fig" rid="F3">Figure 3</xref>, is composed of five plain colored images. Real-world background images are used in B2 and B3 (second and third row, respectively) to simulate different environments where long-term HRI can occur. While complexity B2 consists of images with low variability (i.e., simple walls with little content), B3 contains more complex indoor and outdoor scenes with different visible objects, like furniture or plants. The real robot&#x00027;s total height is 101 cm (Kerzel et al., <xref ref-type="bibr" rid="B14">2017</xref>), such that the cameras of NICO are approximately 89 cm above the ground. Hence, each background image of B2 and B3 is captured from a height of 89 cm to create more realistic recordings. B4 (fourth row) comprises artificial images with a cluttered dynamic structure where drawn objects move in the background. Each instance of B2, B3, and B4 has a different lighting condition in the virtual environment. In B2 and B3, the real-world lighting conditions are tried to be reproduced. By grouping the background images into different complexity levels, the influence of different degrees of environmental conditions on a model can be investigated, making it the key feature of the v-NICO-World-LL dataset. Ten example images of the dataset captured from the moving camera&#x00027;s perspective are shown in <xref ref-type="fig" rid="F4">Figure 4</xref>. It illustrates the dataset&#x00027;s diversity regarding the background, lighting condition, and NICO&#x00027;s hand position. In total, the dataset consists of two times 2,000 videos (cropped and non-cropped), each having 240 frames, resulting in two times 480,000 RGB images.<xref ref-type="fn" rid="fn0002"><sup>2</sup></xref></p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>The 20 different backgrounds of the dataset. The images are split into four complexity groups (B1, B2, B3, and B4), each shown in one row with increasing complexity from top to bottom. The first row shows the plain colored background images of B1. B2 and B3 (second and third row, respectively) are showing real-world backgrounds. A camera captured the B2 and B3 examples from a height of 89 cm. B4 (fourth row) consists of artificial background images with different colors and drawn objects.</p></caption>
<graphic xlink:href="fnbot-15-669534-g0003.tif"/>
</fig>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>This figure presents 10 example images of the dataset captured from the moving camera&#x00027;s perspective. Ten objects are shown in front of different backgrounds with different lighting conditions. Additionally, NICO is holding the objects in different positions. <bold>(A)</bold> 3D model of <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/3d-models/tennis-ball-11e572d8b4214f02955ed4f97dfe2afc">&#x0201C;Tennis ball&#x0201D;</ext-link> by <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/tr3vinj">&#x0201C;tr3vinj&#x0201D;</ext-link>, <bold>(B)</bold> <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/3d-models/metal-cup-ww2-style-cup-vintage-d70c2777df9e4971a7303a6d9da2dd97">&#x0201C;Metal Cup&#x0201D;</ext-link> by <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/warkarma">&#x0201C;Warkarma&#x0201D;</ext-link>, <bold>(C)</bold> <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/3d-models/mug-2933d41e99fa4a0882a67fd123f84de2">&#x0201C;Mug&#x0201D;</ext-link> by <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/kriszcvc">&#x0201C;Kristijan Zecevic&#x0201D;</ext-link>, <bold>(D)</bold> <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/3d-models/beach-ball-free-download-c915a99e9bae4dbe8dd7be3215e19ba0">&#x0201C;Beach Ball&#x0201D;</ext-link> by <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/tommyleenev">&#x0201C;Tommyleenev&#x0201D;</ext-link>, <bold>(E)</bold> <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/3d-models/3d-wine-bottle-model-399587f06b894b00bea59ee03841dbb6">&#x0201C;3D Wine Bottle Model&#x0201D;</ext-link> by <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/Johana-PS">&#x0201C;Johana-PS&#x0201D;</ext-link>, <bold>(F)</bold> <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/3d-models/glass-bottle-00beb4f7241e496e8a2ebb47fdb4f93c">&#x0201C;Glass Bottle&#x0201D;</ext-link> by <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/Aullwen">&#x0201C;Aullwen&#x0201D;</ext-link>, <bold>(G)</bold> <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/3d-models/gift-3december-742efa7beef24cbfb2ffe4bef1d2cbc1">&#x0201C;Gift - 3December&#x0201D;</ext-link> by <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/zhixson">&#x0201C;zhixson&#x0201D;</ext-link>, <bold>(H)</bold> Derivative of <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/3d-models/coffee-and-donuts-0279c25ef4884f99af7894f9a0710331">&#x0201C;Coffee and donuts&#x0201D;</ext-link> by <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/ivanlynx">&#x0201C;Ivan Dnistrian&#x0201D;</ext-link>, <bold>(I)</bold> Derivative of <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/3d-models/antique-pocket-watch-fa99fca27aa54adaabb03473616ff117">&#x0201C;Antique Pocket Watch&#x0201D;</ext-link> by <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/michael_grodkowski">&#x0201C;michael_grodkowski&#x0201D;</ext-link>, and <bold>(J)</bold> Derivative of <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/3d-models/fountain-pen-196c00b1bb5d4958816a17e8947bd65a">&#x0201C;Fountain Pen&#x0201D;</ext-link> by <ext-link ext-link-type="uri" xlink:href="https://sketchfab.com/etherlyte">&#x0201C;Etherlyte&#x0201D;</ext-link>. All 10 models are used under <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</ext-link> and also licensed under <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">CC BY 4.0</ext-link>.</p></caption>
<graphic xlink:href="fnbot-15-669534-g0004.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<title>4. Experiments and Results</title>
<sec>
<title>4.1. Experimental Setup</title>
<p>Five different experiments, organized into two parts, are conducted. Each experiment considers different data samples. The background complexities B1, B2, B3, and B4, are utilized in the first four experiments. In the last experiment, the CORe50 is used to determine the robustness of the A-SOINN&#x0002B; against a different and well-known dataset. In each experiment, three GDM models with different maximal ages are compared to two versions of the A-SOINN&#x0002B;. The considered maximal ages are 6,000, 12,000, and 28,000. These models are referred to as <italic>GDM 6,000, GDM 12,000</italic>, and <italic>GDM 28,000</italic>, respectively. Empirical tests have shown that these maximal age values exhibit a good trade-off between accuracy and neuron number. Additionally, we investigate the performance of the A-SOINN&#x0002B; with and without the node creation constraint to examine this constraint&#x00027;s influence on neuron number and accuracy. The A-SOINN&#x0002B; without creation constraint is referred to as (ncc)A-SOINN&#x0002B;. We focus on analyzing this constraint as it is directly linked to the two criteria, which we apply to evaluate the different approaches: the task performance and the number of created nodes. In each experiment, the models are required to learn 10 different tasks. A task <italic>t</italic><sub><italic>i</italic></sub> of time step <italic>i</italic> refers to learning to classify a particular category of the v-NICO-World-LL or the CORe50 dataset. We compare the A-SOINN&#x0002B; only to the GDM due to two reasons. First, the GDM shows state-of-the-art LL performance compared to other LL approaches (Parisi et al., <xref ref-type="bibr" rid="B31">2018</xref>). Second, both A-SOINN&#x0002B; and GDM show a strong architectural resemblance as they are based on the GWR (Marsland et al., <xref ref-type="bibr" rid="B27">2002</xref>) approach. In future work, further comparisons can be examined.</p>
<sec>
<title>v-NICO-World-LL Experiments</title>
<p>For the first four experiments, each frame is transformed into a feature vector using a VGG-16 (Simonyan and Zisserman, <xref ref-type="bibr" rid="B42">2014</xref>) feature extractor. It is pre-trained on the ImageNet (Russakovsky et al., <xref ref-type="bibr" rid="B38">2015</xref>) dataset and fine-tuned using the v-NICO-World-LL. The original VGG-16 is adapted to have two 2048-dimensional and one 128-dimensional fully connected hidden layers after the last pooling layer. The output layer is 10-dimensional (one for each category). Each frame is transformed using the third hidden layer&#x00027;s 128-dimensional output. In contrast to Parisi et al. (<xref ref-type="bibr" rid="B31">2018</xref>), we use 128-dimensional vectors instead of 256-dimensional to reduce the training time. The first three instances (<italic>I</italic><sub>0</sub>-<italic>I</italic><sub>2</sub>) of each category are used to train the four mentioned fully connected layers. Instances <italic>I</italic><sub>3</sub> to <italic>I</italic><sub>7</sub> are used to validate the model and <italic>I</italic><sub>8</sub> and <italic>I</italic><sub>9</sub> to test it. Five runs are performed with a particular hyperparameter assignment, resulting in a test accuracy of 84.2 &#x000B1; 0.32% averaged across these runs. No hyperparameter optimization is done here since it is not essential to find an optimal accuracy in this work. Instead, sufficiently high accuracy results for reliable feature vector creation are required. A video is represented by a <italic>feature vector sequence</italic> (FVS) of the corresponding frames. Only each second frame of a video is considered for training to reduce the training time further.</p>
<p>At each iteration <italic>i</italic>, a task-dependent training set <inline-formula><mml:math id="M68"><mml:mi>T</mml:mi><mml:msubsup><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is used to train the models. It contains the FVS of the training instances <italic>I</italic><sub><italic>i</italic>,3</sub> to <italic>I</italic><sub><italic>i</italic>,7</sub> of the category <italic>C</italic><sub><italic>i</italic></sub>. The first three instances of each category (<italic>I</italic><sub>0</sub>-<italic>I</italic><sub>2</sub>) are not considered in the LL experiments since the feature extractor is trained on these. This strategy avoids that the feature extractor learns &#x0201C;future&#x0201D; objects on which the LL models are later trained. This contrasts with the work of Parisi et al. (<xref ref-type="bibr" rid="B31">2018</xref>), where the feature extractor was trained on all object instances. After each learned task, task-dependent test sets are used to measure the performance of each model. Single-head evaluation (Chaudhry et al., <xref ref-type="bibr" rid="B5">2018</xref>) is performed, i.e., a task-dependent test set <italic>Te</italic><sub><italic>i</italic></sub> contains the instances <italic>I</italic><sub><italic>i</italic>,8</sub> and <italic>I</italic><sub><italic>i</italic>,9</sub>, as well as <italic>I</italic><sub><italic>j</italic>,8</sub> and <italic>I</italic><sub><italic>j</italic>,9</sub> of all previous tasks <italic>t</italic><sub><italic>j</italic></sub>, with <italic>j</italic> &#x0003C; <italic>i</italic>. Hence, the models are required to predict the category of the current task and the previous tasks. Additionally, the networks are not only tested on the test set <italic>Te</italic><sub><italic>i</italic></sub> after learning <italic>t</italic><sub><italic>i</italic></sub> but also all previous <italic>Te</italic><sub><italic>j</italic></sub>, with <italic>j</italic> &#x0003C; <italic>i</italic>. Three metrics are investigated: the average accuracy (Definition 1), the average forgetting (Definition 2), and the number of neurons. For the GDM models, the units of the G-EM and G-SM layers are summed up.</p>
<p><bold>Definition 1</bold>. (Average Accuracy Chaudhry et al., <xref ref-type="bibr" rid="B5">2018</xref>)</p>
<p>Let <italic>a</italic><sub><italic>k,j</italic></sub> &#x02208; [0, 1] be the accuracy evaluated on the test set <italic>Te</italic><sub><italic>j</italic></sub> of task <italic>t</italic><sub><italic>j</italic></sub> with <italic>j</italic> &#x02264; <italic>k</italic> after training the model incrementally from <italic>t</italic><sub>1</sub> to <italic>t</italic><sub><italic>k</italic></sub>. The average accuracy at task <italic>t</italic><sub><italic>k</italic></sub> is:</p>
<disp-formula id="E13"><mml:math id="M69"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p><bold>Definition 2</bold>. (Average Forgetting Chaudhry et al., <xref ref-type="bibr" rid="B5">2018</xref>) Forgetting <inline-formula><mml:math id="M70"><mml:msubsup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> for a task <italic>t</italic><sub><italic>j</italic></sub>, after the model learned from task <italic>t</italic><sub>1</sub> to <italic>t</italic><sub><italic>k</italic></sub>, is the difference between the maximum knowledge gained about the task in the past and the knowledge the model currently has about it: <inline-formula><mml:math id="M71"><mml:msubsup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:mspace width="0.3em" class="thinspace"/><mml:mo>,</mml:mo><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:munder><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02200;</mml:mo><mml:mi>j</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mi>k</mml:mi><mml:mo>.</mml:mo></mml:math></inline-formula> The average forgetting at task <italic>t</italic><sub><italic>k</italic></sub> is, therefore:</p>
<disp-formula id="E14"><mml:math id="M72"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover></mml:mstyle><mml:msubsup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msubsup><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Hence, <italic>A</italic><sub><italic>k</italic></sub> represents the performance of a model on the current task <italic>t</italic><sub><italic>k</italic></sub> and all previous tasks. <italic>F</italic><sub><italic>k</italic></sub> represents how much a model forgot about all previous tasks given the current time step <italic>k</italic>. The order of the tasks can affect the final results (Lomonaco and Maltoni, <xref ref-type="bibr" rid="B22">2017</xref>; Parisi et al., <xref ref-type="bibr" rid="B31">2018</xref>). Hence, each experiment is repeated four times, and the results are averaged. Each repetition has a different, randomly generated, category order:</p>
<list list-type="bullet">
<list-item><p>R1: ball, bottle, cup, doughnut, glasses, pen, pocket watch, present, vase, book</p></list-item>
<list-item><p>R2: ball, book, bottle, cup, doughnut, glasses, pen, pocket watch, present, vase</p></list-item>
<list-item><p>R3: doughnut, ball, pocket watch, glasses, pen, book, bottle, present, vase, cup</p></list-item>
<list-item><p>R4: present, cup, pocket watch, vase, doughnut, pen, glasses, book, ball, bottle</p></list-item>
</list>
<p>The GDM, and the A-SOINN&#x0002B; approach, have hyperparameters required to be tuned to preserve reliable models. Therefore, a grid search is performed for each model. Since each experiment requires 20 models (three GDM and two A-SOINN&#x0002B; for each repetition), 20 grid searches are executed for a single experiment. The considered hyperparameters and assignments of the GDM and A-SOINN&#x0002B; are shown in <xref ref-type="table" rid="T1">Tables 1</xref>, <xref ref-type="table" rid="T2">2</xref>, respectively. Memory replay is used for each GDM model since Parisi et al. (<xref ref-type="bibr" rid="B31">2018</xref>) already showed that it improves the performance of the GDM approach. Preliminary tests confirmed this for the proposed dataset. According to Parisi et al. (<xref ref-type="bibr" rid="B31">2018</xref>), <italic>K</italic> values larger than two do not improve the GDM model&#x00027;s performance. Therefore, only two assignments are considered.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>This table shows the considered GDM hyperparameter assignments for the grid searches of the first four experiments.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Parameter</bold></th>
<th valign="top" align="left"><bold>Description</bold></th>
<th valign="top" align="left"><bold>Assignments</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic>K</italic></td>
<td valign="top" align="left">No. of context vectors</td>
<td valign="top" align="left"><bold>2</bold>, 4</td>
</tr>
<tr>
<td valign="top" align="left">[&#x003F5;<sub><italic>b</italic></sub>, &#x003F5;<sub><italic>n</italic></sub>]</td>
<td valign="top" align="left">Learning rates</td>
<td valign="top" align="left">[0.1, 0.001], <bold>[0.3, 0.003]</bold>, [0.5, 0.005]</td>
</tr>
<tr>
<td valign="top" align="left"><inline-formula><mml:math id="M73"><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula></td>
<td valign="top" align="left">Activation thresholds</td>
<td valign="top" align="left"><bold>[0.1, 0.01]</bold>, [0.3, 0.03], [0.5, 0.05]</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>K denotes the number of context vectors. &#x003F5;<sub>b</sub> and &#x003F5;<sub>n</sub> represent the learning rates of the BMU and its neighbors, respectively. The activation thresholds of the G-EM and the G-SM layers are denoted with <inline-formula><mml:math id="M74"><mml:msubsup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M75"><mml:msubsup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, respectively. The best results averaged across all grid search iterations are shown in bold letters</italic>.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>This table shows the considered A-SOINN&#x0002B; hyperparameter assignments for the grid searches of the first four experiments.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Parameter</bold></th>
<th valign="top" align="left"><bold>Description</bold></th>
<th valign="top" align="left"><bold>Assignments</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic>K</italic></td>
<td valign="top" align="left">No. of context vectors</td>
<td valign="top" align="left"><bold>2</bold>, 4</td>
</tr>
<tr>
<td valign="top" align="left">[&#x003F5;<sub><italic>b</italic></sub>, &#x003F5;<sub><italic>n</italic></sub>]</td>
<td valign="top" align="left">Learning rates</td>
<td valign="top" align="left">[0.3, 0.003], [0.5, 0.005], [1, 0.001], [2, 0.001], <bold>[3, 0.002]</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>K denotes the number of context vectors. &#x003F5;<sub>b</sub> and &#x003F5;<sub>n</sub> represent the learning rates of the BMU and its neighbors, respectively. The best results averaged across all grid search iterations are shown in bold letters</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>Both tables show the best assignments averaged across all grid search iterations in bold letters. We considered context learning for the A-SOINN&#x0002B; in the first place. Therefore, we utilize the number <italic>K</italic> of context vectors in our grid searches. However, as preliminary tests showed, no improvements were observable if including context learning into the A-SOINN&#x0002B;. This grid search confirms this assumption since both assignments show similar average accuracy values (<italic>K</italic> = 2: 83.9%, <italic>K</italic> = 4: 83.7%). The best models of each grid search are used for a more detailed analysis by investigating the three mentioned metrics.</p>
</sec>
<sec>
<title>CORe50 Experiment</title>
<p>In the last experiment, we aim to maintain consistency with (Parisi et al., <xref ref-type="bibr" rid="B31">2018</xref>), resulting in several differences compared to the previous four experiments. A new feature extractor is fine-tuned on the CORe50 to predict the 50 object instances of the dataset. The last hidden layer is 256-dimensional, resulting in 256-dimensional feature vectors. The CORe50 contains 11 backgrounds/sessions. Samples from sessions &#x00023;3, &#x00023;7, and &#x00023;10 are used for testing the feature extractor and the remaining for training. After five runs with one particular parametrization, an average accuracy of 76.47% &#x000B1; 0.21% is achieved. The same train/test split is used for LL afterward. Therefore, the task-dependent training set of a particular iteration contains all five instances of one object category from all sessions except &#x00023;3, &#x00023;7, and &#x00023;10. Again four different repetitions are performed to reduce the category order influence. The randomly generated category orders are:</p>
<list list-type="bullet">
<list-item><p>R1: plug adapter, glasses, light bulb, scissors, cup, remote control, marker, can, mobile phone, ball</p></list-item>
<list-item><p>R2: scissors, cup, plug adapter, light bulb, can, glasses, marker, ball, mobile phone, remote control</p></list-item>
<list-item><p>R3: mobile phone, scissors, remote control, ball, light bulb, can, plug adapter, glasses, cup, marker</p></list-item>
<list-item><p>R4: glasses, scissors, can, light bulb, cup, plug adapter, ball, mobile phone, remote control, marker</p></list-item>
</list>
<p>The same assignments as shown before (<xref ref-type="table" rid="T1">Tables 1</xref>, <xref ref-type="table" rid="T2">2</xref>) are used for hyperparameter optimization (see <xref ref-type="table" rid="T3">Tables 3</xref>, <xref ref-type="table" rid="T4">4</xref>). However, the activation threshold assignments are extended by three values: [0.01, 0.01], [0.05, 0.05], and [0.07, 0.07] (see <xref ref-type="table" rid="T3">Table 3</xref>) to reduce the node creation probability, which can lead to smaller GDM networks. The best assignments (bold letters) reveal that the GDM can benefit from having a smaller activation threshold (i.e., [0.05, 0.05] with 85.3% average accuracy). However, smaller values, like [0.01, 0.01], demonstrate only 39% average accuracy, indicating that even smaller values will show worse results and that we found a reliable assignment for this hyperparameter.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>This table shows the considered GDM hyperparameter assignments for the grid searches of the CORe50 experiment.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Parameter</bold></th>
<th valign="top" align="left"><bold>Description</bold></th>
<th valign="top" align="left"><bold>Assignments</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic>K</italic></td>
<td valign="top" align="left">No. of context vectors</td>
<td valign="top" align="left"><bold>2</bold>, 4</td>
</tr>
<tr>
<td valign="top" align="left">[&#x003F5;<sub><italic>b</italic></sub>, &#x003F5;<sub><italic>n</italic></sub>]</td>
<td valign="top" align="left">Learning rates</td>
<td valign="top" align="left">[0.1, 0.001], <bold>[0.3, 0.003]</bold>, [0.5, 0.005]</td>
</tr>
<tr>
<td valign="top" align="left"><inline-formula><mml:math id="M76"><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>E</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula></td>
<td valign="top" align="left">Activation thresholds</td>
<td valign="top" align="left">[0.1, 0.01], [0.3, 0.03], [0.5, 0.05], [0.01, 0.01], <bold>[0.05, 0.05]</bold>, [0.07, 0.07]</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The best results averaged across all grid search iterations are shown in bold letters. For further descriptions of the parameters, see <xref ref-type="table" rid="T1">Table 1</xref></italic>.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>This table shows the considered A-SOINN&#x0002B; hyperparameter assignments for the grid searches of the CORe50 experiment.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Parameter</bold></th>
<th valign="top" align="left"><bold>Description</bold></th>
<th valign="top" align="left"><bold>Assignments</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic>K</italic></td>
<td valign="top" align="left">No. of context vectors</td>
<td valign="top" align="left"><bold>2</bold>, 4</td>
</tr>
<tr>
<td valign="top" align="left">[&#x003F5;<sub><italic>b</italic></sub>, &#x003F5;<sub><italic>n</italic></sub>]</td>
<td valign="top" align="left">Learning rates</td>
<td valign="top" align="left">[0.3, 0.003], [0.5, 0.005], [1, 0.001], <bold>[2, 0.001]</bold>, [3, 0.002]</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The best results averaged across all grid search iterations are shown in bold letters. For further descriptions of the parameters, see <xref ref-type="table" rid="T2">Table 2</xref></italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec>
<title>4.2. Results</title>
<p><xref ref-type="fig" rid="F5">Figures 5</xref>&#x02013;<xref ref-type="fig" rid="F9">9</xref> show the results of all five experiments. In the presented graphs, the results of the models GDM 6,000, GDM 12,000, GDM 28,000, A-SOINN&#x0002B;, and (ncc)A-SOINN&#x0002B; are shown in blue, green, red, purple, and yellow, respectively. The shaded areas represent the standard deviation created by computing the arithmetic mean across the four repetitions. The x-axes show the number of tasks already learned by the models, hence the seen categories. Sub-figures labeled with the letter A are showing the average accuracy results in the range 0.4 to 1.0, B the average forgetting results in the range &#x02212;0.04 to 0.4, C the number of units of each model in the range 0&#x02013;10, 000, and D the number of units focused on the A-SOINN&#x0002B; in the range zero to 100.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>The results of the first experiment. GDM 6,000, GDM 12,000, GDM 28,000, A-SOINN&#x0002B;, and A-SOINN&#x0002B; without node creation constraint are represented by blue, green, red, purple, and yellow lines. The x-axes represent the number of learned tasks and, therefore, the number of seen categories over time from 1 to 10. <bold>(A)</bold> is showing the average accuracy of each model on the y-axis in the interval 0.4 to 1.0. <bold>(B)</bold> is presenting the average forgetting in the interval &#x02212;0.04 to 0.4, <bold>(C)</bold> the number of units in the interval 0 to 10,000, and <bold>(D)</bold> the number of units focused on the A-SOINN&#x0002B; model in the interval zero to 100. All results are averaged across the four repetitions. The shaded areas represent the resulting standard deviation values.</p></caption>
<graphic xlink:href="fnbot-15-669534-g0005.tif"/>
</fig>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>The results of the second experiment. For further descriptions, see <xref ref-type="fig" rid="F5">Figure 5</xref>.</p></caption>
<graphic xlink:href="fnbot-15-669534-g0006.tif"/>
</fig>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>The results of the third experiment. For further descriptions, see <xref ref-type="fig" rid="F5">Figure 5</xref>.</p></caption>
<graphic xlink:href="fnbot-15-669534-g0007.tif"/>
</fig>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>The results of the fourth experiment. For further descriptions, see <xref ref-type="fig" rid="F5">Figure 5</xref>.</p></caption>
<graphic xlink:href="fnbot-15-669534-g0008.tif"/>
</fig>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p>The results of the fifth experiment, using the CORe50 dataset. For further descriptions, see <xref ref-type="fig" rid="F5">Figure 5</xref>.</p></caption>
<graphic xlink:href="fnbot-15-669534-g0009.tif"/>
</fig>
<sec>
<title>v-NICO-World-LL Experiments</title>
<p>By considering <xref ref-type="fig" rid="F5">Figures 5A</xref>, <xref ref-type="fig" rid="F6">6A</xref>, <xref ref-type="fig" rid="F7">7A</xref>, <xref ref-type="fig" rid="F8">8A</xref>, it can be observed that the average accuracy value generally decreases with an increasing amount of learned tasks, independent of the model or the background complexity. Additionally, the curves are lower for models trained on the two most complex experiments. Among the other GDM models, the GDM 28,000 reaches the highest average accuracy values, followed by GDM 12,000 and GDM 6,000. Hence, the higher the maximal age, the higher the average accuracy. In the first three experiments, the A-SOINN&#x0002B; approach displays similar or even better average accuracy progress than the GDM 28,000. Between tasks three and five of the fourth and most challenging experiment, the accuracy results of the A-SOINN&#x0002B; are lower than those of the GDM 12,000. The standard deviation values of the A-SOINN&#x0002B; are higher in the experiments of <italic>B</italic>1, <italic>B</italic>2, and <italic>B</italic>4. Especially after learning task eight, an increase of the standard deviation is observable, accompanied by a decrease in the average accuracy. This behavior is not visible in the results of the GDM models. The A-SOINN&#x0002B; without node creation constraint shows the highest average accuracies in the first and fourth experiment, together with a lower standard deviation than the A-SOINN&#x0002B; with node creation constraint. After the last task, the (ncc)A-SOINN&#x0002B; and the A-SOINN&#x0002B; show the best average accuracy results in each experiment.</p>
<p>No substantial average forgetting (<xref ref-type="fig" rid="F5">Figures 5B</xref>, <xref ref-type="fig" rid="F6">6B</xref>, <xref ref-type="fig" rid="F7">7B</xref>, <xref ref-type="fig" rid="F8">8B</xref>) increases are observable in the GDM model curves, especially not in the first three experiments. In the fourth one, the increase and the standard deviation are higher than in the other three experiments. However, the average forgetting is comparably low and does not exceed the value 0.17. In contrast to that, the average forgetting of the A-SOINN&#x0002B; is higher in each experiment. It increases after the model learned task eight. This increase corresponds to the average accuracy drop after this task mentioned earlier. The (ncc)A-SOINN&#x0002B; shows the second-highest forgetting curves. However, it has no substantial forgetting increase at task eight.</p>
<p>An essential part of this work is comparing each model&#x00027;s memory requirements, which are directly linked to the number of nodes created by the models. It is also related to the computational requirements since for BMU determination, an iteration over the whole network is required. In each experiment (<xref ref-type="fig" rid="F5">Figures 5C</xref>, <xref ref-type="fig" rid="F6">6C</xref>, <xref ref-type="fig" rid="F7">7C</xref>, <xref ref-type="fig" rid="F8">8C</xref>), each GDM model&#x00027;s unit number reaches a maximal point, slightly declines, and converges afterward. The maximum unit number of the GDM 6,000 is approximately 2,000 in each experiment, and the value converges at approximately 1,000 units. In GDM 12,000, the maximal point is approximately 4,000 and converges at approximately 2,800. GDM 28,000 shows a maximal number of approximately 7,000 and converges at approximately 6,000 units in the first, second, and fourth experiment. However, the fourth experiment&#x00027;s standard deviation is higher than in the other. An exception is the third experiment, where the GDM 28,000 shows a maximal value of approximately 9,000, and no convergence can be observed until task 10. However, due to the other results, it is expected that convergence would take place with additional tasks. In general, a GDM model with a higher maximal age shows a larger number of units during LL. Additionally, it reaches the maximum number at a later point in time than GDM models with a lower maximal age. The (ncc)A-SOINN&#x0002B; shows a linear unit number increase in all experiments. The maximum number at the end of the training is approximately 4,000, 2,200, 2,500, and 7,000 in the four experiments, respectively. <xref ref-type="fig" rid="F5">Figures 5D</xref>, <xref ref-type="fig" rid="F6">6D</xref>, <xref ref-type="fig" rid="F7">7D</xref>, <xref ref-type="fig" rid="F8">8D</xref> focus on the unit number of the A-SOINN&#x0002B; approach. The first and the second experiment show a maximal number of approximately 20 nodes with a comparably low standard deviation. The maximal number grows to approximately 30 in the third experiment, and the standard deviation increases. In the fourth experiment, the maximal number of nodes increases to approximately 70. Furthermore, the curve shows a generally higher standard deviation than in the other three experiments. In general, the A-SOINN&#x0002B; is showing the lowest number of neurons. However, this number does not converge like in the GDM approaches. In the first three experiments, it grows roughly linear. In the fourth experiment, the slope increases over time, resulting in a stronger than linear behavior.</p>
</sec>
<sec>
<title>CORe50 Experiment</title>
<p><xref ref-type="fig" rid="F9">Figure 9A</xref> shows that except for the GDM 6,000, all models have a high and nearly constant average accuracy of at least 0.97. Similar behavior is observable in <xref ref-type="fig" rid="F9">Figure 9B</xref>. The average forgetting of all models, except for GDM 6,000, is constantly lower than 0.06. As shown in <xref ref-type="fig" rid="F9">Figure 9C</xref>, the number of units converges for the GDM approaches with a maximum of around 1, 000, 1, 500, and 2, 300 nodes for the GDM 6,000, GDM 12,000, and GDM 2,8000, respectively. The (ncc)A-SOINN&#x0002B; unit number increases linearly and reaches a maximum value of approximately 9, 500. A linear unit number increase is also observable for the A-SOINN&#x0002B; (<xref ref-type="fig" rid="F9">Figure 9D</xref>). However, as in the previous experiments, the slope is much lower than in the (ncc)A-SOINN&#x0002B; with a maximal value of approximately 75 units and, therefore, 30 times fewer than the best GDM model.</p>
<p><xref ref-type="table" rid="T5">Table 5</xref> is summarizing all five experiments. The average accuracy of each model is averaged across all 10 tasks. The A-SOINN&#x0002B; approach achieves the highest values (gray shaded cells) either with or without the node creation constraint. The A-SOINN&#x0002B; with node creation constraint shows either the best or second-best results in the first four experiments. In the CORe50 experiment, the GDM 12,000, GDM 28,000, and both A-SOINN&#x0002B; models show similar high results with an average accuracy of 97.26% for the proposed model. <xref ref-type="table" rid="T6">Table 6</xref> shows the training time (in minutes) and the file size (in megabytes) of the different models in the B4 experiment without considering the testing time. The models are trained on an Intel Core i7-4930 K CPU with 3.40 GHz. Due to the sequential nature of LL, we are not able to parallelize the training. The A-SOINN&#x0002B; shows the lowest training time of around 4 min averaged across all four repetitions. This is followed by the GDM 6,000. GDM 12,000 and (ncc)A-SOINN&#x0002B; show similar training times of 363 &#x000B1; 58 and 370 &#x000B1; 22 min, respectively. The GDM 28,000 exhibits the highest training time of 1,704 &#x000B1; 437 min (&#x02248;28.4 &#x000B1; 7.3 h). The lowest memory requirements of around 0.7 MB are observable in the A-SOINN&#x0002B;, followed by GDM 6,000, GDM 12,000, GDM 28,000, and the (ncc)A-SOINN&#x0002B;, which exhibits the highest memory requirements of around 741 MB.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>This table shows the arithmetic mean of the average accuracies across all 10 tasks (in percent) for each model (columns) in each experiment (rows).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Dataset</bold></th>
<th valign="top" align="center"><bold>GDM 6,000</bold></th>
<th valign="top" align="center"><bold>GDM 12,000</bold></th>
<th valign="top" align="center"><bold>GDM 28,000</bold></th>
<th valign="top" align="center"><bold>A-SOINN&#x0002B;</bold></th>
<th valign="top" align="center"><bold>(ncc)A-SOINN&#x0002B;</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">B1</td>
<td valign="top" align="center">73.95 &#x000B1; 1.78</td>
<td valign="top" align="center">83.88 &#x000B1; 0.66</td>
<td valign="top" align="center">92.60 &#x000B1; 1.31</td>
<td valign="top" align="center">93.07 &#x000B1; 3.47</td>
<td valign="top" align="center" style="background-color:#bbbdc0">95.71 &#x000B1; 1.83</td>
</tr>
<tr>
<td valign="top" align="left">B2</td>
<td valign="top" align="center">73.65 &#x000B1; 1.68</td>
<td valign="top" align="center">81.95 &#x000B1; 2.02</td>
<td valign="top" align="center">89.86 &#x000B1; 2.21</td>
<td valign="top" align="center" style="background-color:#bbbdc0">90.7 &#x000B1; 3.62</td>
<td valign="top" align="center">90.06 &#x000B1; 2.94</td>
</tr>
<tr>
<td valign="top" align="left">B3</td>
<td valign="top" align="center">72.09 &#x000B1; 1.58</td>
<td valign="top" align="center">80.56 &#x000B1; 2.08</td>
<td valign="top" align="center">84.88 &#x000B1; 1.73</td>
<td valign="top" align="center" style="background-color:#bbbdc0">88.51 &#x000B1; 2.21</td>
<td valign="top" align="center">88.28 &#x000B1; 2.00</td>
</tr>
<tr>
<td valign="top" align="left">B4</td>
<td valign="top" align="center">70.96 &#x000B1; 2.86</td>
<td valign="top" align="center">78.25 &#x000B1; 2.63</td>
<td valign="top" align="center">80.69 &#x000B1; 2.74</td>
<td valign="top" align="center">81.07 &#x000B1; 6.30</td>
<td valign="top" align="center" style="background-color:#bbbdc0">86.93 &#x000B1; 1.54</td>
</tr>
<tr>
<td valign="top" align="left">CORe50</td>
<td valign="top" align="center">85.58 &#x000B1; 3.02</td>
<td valign="top" align="center">98.73 &#x000B1; 0.97</td>
<td valign="top" align="center">99.19 &#x000B1; 0.46</td>
<td valign="top" align="center">97.26 &#x000B1; 1.76</td>
<td valign="top" align="center" style="background-color:#bbbdc0">99.33 &#x000B1; 0.30</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The gray shaded entries represent the highest values for each experiment</italic>.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>This table shows the training time (in minutes) and the file size (in megabytes) of each model in the B4 experiment.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>Training time (min)</bold></th>
<th valign="top" align="center"><bold>File size (in MB)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">GDM 6,000</td>
<td valign="top" align="center">129 &#x000B1; 25</td>
<td valign="top" align="center">39 &#x000B1; 4</td>
</tr>
<tr>
<td valign="top" align="left">GDM 12,000</td>
<td valign="top" align="center">363 &#x000B1; 58</td>
<td valign="top" align="center">146 &#x000B1; 10</td>
</tr>
<tr>
<td valign="top" align="left">GDM 28,000</td>
<td valign="top" align="center">1704 &#x000B1; 437</td>
<td valign="top" align="center">649 &#x000B1; 226</td>
</tr>
<tr>
<td valign="top" align="left">A-SOINN&#x0002B;</td>
<td valign="top" align="center" style="background-color:#bbbdc0">4 &#x000B1; 0.7</td>
<td valign="top" align="center" style="background-color:#bbbdc0">0.7 &#x000B1; 0.4</td>
</tr>
<tr>
<td valign="top" align="left">(ncc)A-SOINN&#x0002B;</td>
<td valign="top" align="center">370 &#x000B1; 22</td>
<td valign="top" align="center">741 &#x000B1; 77</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The values are averaged across the four repetitions and shown together with the corresponding standard deviations. The gray shaded areas represent the lowest values</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
</sec>
<sec sec-type="discussion" id="s5">
<title>5. Discussion</title>
<p>Although the A-SOINN&#x0002B; approach exhibits higher average forgetting curves, it reaches average accuracy values comparably high to the best GDM model. Additionally, it shows considerably lower memory requirements regarding the number of units. In the fourth experiment, the difference consists of approximately 28 times fewer units than the GDM 6,000 and approximately 100 times fewer than the GDM 28,000. Here, the maximum of each unit curve is used for comparison. In the first experiment (<xref ref-type="fig" rid="F5">Figure 5</xref>), the highest difference between the two approaches is observable. The A-SOINN&#x0002B; consists of around 100 and 350 times fewer nodes than the GDM 6,000 and GDM 28,000, respectively. Removing the node creation constraint can improve the performance of the A-SOINN&#x0002B;. However, it also increases the memory requirements since nodes are created even if the input has the same label as the prediction.</p>
<p>Nevertheless, the proposed model also shows drawbacks. It exhibits a generally higher standard deviation, especially after task eight in the v-NICO-World-LL experiments. Hence, the different category orders (repetitions) cause different performances. We call this effect <italic>category order effect</italic>. After analyzing each repetition of the first four experiments individually, we observe that especially R1 is causing poor results. However, sometimes R2 and R3 cause poor results as well. With a more in-depth analysis of the training procedures, we observe that as soon as the categories &#x0201C;pocket watch&#x0201D; and &#x0201C;present&#x0201D; become available, the performance decreases. In R1, R2, and R3, these objects become available at tasks 8 and 9, respectively. At the same time, the unit number for both categories increases. We further observe that the categories &#x0201C;present&#x0201D; and &#x0201C;pocket watch&#x0201D; are often misclassified as &#x0201C;ball&#x0201D; or &#x0201C;doughnut.&#x0201D; We assume that the categories&#x00027; intra-category dissimilarity is large due to high differences among the corresponding instances. This dissimilarity and the mentioned misclassification as &#x0201C;ball&#x0201D; or &#x0201C;doughnut&#x0201D; can explain the higher number of units. Due to the comparably high number of &#x0201C;present&#x0201D; units and its similarity to existing neurons, the task-dependent test sets of previous similar categories are more likely to be misclassified as a &#x0201C;present,&#x0201D; resulting in low accuracy and high forgetting. As we never train the model on previous samples again, the nodes remain unchanged, and the performance stays low. We further assume that this category order effect is amplified with an increasing background complexity, resulting in a higher standard deviation. It is assumed that the inter-category distance is larger in higher background complexities due to a more substantial influence of the background on the feature vectors. We hypothesize that this effect occurs due to the small number of neurons in the network and because already trained neurons are never retrained afterward. Therefore, we suggest two extensions for future work: (1) removing the creation constraint or (2) implementing a (pseudo-)rehearsal technique. The first allows the creation of more units for one category, which can lead to a higher probability of being predicted correctly. The second can retrain existing neurons to represent the corresponding category better. Our improvement suggestions are supported by the results of the (ncc)A-SOINN&#x0002B; and the GDM models. The former has no node creation constraint, and the latter use memory replay (pseudo-rehearsal), and in all four, the category order effect does not occur. However, the (ncc)A-SOINN&#x0002B; shows an infeasible unit number increase. Therefore, we recommend incorporating (pseudo-)rehearsal in the A-SOINN&#x0002B; for future work.</p>
<p>An essential aspect of LL is a high performance even after thousands or hundreds of thousands of tasks, including the memory requirements. In all five experiments, the unit number of the GDMs converges and increases linearly or stronger than linearly in the proposed model. In the CORe50 experiment, the A-SOINN&#x0002B; would become less efficient than the best GDM after approximately 30 times more tasks (around task 310), assuming a naive continuation of the unit number curves and a constant average accuracy after task 10. In the v-NICO-World-LL experiments, this would happen after 200 times more tasks (around task 2,000) at the earliest. However, due to the decreasing average accuracy behavior in the first four experiments, such a prediction is unreliable. It is much more likely that the GDM will increase its neuron number or drop in performance. On the other hand, the A-SOINN&#x0002B; can grow continually to learn a theoretically unlimited amount of tasks. However, this, in turn, can lead to infeasible memory requirements. Therefore, we recommend further investigations in this direction to evaluate the proposed model&#x00027;s long-term usability on a larger dataset, e.g., by expanding the v-NICO-World-LL with additional categories. Two further aspects of the dataset can be tackled in future work to create a higher real-world resemblance. The robot&#x00027;s arm movements are smooth to create sharp images, and jittery movements are not included. However, they take place in a real-world HRI scenario. Furthermore, the object&#x00027;s physical properties like weight or texture do not influence the arm movement in our scenario. Due to the adaptability of a virtual 3D scenario, jittery movements and physical object properties can be included in future work.</p>
<p>The results of our experiments demonstrate that for the two tested datasets, the pruning strategy of the original SOINN&#x0002B; approach (Wiwatcharakoses and Berrar, <xref ref-type="bibr" rid="B45">2019</xref>) and the adaptations made in this work lead to an efficient LL model for classification tasks in terms of average accuracy and memory requirements. The latter is observable in the unit number and in the file size of each model. We further observe that the training time of the A-SOINN&#x0002B; is approximately 268 times lower than the one of the GDM 2,8000 in the fourth experiment. We attribute this to the fact that for BMU determination, an iteration over the whole network is required, making the computation time grow linearly with the number of units. Although the A-SOINN&#x0002B; without creation constraint shows similar high unit numbers as the GDM 28,000, its training time is four times lower. We assume that GDM&#x00027;s memory replay is causing higher training time as additional iterations over the network are performed. These aspects indicate higher suitability of the A-SOINN&#x0002B; for autonomous robots with limited computational resources.</p>
</sec>
<sec id="s6">
<title>6. Conclusion and Future Work</title>
<sec>
<title>6.1. Conclusion</title>
<p>This work investigates whether the A-SOINN&#x0002B; approach can reach state-of-the-art classification accuracy results compared to the GDM architecture while showing fewer memory requirements in terms of created neurons. Two main adaptations are made compared to the original SOINN&#x0002B; (Wiwatcharakoses and Berrar, <xref ref-type="bibr" rid="B45">2019</xref>). (1) An associative matrix is used that stores for each node a frequency-based distribution of input labels to enable classification, and (2) top-down cues to regulate the structural plasticity of the network are introduced in the form of additional constraints for node creation and weight adaptation. The models are tested on two LL object recognition datasets, namely CORe50 and the v-NICO-World-LL. The latter is a novel LL dataset proposed in this work. This dataset exhibits three novel features that are, to our knowledge, currently not considered as a whole by other LL object recognition datasets: (1) four different background groups of different levels of complexity are defined, (2) a virtual robot is manipulating objects instead of a human to simulate a long-term HRI scenario where the robot receives different objects over time, and (3) it is recorded in a nearly photorealistic virtual environment making it highly controlled. The dataset consists of 100 objects belonging to 10 categories. These categories could also appear in an HRI scenario. The results of five experiments show that the A-SOINN&#x0002B; approach reaches high average accuracy results during LL and, in general, it is as accurate as the best GDM model. Furthermore, it requires fewer neurons than the other models, with at least approximately 30 and at most 350 times fewer units than the best GDM of each experiment. Additionally, its training time is approximately 268 times lower.</p>
<p>In summary, our contributions are:</p>
<list list-type="order">
<list-item><p>The A-SOINN&#x0002B; approach is developed, which is a novel and efficient version of an existing unsupervised LL approach for classification tasks.</p></list-item>
<list-item><p>A novel, nearly photorealistic LL object recognition dataset is created using a virtual humanoid robot. This dataset&#x00027;s main feature is the grouping of the environments into different levels of complexity. The dataset and methodology for synthetic data generation can be made available to the research community.</p></list-item>
</list>
</sec>
<sec>
<title>6.2. Future Work</title>
<p>Future work can tackle the long-term behavior of the A-SOINN&#x0002B;. In this context, the v-NICO-World-LL can be expanded with additional categories. A more in-depth analysis of different feature extractor architectures and data splits for transfer learning can also be examined in future works. Furthermore, it would be interesting to examine whether the proposed model can continuously learn from a multi-modal signal, like an audio-visual stream. Different aspects of a real-world HRI scenario are not considered in this work but can influence the performance of the A-SOINN&#x0002B; and the long-term HRI experience in general. Grasping objects is not an easy task for real robots, and incorrect grasping can cause the object to slip out of the robot&#x00027;s hand. Furthermore, depending on the hardware, cameras might create blurred images due to jittery robot arm movements. Future work can examine whether the A-SOINN&#x0002B; behaves differently on such images by applying it to a real robot. The results shown in this work indicate higher suitability of the A-SOINN&#x0002B; approach for autonomous robots with limited computational resources. However, whether this applies to thousands or hundreds of thousands of tasks remains an open question. Nevertheless, this work is a further step toward social robots that continually acquire knowledge through long-term human-robot interactions.</p>
</sec>
</sec>
<sec sec-type="data-availability-statement" id="s7">
<title>Data Availability Statement</title>
<p>The datasets presented in this article are not readily available because not all 3D models are licensed under CC BY. Requests to access the datasets should be directed to <email>a.logacjov&#x00040;gmail.com</email>. We would like to publish a partial dataset, including all objects licensed under CC BY. Once available it will be placed on <ext-link ext-link-type="uri" xlink:href="https://www.inf.uni-hamburg.de/en/inst/ab/wtm/research/corpora.html">https://www.inf.uni-hamburg.de/en/inst/ab/wtm/research/corpora.html</ext-link>.</p>
</sec>
<sec id="s8">
<title>Author Contributions</title>
<p>AL and MK conceived the presented idea. AL developed the neural architectures and conducted and evaluated the experiments with support from MK. AL was the primary contributor to the final version of the manuscript. SW and MK supervised the project and revised the manuscript. All authors provided critical feedback and helped to shape the research, analysis, and manuscript.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. The handling editor declared a shared affiliation with several of the authors AL, MK, and SW at time of review.</p>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Adel</surname> <given-names>T.</given-names></name> <name><surname>Zhao</surname> <given-names>H.</given-names></name> <name><surname>Turner</surname> <given-names>R. E.</given-names></name></person-group> (<year>2020</year>). <article-title>Continual learning with adaptive weights (CLAW)</article-title>. <source>arXiv:1911.09514 [cs, stat]</source>. arXiv: 1911.09514.</citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aljundi</surname> <given-names>R.</given-names></name> <name><surname>Babiloni</surname> <given-names>F.</given-names></name> <name><surname>Elhoseiny</surname> <given-names>M.</given-names></name> <name><surname>Rohrbach</surname> <given-names>M.</given-names></name> <name><surname>Tuytelaars</surname> <given-names>T.</given-names></name></person-group> (<year>2018</year>). <article-title>Memory aware synapses: learning what (not) to forget</article-title>. <source>arXiv:1711.09601 [cs, stat]</source>. arXiv: 1711.09601. <pub-id pub-id-type="doi">10.1007/978-3-030-01219-9_9</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Belouadah</surname> <given-names>E.</given-names></name> <name><surname>Popescu</surname> <given-names>A.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;IL2M: class incremental learning with dual memory,&#x0201D;</article-title> in <source>2019 IEEE/CVF International Conference on Computer Vision (ICCV)</source> (<publisher-loc>Seoul</publisher-loc>), <fpage>583</fpage>&#x02013;<lpage>592</lpage>.</citation></ref>
<ref id="B4">
<citation citation-type="web"><person-group person-group-type="author"><collab>Blender Foundation</collab></person-group> (<year>2020</year>). <source>Blender (v2.82.7)</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.blender.org">https://www.blender.org</ext-link>.</citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chaudhry</surname> <given-names>A.</given-names></name> <name><surname>Dokania</surname> <given-names>P. K.</given-names></name> <name><surname>Ajanthan</surname> <given-names>T.</given-names></name> <name><surname>Torr</surname> <given-names>P. H. S.</given-names></name></person-group> (<year>2018</year>). <article-title>Riemannian walk for incremental learning: understanding forgetting and intransigence</article-title>. <source>CoRR</source> abs/1801.10112. <pub-id pub-id-type="doi">10.1007/978-3-030-01252-6_33</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dautenhahn</surname> <given-names>K.</given-names></name></person-group> (<year>2004</year>). <article-title>Robots we like to live with?! a developmental perspective on a personalized, life-long robot companion</article-title>. <fpage>17</fpage>&#x02013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.1109/ROMAN.2004.1374720</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Douillard</surname> <given-names>A.</given-names></name> <name><surname>Cord</surname> <given-names>M.</given-names></name> <name><surname>Ollion</surname> <given-names>C.</given-names></name> <name><surname>Robert</surname> <given-names>T.</given-names></name> <name><surname>Valle</surname> <given-names>E.</given-names></name></person-group> (<year>2020</year>). <article-title>PODNet: pooled outputs distillation for small-tasks incremental learning</article-title>. <source>arXiv:2004.13513 [cs]</source>. arXiv: 2004.13513. <pub-id pub-id-type="doi">10.1007/978-3-030-58565-5_6</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Fernando</surname> <given-names>C.</given-names></name> <name><surname>Banarse</surname> <given-names>D.</given-names></name> <name><surname>Blundell</surname> <given-names>C.</given-names></name> <name><surname>Zwols</surname> <given-names>Y.</given-names></name> <name><surname>Ha</surname> <given-names>D.</given-names></name> <name><surname>Rusu</surname> <given-names>A. A.</given-names></name> <etal/></person-group>. (<year>2017</year>). <source>Pathnet: Evolution channels gradient descent in super neural networks. <italic>CoRR</italic> abs/1701.08734</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1701.08734">https://arxiv.org/abs/1701.08734</ext-link></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ghesmoune</surname> <given-names>M.</given-names></name> <name><surname>Lebbah</surname> <given-names>M.</given-names></name> <name><surname>Azzag</surname> <given-names>H.</given-names></name></person-group> (<year>2016</year>). <article-title>State-of-the-art on clustering data streams</article-title>. <source>Big Data Analyt.</source> <volume>1</volume>, <fpage>1</fpage>&#x02013;<lpage>27</lpage>. <pub-id pub-id-type="doi">10.1186/s41044-016-0011-3</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Oerlemans</surname> <given-names>A.</given-names></name> <name><surname>Lao</surname> <given-names>S.</given-names></name> <name><surname>Wu</surname> <given-names>S.</given-names></name> <name><surname>Lew</surname> <given-names>M. S.</given-names></name></person-group> (<year>2016</year>). <article-title>Deep learning for visual understanding: A review</article-title>. <source>Neurocomputing</source> <volume>187</volume>, <fpage>27</fpage>&#x02013;<lpage>48</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2015.09.116</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hou</surname> <given-names>S.</given-names></name> <name><surname>Pan</surname> <given-names>X.</given-names></name> <name><surname>Loy</surname> <given-names>C. C.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Lin</surname> <given-names>D.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Lifelong learning via progressive distillation and retrospection,&#x0201D;</article-title> in <source>Computer Vision &#x02013; ECCV 2018</source>, Lecture Notes in Computer Science, (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <fpage>452</fpage>&#x02013;<lpage>467</lpage>.</citation></ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hou</surname> <given-names>S.</given-names></name> <name><surname>Pan</surname> <given-names>X.</given-names></name> <name><surname>Loy</surname> <given-names>C. C.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Lin</surname> <given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Learning a unified classifier incrementally via rebalancing,&#x0201D;</article-title> in <source>2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Long Beach, CA</publisher-loc>), <fpage>831</fpage>&#x02013;<lpage>839</lpage>.</citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kerzel</surname> <given-names>M.</given-names></name> <name><surname>Pekarek-Rosin</surname> <given-names>T.</given-names></name> <name><surname>Strahl</surname> <given-names>E.</given-names></name> <name><surname>Heinrich</surname> <given-names>S.</given-names></name> <name><surname>Wermter</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Teaching nico how to grasp: an empirical study on crossmodal social interaction as a key factor for robots learning from humans</article-title>. <source>Front. Neurorobot.</source> <volume>14</volume>:<fpage>28</fpage>. <pub-id pub-id-type="doi">10.3389/fnbot.2020.00028</pub-id><pub-id pub-id-type="pmid">32581759</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kerzel</surname> <given-names>M.</given-names></name> <name><surname>Strahl</surname> <given-names>E.</given-names></name> <name><surname>Magg</surname> <given-names>S.</given-names></name> <name><surname>Navarro-Guerrero</surname> <given-names>N.</given-names></name> <name><surname>Heinrich</surname> <given-names>S.</given-names></name> <name><surname>Wermter</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Nico&#x02014;neuro-inspired companion: a developmental humanoid robot platform for multimodal interaction,&#x0201D;</article-title> in <source>2017 26th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN)</source> (<publisher-loc>Lisbon</publisher-loc>), <fpage>113</fpage>&#x02013;<lpage>120</lpage>.</citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kirkpatrick</surname> <given-names>J.</given-names></name> <name><surname>Pascanu</surname> <given-names>R.</given-names></name> <name><surname>Rabinowitz</surname> <given-names>N. C.</given-names></name> <name><surname>Veness</surname> <given-names>J.</given-names></name> <name><surname>Desjardins</surname> <given-names>G.</given-names></name> <name><surname>Rusu</surname> <given-names>A. A.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Overcoming catastrophic forgetting in neural networks</article-title>. <source>CoRR</source> abs/1612.00796. <pub-id pub-id-type="doi">10.1073/pnas.1611835114</pub-id><pub-id pub-id-type="pmid">28292907</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeCun</surname> <given-names>Y.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep learning</article-title>. <source>Nature</source> <volume>521</volume>, <fpage>436</fpage>&#x02013;<lpage>444</lpage>. <pub-id pub-id-type="doi">10.1038/nature14539</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leite</surname> <given-names>I.</given-names></name> <name><surname>Martinho</surname> <given-names>C.</given-names></name> <name><surname>Paiva</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Social robots for long-term interaction: a survey</article-title>. <source>Int. J. Soc. Robot.</source> <volume>5</volume>, <fpage>291</fpage>&#x02013;<lpage>308</lpage>. <pub-id pub-id-type="doi">10.1007/s12369-013-0178-y</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lesort</surname> <given-names>T.</given-names></name> <name><surname>Lomonaco</surname> <given-names>V.</given-names></name> <name><surname>Stoian</surname> <given-names>A.</given-names></name> <name><surname>Maltoni</surname> <given-names>D.</given-names></name> <name><surname>Filliat</surname> <given-names>D.</given-names></name> <name><surname>Diaz Rodriguez</surname> <given-names>N.</given-names></name></person-group> (<year>2020</year>). <article-title>Continual learning for robotics: Definition, framework, learning strategies, opportunities and challenges</article-title>. <source>Inform. Fusion</source> <volume>58</volume>, <fpage>52</fpage>&#x02013;<lpage>68</lpage>. <pub-id pub-id-type="doi">10.1016/j.inffus.2019.12.004</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leys</surname> <given-names>C.</given-names></name> <name><surname>Ley</surname> <given-names>C.</given-names></name> <name><surname>Klein</surname> <given-names>O.</given-names></name> <name><surname>Bernard</surname> <given-names>P.</given-names></name> <name><surname>Licata</surname> <given-names>L.</given-names></name></person-group> (<year>2013</year>). <article-title>Detecting outliers: Do not use standard deviation around the mean, use absolute deviation around the median</article-title>. <source>J. Exp. Soc. Psychol.</source> <volume>49</volume>, <fpage>764</fpage>&#x02013;<lpage>766</lpage>. <pub-id pub-id-type="doi">10.1016/j.jesp.2013.03.013</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Hoiem</surname> <given-names>D.</given-names></name></person-group> (<year>2018</year>). <article-title>Learning without forgetting</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>40</volume>, <fpage>2935</fpage>&#x02013;<lpage>2947</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2017.2773081</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liew</surname> <given-names>W. S.</given-names></name> <name><surname>Kiong Loo</surname> <given-names>C.</given-names></name> <name><surname>Gryshchuk</surname> <given-names>V.</given-names></name> <name><surname>Weber</surname> <given-names>C.</given-names></name> <name><surname>Wermter</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;Effect of pruning on catastrophic forgetting in growing dual memory networks,&#x0201D;</article-title> in <source>2019 International Joint Conference on Neural Networks (IJCNN)</source> (<publisher-loc>Hungary</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</citation></ref>
<ref id="B22">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Lomonaco</surname> <given-names>V.</given-names></name> <name><surname>Maltoni</surname> <given-names>D.</given-names></name></person-group> (<year>2017</year>). <source>Core50: a new dataset and benchmark for continuous object recognition. <italic>CoRR</italic> abs/1705.03550</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1705.03550">https://arxiv.org/abs/1705.03550</ext-link></citation></ref>
<ref id="B23">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Lomonaco</surname> <given-names>V.</given-names></name> <name><surname>Maltoni</surname> <given-names>D.</given-names></name> <name><surname>Pellegrini</surname> <given-names>L.</given-names></name></person-group> (<year>2019</year>). <source>Fine-grained continual learning. <italic>CoRR</italic> abs/1907.03799</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1907.03799">https://arxiv.org/abs/1907.03799</ext-link></citation></ref>
<ref id="B24">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Lopez-Paz</surname> <given-names>D.</given-names></name> <name><surname>Ranzato</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <source>Gradient episodic memory for continual learning. <italic>CoRR</italic> abs/1706.08840</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1706.08840">https://arxiv.org/abs/1706.08840</ext-link></citation></ref>
<ref id="B25">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mallya</surname> <given-names>A.</given-names></name> <name><surname>Davis</surname> <given-names>D.</given-names></name> <name><surname>Lazebnik</surname> <given-names>S.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Piggyback: adapting a single network to multiple tasks by learning to mask weights,&#x0201D;</article-title> in <source>Computer Vision&#x02013;ECCV 2018</source>, eds V. Ferrari, M. Hebert, C. Sminchisescu, and Y. Weiss (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <fpage>72</fpage>&#x02013;<lpage>88</lpage>.</citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mallya</surname> <given-names>A.</given-names></name> <name><surname>Lazebnik</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>Packnet: adding multiple tasks to a single network by iterative pruning</article-title>. <source>CoRR</source> abs/1711.05769. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00810</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marsland</surname> <given-names>S.</given-names></name> <name><surname>Shapiro</surname> <given-names>J.</given-names></name> <name><surname>Nehmzow</surname> <given-names>U.</given-names></name></person-group> (<year>2002</year>). <article-title>A self-organising network that grows when required</article-title>. <source>Neural Netw.</source> <volume>15</volume>, <fpage>1041</fpage>&#x02013;<lpage>1058</lpage>. <pub-id pub-id-type="doi">10.1016/S0893-6080(02)00078-3</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Metta</surname> <given-names>G.</given-names></name> <name><surname>Sandini</surname> <given-names>G.</given-names></name> <name><surname>Vernon</surname> <given-names>D.</given-names></name> <name><surname>Natale</surname> <given-names>L.</given-names></name> <name><surname>Nori</surname> <given-names>F.</given-names></name></person-group> (<year>2008</year>). <article-title>&#x0201C;The icub humanoid robot: an open platform for research in embodied cognition,&#x0201D;</article-title> in <source>Proceedings of the 8th Workshop on Performance Metrics for Intelligent Systems</source>, PerMIS &#x00027;08 (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Association for Computing Machinery0</publisher-name>, <fpage>50</fpage>&#x02013;<lpage>56</lpage>.</citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Najafabadi</surname> <given-names>M.</given-names></name> <name><surname>Villanustre</surname> <given-names>F.</given-names></name> <name><surname>Khoshgoftaar</surname> <given-names>T.</given-names></name> <name><surname>Seliya</surname> <given-names>N.</given-names></name> <name><surname>Wald</surname> <given-names>R.</given-names></name> <name><surname>Muharemagic</surname> <given-names>E.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep learning applications and challenges in big data analytics</article-title>. <source>J. Big Data</source> <volume>2</volume>, <fpage>1</fpage>&#x02013;<lpage>21</lpage>. <pub-id pub-id-type="doi">10.1186/s40537-014-0007-7</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ostapenko</surname> <given-names>O.</given-names></name> <name><surname>Puscas</surname> <given-names>M.</given-names></name> <name><surname>Klein</surname> <given-names>T.</given-names></name> <name><surname>J&#x000E4;hnichen</surname> <given-names>P.</given-names></name> <name><surname>Nabi</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <article-title>Learning to remember: a synaptic plasticity driven framework for continual learning</article-title>. <source>arXiv:1904.03137 [cs]</source>. arXiv: 1904.03137. <pub-id pub-id-type="doi">10.1109/CVPR.2019.01158</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Parisi</surname> <given-names>G.</given-names></name> <name><surname>Tani</surname> <given-names>J.</given-names></name> <name><surname>Weber</surname> <given-names>C.</given-names></name> <name><surname>Wermter</surname> <given-names>S.</given-names></name></person-group> (<year>2018</year>). <article-title>Lifelong learning of spatiotemporal representations with dual-memory recurrent self-organization</article-title>. <source>Front. Neurorobot.</source> <volume>12</volume>:<fpage>78</fpage>. <pub-id pub-id-type="doi">10.3389/fnbot.2018.00078</pub-id><pub-id pub-id-type="pmid">30546302</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Parisi</surname> <given-names>G. I.</given-names></name> <name><surname>Kemker</surname> <given-names>R.</given-names></name> <name><surname>Part</surname> <given-names>J. L.</given-names></name> <name><surname>Kanan</surname> <given-names>C.</given-names></name> <name><surname>Wermter</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>Continual lifelong learning with neural networks: a review</article-title>. <source>Neural Netw.</source> <volume>113</volume>, <fpage>54</fpage>&#x02013;<lpage>71</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2019.01.012</pub-id><pub-id pub-id-type="pmid">30780045</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Parisi</surname> <given-names>G. I.</given-names></name> <name><surname>Tani</surname> <given-names>J.</given-names></name> <name><surname>Weber</surname> <given-names>C.</given-names></name> <name><surname>Wermter</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>Lifelong learning of human actions with deep neural network self-organization</article-title>. <source>Neural Netw.</source> <volume>96</volume>, <fpage>137</fpage>&#x02013;<lpage>149</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2017.09.001</pub-id><pub-id pub-id-type="pmid">29017140</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pasquale</surname> <given-names>G.</given-names></name> <name><surname>Ciliberto</surname> <given-names>C.</given-names></name> <name><surname>Odone</surname> <given-names>F.</given-names></name> <name><surname>Rosasco</surname> <given-names>L.</given-names></name> <name><surname>Natale</surname> <given-names>L.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Teaching icub to recognize objects using deep convolutional neural networks,&#x0201D;</article-title> in <source>Proceedings of the 4th International Conference on Machine Learning for Interactive Systems</source>, MLIS&#x00027;15, <volume>Vol. 43</volume>, (<publisher-loc>Lille</publisher-loc>: <publisher-name>JMLR.org</publisher-name>), <fpage>21</fpage>&#x02013;<lpage>25</lpage>.</citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pasquale</surname> <given-names>G.</given-names></name> <name><surname>Ciliberto</surname> <given-names>C.</given-names></name> <name><surname>Odone</surname> <given-names>F.</given-names></name> <name><surname>Rosasco</surname> <given-names>L.</given-names></name> <name><surname>Natale</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). <article-title>Are we done with object recognition? the icub robot&#x00027;s perspective</article-title>. <source>CoRR</source> Lille. abs/1709.09882. <pub-id pub-id-type="doi">10.1016/j.robot.2018.11.001</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pasquale</surname> <given-names>G.</given-names></name> <name><surname>Ciliberto</surname> <given-names>C.</given-names></name> <name><surname>Rosasco</surname> <given-names>L.</given-names></name> <name><surname>Natale</surname> <given-names>L.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Object identification from few examples by improving the invariance of a deep convolutional neural network,&#x0201D;</article-title> in <source>2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)</source> (<publisher-loc>Daejeon</publisher-loc>), <fpage>4904</fpage>&#x02013;<lpage>4911</lpage>.</citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rebuffi</surname> <given-names>S.</given-names></name> <name><surname>Kolesnikov</surname> <given-names>A.</given-names></name> <name><surname>Lampert</surname> <given-names>C. H.</given-names></name></person-group> (<year>2016</year>). <article-title>icarl: incremental classifier and representation learning</article-title>. <source>CoRR</source> abs/1611.07725. <pub-id pub-id-type="doi">10.1109/CVPR.2017.587</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Russakovsky</surname> <given-names>O.</given-names></name> <name><surname>Deng</surname> <given-names>J.</given-names></name> <name><surname>Su</surname> <given-names>H.</given-names></name> <name><surname>Krause</surname> <given-names>J.</given-names></name> <name><surname>Satheesh</surname> <given-names>S.</given-names></name> <name><surname>Ma</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Imagenet large scale visual recognition challenge</article-title>. <source>Int. J. Comput. Vis.</source> <volume>115</volume>, <fpage>1</fpage>&#x02013;<lpage>42</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-015-0816-y</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Rusu</surname> <given-names>A. A.</given-names></name> <name><surname>Rabinowitz</surname> <given-names>N. C.</given-names></name> <name><surname>Desjardins</surname> <given-names>G.</given-names></name> <name><surname>Soyer</surname> <given-names>H.</given-names></name> <name><surname>Kirkpatrick</surname> <given-names>J.</given-names></name> <name><surname>Kavukcuoglu</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2016</year>). <source>Progressive neural networks. <italic>CoRR</italic> abs/1606.04671</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1606.04671">http://arxiv.org/abs/1606.04671</ext-link></citation></ref>
<ref id="B40">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Serr&#x000E0;</surname> <given-names>J.</given-names></name> <name><surname>Sur&#x000ED;s</surname> <given-names>D.</given-names></name> <name><surname>Miron</surname> <given-names>M.</given-names></name> <name><surname>Karatzoglou</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <source>Overcoming catastrophic forgetting with hard attention to the task. <italic>CoRR</italic> abs/1801.01423</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1801.01423">https://arxiv.org/abs/1801.01423</ext-link></citation></ref>
<ref id="B41">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Shin</surname> <given-names>H.</given-names></name> <name><surname>Lee</surname> <given-names>J. K.</given-names></name> <name><surname>Kim</surname> <given-names>J.</given-names></name> <name><surname>Kim</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <source>Continual learning with deep generative replay. <italic>CoRR</italic> abs/1705.08690</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1705.08690">https://arxiv.org/abs/1705.08690</ext-link></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Simonyan</surname> <given-names>K.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Very deep convolutional networks for large-scale image recognition</article-title>. <source>arXiv 1409.1556</source>.</citation>
</ref>
<ref id="B43">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Stojanov</surname> <given-names>S.</given-names></name> <name><surname>Mishra</surname> <given-names>S.</given-names></name> <name><surname>Thai</surname> <given-names>N. A.</given-names></name> <name><surname>Dhanda</surname> <given-names>N.</given-names></name> <name><surname>Humayun</surname> <given-names>A.</given-names></name> <name><surname>Yu</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>&#x0201C;Incremental object learning from contiguous views,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Long Beach, CA</publisher-loc>).</citation></ref>
<ref id="B44">
<citation citation-type="book"><person-group person-group-type="author"><collab>Upton G. and Cook, I.</collab></person-group> (<year>1996</year>). <source>Understanding Statistics</source>. <publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>OUP Oxford, 55</publisher-name>.</citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wiwatcharakoses</surname> <given-names>C.</given-names></name> <name><surname>Berrar</surname> <given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>Soinn&#x0002B;, a self-organizing incremental neural network for unsupervised learning from noisy data streams</article-title>. <source>Exp. Syst. Appl.</source> <fpage>143</fpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2019.113069</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>Y.</given-names></name> <name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Ye</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Guo</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Large Scale Incremental Learning</article-title>. <source>arXiv:1905.13260 [cs]</source>. arXiv: 1905.13260 version: 1. <pub-id pub-id-type="doi">10.1109/CVPR.2019.00046</pub-id></citation></ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yoon</surname> <given-names>J.</given-names></name> <name><surname>Yang</surname> <given-names>E.</given-names></name> <name><surname>Lee</surname> <given-names>J.</given-names></name> <name><surname>Hwang</surname> <given-names>S. J.</given-names></name></person-group> (<year>2018</year>). <article-title>Lifelong learning with dynamically expandable networks</article-title>. <source>arXiv:1708.01547 [cs]</source>. arXiv: 1708.01547.</citation></ref>
<ref id="B48">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Zenke</surname> <given-names>F.</given-names></name> <name><surname>Poole</surname> <given-names>B.</given-names></name> <name><surname>Ganguli</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <source>Improved multitask learning through synaptic intelligence. <italic>CoRR</italic> abs/1703.04200</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1703.04200v1">https://arxiv.org/abs/1703.04200v1</ext-link></citation></ref>
</ref-list>
<fn-group>
<fn id="fn0001"><p><sup>1</sup>The Algorithm and all subroutines are created by following the explanations and notations of Wiwatcharakoses and Berrar (<xref ref-type="bibr" rid="B45">2019</xref>) and Parisi et al. (<xref ref-type="bibr" rid="B33">2017</xref>).</p></fn>
<fn id="fn0002"><p><sup>2</sup>The dataset is available at <ext-link ext-link-type="uri" xlink:href="https://www.inf.uni-hamburg.de/en/inst/ab/wtm/research/corpora.html">https://www.inf.uni-hamburg.de/en/inst/ab/wtm/research/corpora.html</ext-link>.</p></fn>
</fn-group>
<fn-group>
<fn fn-type="financial-disclosure"><p><bold>Funding.</bold> We gratefully acknowledge partial support from the German Research Foundation DFG under project CML (TRR 169).</p>
</fn>
</fn-group>
</back>
</article>