<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Phys.</journal-id>
<journal-title>Frontiers in Physics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Phys.</abbrev-journal-title>
<issn pub-type="epub">2296-424X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">863291</article-id>
<article-id pub-id-type="doi">10.3389/fphy.2022.863291</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Physics</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Neural Network Model Based on the Tensor Network for Audio Tagging of Domestic Activities</article-title>
<alt-title alt-title-type="left-running-head">Yang et al.</alt-title>
<alt-title alt-title-type="right-running-head">Neural Network Model for Tagging</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Yang</surname>
<given-names>LiDong</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1260742/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Yue</surname>
<given-names>RenBo</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1621518/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Wang</surname>
<given-names>Jing</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1663260/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Liu</surname>
<given-names>Min</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>School of Information Engineering</institution>, <institution>Inner Mongolia University of Science and Technology</institution>, <addr-line>Baotou</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>School of Information and Electronics</institution>, <institution>Beijing Institute of Technology</institution>, <addr-line>Beijing</addr-line>, <country>China</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>China Mobile Research Institute</institution>, <addr-line>Beijing</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/966698/overview">Shengchen Li</ext-link>, Xi&#x2019;an Jiaotong-Liverpool University, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1183003/overview">Qiuqiang Kong</ext-link>, University of Surrey, United Kingdom</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1674762/overview">Xi Shao</ext-link>, Nanjing University of Posts and Telecommunications, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Jing Wang, <email>wangjing@bit.edu.cn</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Interdisciplinary Physics, a section of the journal Frontiers in Physics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>12</day>
<month>04</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>10</volume>
<elocation-id>863291</elocation-id>
<history>
<date date-type="received">
<day>27</day>
<month>01</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>15</day>
<month>03</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Yang, Yue, Wang and Liu.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Yang, Yue, Wang and Liu</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Due to the serious problem of population aging, monitoring of domestic activities is increasingly important. Audio tagging of domestic activities is very suitable when the visual data are unavailable due to the interference from light and the environment. Aiming at solving this problem, a neural network model based on the tensor network is proposed for audio tagging of domestic activities that is more interpretable than traditional neural networks. The introduction of the tensor network can compress the network parameters and reduce the redundancy of the training model while maintaining a good performance. First, the important features of a Mel spectrogram of the input audio are extracted through the convolutional neural networks (CNNs). Then, they are converted into the high-order space corresponding with the tensor network. The spatial structure information and important features can be further extracted and retained through the matrix product state (MPS). Large patches of the featured data are divided into small local orderless patches when using the tensor network. The final tagging results are obtained through the MPS layers which is just a tensor network structure based on the tensor train decomposition. In order to evaluate the proposed method, the DCASE 2018 challenge task 5 dataset for monitoring domestic activities is selected. The results showed that the average F1-score of the proposed model in the test set of the development dataset and validation dataset reached 87.7 and 85.9%, which are 3.2 and 2.8% higher than the baseline system, respectively. It is verified that the proposed model can perform better and more efficiently for audio tagging of domestic activities.</p>
</abstract>
<kwd-group>
<kwd>tensor network</kwd>
<kwd>matrix product state (MPS)</kwd>
<kwd>tensor train decomposition</kwd>
<kwd>audio tagging</kwd>
<kwd>neural network</kwd>
</kwd-group>
<contract-num rid="cn001">62071039 62161040</contract-num>
<contract-num rid="cn002">2021MS06030</contract-num>
<contract-num rid="cn003">2021GG0023</contract-num>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content>
</contract-sponsor>
<contract-sponsor id="cn002">Natural Science Foundation of Inner Mongolia<named-content content-type="fundref-id">10.13039/501100004763</named-content>
</contract-sponsor>
<contract-sponsor id="cn003">Science and Technology Major Project of Inner Mongolia<named-content content-type="fundref-id">10.13039/501100020788</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>The world is facing the problem of population aging. It is estimated that by 2050, the number of people over 64&#xa0;years will exceed 20% of the world&#x2019;s population. According to the survey, 40% of the elderly will live alone at home [<xref ref-type="bibr" rid="B1">1</xref>]. This will lead to many social problems, such as the increase in diseases and healthcare costs, the shortage of nursing staff, and the increase in the number of people unable to live independently. Therefore, it is imperative to develop ambient intelligence-assisted living tools to help the elderly live independently at home [<xref ref-type="bibr" rid="B2">2</xref>]. The first task is to detect what is happening at home. Audio tagging is very suitable when the visual data are unavailable due to the interference from light and the environment. Audio tagging associate tags with the audio and identifies the events that generate the audio. Audio tagging of domestic activities has important applications in smart home robots, monitoring of domestic activities, and the lives of the elderly [<xref ref-type="bibr" rid="B3">3</xref>]. For the problem of audio tagging, Gong [<xref ref-type="bibr" rid="B4">4</xref>] proposed PSLA, a collection of model-agnostic training techniques. It includes ImageNet pre-training, balanced sampling, data augmentation, label augmentation, and model aggregation. The results we obtained outstripped the best previous systems. Puy [<xref ref-type="bibr" rid="B5">5</xref>] proposed a model based on separable convolutions, which uses separable convolutions in channel, time, and frequency dimensions to control the complexity of the network and achieved good results in terms of effect and complexity. The widely used dataset for audio signals is DCASE (Detection and Classification of Acoustic Scenes and Events). DCASE 2018 Challenge Task 5 [<xref ref-type="bibr" rid="B6">6</xref>] is specifically used for audio tagging for domestic activities. This tagging task provides the development and validation datasets and baseline system and requires identifying nine classes of events in domestic activities within 10-s clips. The audio data are collected by four linearly arranged microphones. There are many ways to process microphone array audio, among which Wang [<xref ref-type="bibr" rid="B7">7</xref>] proposed a modeling method that uses the channel mode, time mode, and frequency mode as the three dimensions to construct a three-dimensional tensor space, which has achieved good results. In the tensor completion method proposed by Yang [<xref ref-type="bibr" rid="B8">8</xref>], tensor modeling of multi-channel audio signals with the missing data has achieved good results.</p>
<p>Among the submitted systems in DCASE 2018 Challenge Task 5, the baseline system of this task trains a single classifier model that takes a single channel as the input. The learner in the baseline system is based on a neural network architecture using convolutional and dense layers. As input, log Mel-band energies are provided to the network for each microphone channel separately [<xref ref-type="bibr" rid="B9">9</xref>]. Inoue [<xref ref-type="bibr" rid="B10">10</xref>] put forward a combination method of a data-enhanced front-end module and a back-end module based on the CNN classification method. First, it enhances the input data by shuffling and mixing the sound clips. Its data enhancement method helped increase the variation of training samples and reduce the impact of unbalanced datasets. Then, the input of the CNN, as a classifier, is the log-Mel spectrogram of the enhanced data. The system proposed by Tanabe [<xref ref-type="bibr" rid="B11">11</xref>] is a combination of the front-end modules based on blind signal processing and the back-end modules based on machine learning. The front-end modules employ blind dereverberation and blind source separation. They use spatial cues without machine learning to avoid overfitting. The back-end modules employ one-dimensional convolutional neural network (1DCNN)-based architecture and VGG16-based architecture for the individual front-end modules. All of the probability outputs are ensembled. In addition, through mix-up-based data augmentation, overfitting is avoided in the back-end modules. TC2DCNN [<xref ref-type="bibr" rid="B12">12</xref>] is extended by operating the convolutions along the two dimensions of time and channel, not along the frequency axis, since similar patterns in different frequency bands do not necessarily belong to the similar audio event. INRC_2D [<xref ref-type="bibr" rid="B13">13</xref>] combines a deep neural network with a scattering transform. Each audio segment is first represented by two layers of scattering transform. The four denoised transforms of each of the two layers are combined together. Each of the fused layers is processed in parallel by two neural network (NN) architectures, RESNET, and a long short-term memory (LSTM) network, with a joint fully connected layer. The VGGish model proposed by Kong [<xref ref-type="bibr" rid="B14">14</xref>], which has an AlexNetish 8-layer CNN with global max pooling, has achieved good results.</p>
<p>The tensor network is a sparse data structure designed for the efficient representation and manipulation of the ultra-high dimensional data to achieve better interpretability of the data. It is similar to the kernel method in machine learning [<xref ref-type="bibr" rid="B15">15</xref>]. Through feature mapping, the original linearly inseparable data are converted to a high-dimensional space. In this space, a hyperplane can be linearly separable. But the number of parameters will be very large. Tensor train decomposition (also called the matrix product state) is a kind of tensor decomposition specifically for high-dimensional data. Wang [<xref ref-type="bibr" rid="B16">16</xref>] uses tensor train decomposition in a compressed HRTF, which is closer to the original HRTF than other methods. Therefore, tensor train decomposition is used to approximate the tensor networks. Matrix product state is the first tensor network to be discovered and used, which can be efficiently used in the simulation of the ground state of an infinite one-dimensional system. In recent years, tensor networks based on matrix product states have shown good performance in classification. For example, Stoudenmire [<xref ref-type="bibr" rid="B17">17</xref>] encoded the MNIST data into a tensor network, and the tensor network was trained to obtain the probability of each class to complete the classification. Efthymiou [<xref ref-type="bibr" rid="B18">18</xref>] proposed a new contraction method for Fashion-MNIST, which realizes the parallel compression of the horizontal edges, and then the vertical compression, which further accelerates the training speed. Selvan [<xref ref-type="bibr" rid="B19">19</xref>] proposed a lonet tensor network, which overcomes the shortcomings of the MPS tensor network, that is, the loss of spatial correlation when used for large resolutions. It is used for the two-dimensional classification of medical images and has achieved good results. While achieving good results, compared with other models, the GPU usage is significantly lower than that of the other models. PEPS [<xref ref-type="bibr" rid="B20">20</xref>] is a two-dimensional extension of the matrix product state. Although it has achieved great success, its algorithmic complexity is much higher than that of the matrix product state. MERA [<xref ref-type="bibr" rid="B20">20</xref>] is an experimental state of the ground state of a one-dimensional quantum system, which is inherently scale-invariant. In the MERA, tensors are connected to reproduce the holographic geometry. There are also other kinds of tensor network structures which have higher complexity than the MPS and can be used in other applications such as applied mathematics, chemistry, physics, machine learning, and many other fields.</p>
<p>In the article, a neural network model based on the tensor network is proposed for audio tagging of domestic activities. This article draws on the research results of the simplest and most mature matrix product state in the tensor network, hoping to achieve a balance between the complexity and effectiveness of the network model. An end-to-end tensor network-based neural network model is constructed and trained with the Mel spectrograms. After going through the convolutional layers, important features are extracted. Then, the MPS tensor network further extracts the features and gives the tagging results. This can not only achieve good tagging results but also compresses the network through tensor train decomposition, which has a smaller number of parameters than the traditional CNN. The F1-score is used to evaluate the performance of the proposed method. In terms of tagging performance, the performance of the proposed model is compared with other models. Compared with the results of the development dataset and the validation dataset of DCASE 2018 challenge task 5, the proposed method achieved better results. This article is a beneficial attempt to combine the tensor networks and neural networks and can also be extended to other deep learning sound signal processing fields.</p>
<p>The rest of this article is organized as follows: <xref ref-type="sec" rid="s2">Section 2</xref> introduces the neural network model based on the tensor network proposed in this article in detail. <xref ref-type="sec" rid="s3">Section 3</xref> introduces the parameter settings and experimental results of the proposed method, which are analyzed in terms of precision, recall, and F1-score, respectively. This article is concluded in <xref ref-type="sec" rid="s4">Section 4</xref>.</p>
</sec>
<sec id="s2">
<title>2 Neural Network Model Based on Tensor Network</title>
<p>As the experimental flowchart shows in <xref ref-type="fig" rid="F1">Figure 1</xref>, the proposed audio tagging method consists of three main stages, namely, data preprocessing, data augmentation, and neural network model based on the tensor network. Data preprocessing first performs channel fusion [<xref ref-type="bibr" rid="B21">21</xref>] on the audio, then takes the log after FFT, and then the Mel spectrogram is obtained by mapping the Mel frequency.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Flow chart of audio tagging.</p>
</caption>
<graphic xlink:href="fphy-10-863291-g001.tif"/>
</fig>
<p>The structure of the neural network model based on the tensor network is shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. Convolutional layers are used for extracting deeper feature representations. Important spatial structure and time information will be retained in the middle MPS layers. Finally, the retained information enters the MPS decision layer after being flattened to obtain the audio tagging results.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Structure of the neural network model based on the tensor network.</p>
</caption>
<graphic xlink:href="fphy-10-863291-g002.tif"/>
</fig>
<sec id="s2-1">
<title>2.1 Data Preprocessing and Augmentation</title>
<p>The Mel spectrogram as the audio feature of the original signal is used in the proposed method. The Mel spectrogram converts the ordinary frequency scale of the spectrogram into the Mel frequency scale. After framing, the fbank feature is extracted through the Mel filter bank [<xref ref-type="bibr" rid="B22">22</xref>]. The energy value distribution range is summarized and is then linearly corresponded to blue-yellow [<xref ref-type="bibr" rid="B23">23</xref>]. In this article, 128 triangular filters are used to form a Mel filter bank, which corresponds to the objective law that the higher the frequency, the duller the human ear is.</p>
<p>Data augmentation uses horizontal flip, vertical flip, and random rotation to enlarge the training data, avoid overfitting, and enhance the robustness of the model.</p>
</sec>
<sec id="s2-2">
<title>2.2 Neural Network Model Based on the Tensor Network</title>
<sec id="s2-2-1">
<title>2.2.1 CNN Feature Extraction</title>
<p>CNN [<xref ref-type="bibr" rid="B24">24</xref>] is used to process the multi-dimensional data, such as the two-dimensional images with many channels. CNN uses shared weights, local connections, pooling, and other layers to organize the attributes of natural signals. The convolutional layer, ReLU layer, and pooling layer are the most commonly used CNN layers.</p>
<p>The basic purpose of the convolutional layer is to determine the local connections between the features and map their information to a specific feature map. The convolution of the input <inline-formula id="inf1">
<mml:math id="m1">
<mml:mi>I</mml:mi>
</mml:math>
</inline-formula> with filter <inline-formula id="inf2">
<mml:math id="m2">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi mathvariant="bold">R</mml:mi>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>2</mml:mn>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is given as follows:<disp-formula id="e1">
<mml:math id="m3">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>I</mml:mi>
<mml:mi mathvariant="normal">&#x2217;</mml:mi>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>m</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:mstyle>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>where <inline-formula id="inf3">
<mml:math id="m4">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf4">
<mml:math id="m5">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> determine the size of the convolution kernel along the <italic>x</italic> and <italic>y</italic> directions. ReLU <inline-formula id="inf5">
<mml:math id="m6">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>z</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>max</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>z</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> [<xref ref-type="bibr" rid="B25">25</xref>] is a non-linear function which is applied to feature mapping created by the convolutional layer. The BN [<xref ref-type="bibr" rid="B26">26</xref>] layer normalizes each mini batch throughout the entire network, reducing the internal covariate shift caused by the progressive transforms. The BN layer is used to reduce the training time of the CNN and the sensitivity of network initialization. Therefore, this layer is used for normalization in the proposed network model.</p>
</sec>
<sec id="s2-2-2">
<title>2.2.2MPS Tensor Network</title>
<p>The tensor network notation is a brief graphical representation of the high-dimensional tensors. It not only makes it easier and more intuitive to process the high-dimensional tensors but also provides an insight into how to achieve more efficient operations. For a more comprehensive introduction to the tensor networks, references in [<xref ref-type="bibr" rid="B27">27</xref>]can be referred.</p>
<p>The MPS (matrix product state) [<xref ref-type="bibr" rid="B17">17</xref>, <xref ref-type="bibr" rid="B18">18</xref>] is a one-dimensional tensor network structure, which is based on tensor train decomposition [<xref ref-type="bibr" rid="B28">28</xref>]. It uses chain-connected small tensors to represent the high-dimensional tensors.</p>
<p>For a neural network model based on the tensor network, the generated Mel spectrograms must first be mapped to the high-dimensional space corresponding to the tensor network. According to <xref ref-type="disp-formula" rid="e2">Eq. 2</xref>, each pixel of the Mel spectrogram is mapped to a two-dimensional space.<disp-formula id="e2">
<mml:math id="m7">
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo>&#x232a;</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>cos</mml:mi>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mi>&#x3c0;</mml:mi>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo>&#x232a;</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>sin</mml:mi>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mi>&#x3c0;</mml:mi>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x232a;</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>where <inline-formula id="inf6">
<mml:math id="m8">
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mo>&#x232a;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the Dirac symbol in physics, representing the state vector. <inline-formula id="inf7">
<mml:math id="m9">
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo>&#x232a;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> means blue with low energy, and <inline-formula id="inf8">
<mml:math id="m10">
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x232a;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> means yellow with high energy, where <inline-formula id="inf9">
<mml:math id="m11">
<mml:mi>l</mml:mi>
</mml:math>
</inline-formula> represents the order of the Mel spectrogram, and <inline-formula id="inf10">
<mml:math id="m12">
<mml:mi>n</mml:mi>
</mml:math>
</inline-formula> represents the pixel order in the Mel spectrogram. The function with cos(&#x3c0;x/2) and sin(&#x3c0;x/2) is one of the mapping methods. After inputting the spectrogram, the data of each pixel are normalized to be between 0 and 1; using cos(&#x3c0;x/2) and sin(&#x3c0;x/2) can accurately represent the information in the pixel. After mapping, <inline-formula id="inf11">
<mml:math id="m13">
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo>&#x232a;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> can represent all the magnitudes of energy in the Mel spectrogram. After all the pixels are mapped, Mel spectrograms can be expressed as <xref ref-type="disp-formula" rid="e3">Eq. 3</xref> and also be expressed as <xref ref-type="disp-formula" rid="e4">Eq. 4</xref> using the tensor network notation.<disp-formula id="e3">
<mml:math id="m14">
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi>X</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo>&#x232a;</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo>&#x232a;</mml:mo>
</mml:mrow>
<mml:mo>&#x2297;</mml:mo>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
<mml:mrow>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo>&#x232a;</mml:mo>
</mml:mrow>
<mml:mo>&#x2297;</mml:mo>
<mml:mn>...</mml:mn>
<mml:mo>&#x2297;</mml:mo>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo>&#x232a;</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>
<disp-formula id="e4">
<mml:math id="m15">
<mml:mrow>
<mml:mi mathvariant="bold">&#x3a6;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold">&#x3d5;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2297;</mml:mo>
<mml:mi mathvariant="bold">&#x3d5;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2297;</mml:mo>
<mml:mn>...</mml:mn>
<mml:mo>&#x2297;</mml:mo>
<mml:mi mathvariant="bold">&#x3d5;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>where <inline-formula id="inf12">
<mml:math id="m16">
<mml:mo>&#x2297;</mml:mo>
</mml:math>
</inline-formula> represents the tensor product. <inline-formula id="inf13">
<mml:math id="m17">
<mml:mi>x</mml:mi>
</mml:math>
</inline-formula> represents the Mel spectrogram of each input, and <inline-formula id="inf14">
<mml:math id="m18">
<mml:mi>N</mml:mi>
</mml:math>
</inline-formula> is the total number of pixels in the Mel spectrogram. <inline-formula id="inf15">
<mml:math id="m19">
<mml:mrow>
<mml:mi mathvariant="italic">&#x3d5;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the representation of the first pixel in the Mel spectrogram mapped to a two-dimensional space, and <inline-formula id="inf16">
<mml:math id="m20">
<mml:mrow>
<mml:mi>&#x3a6;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the high-dimensional mapping form of the Mel spectrogram. Given the high-dimensional features, for the input Mel spectrogram, the decision function of the event tagging can be expressed as<disp-formula id="e5">
<mml:math id="m21">
<mml:mrow>
<mml:msup>
<mml:mi>f</mml:mi>
<mml:mi>m</mml:mi>
</mml:msup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mi mathvariant="italic">&#x3c8;</mml:mi>
<mml:mi>m</mml:mi>
</mml:msup>
<mml:mo>&#x22c5;</mml:mo>
<mml:mi mathvariant="bold">&#x3a6;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>
<disp-formula id="e6">
<mml:math id="m22">
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>arg</mml:mi>
<mml:mi>max</mml:mi>
<mml:msup>
<mml:mi>f</mml:mi>
<mml:mi>m</mml:mi>
</mml:msup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>
</p>
<p>Here, <inline-formula id="inf17">
<mml:math id="m23">
<mml:mi>m</mml:mi>
</mml:math>
</inline-formula> represents M categories, <inline-formula id="inf18">
<mml:math id="m24">
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mn>0,1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>...</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>M</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, where <inline-formula id="inf19">
<mml:math id="m25">
<mml:mrow>
<mml:msup>
<mml:mi mathvariant="italic">&#x3c8;</mml:mi>
<mml:mi>m</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is the trainable weight tensor. The model of the decision module in audio tagging is shown on the left of <xref ref-type="fig" rid="F3">Figure 3</xref> and in <xref ref-type="disp-formula" rid="e5">Eq. 5</xref>. <inline-formula id="inf20">
<mml:math id="m26">
<mml:mrow>
<mml:msup>
<mml:mi mathvariant="italic">&#x3c8;</mml:mi>
<mml:mi>m</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is a weight tensor, and its dimension is as high as <inline-formula id="inf21">
<mml:math id="m27">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mo>&#x22c5;</mml:mo>
<mml:msup>
<mml:mn>2</mml:mn>
<mml:mi>N</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>, which is difficult to be calculated. After decomposing <inline-formula id="inf22">
<mml:math id="m28">
<mml:mrow>
<mml:msup>
<mml:mi mathvariant="italic">&#x3c8;</mml:mi>
<mml:mi>m</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> into the chained small tensors through the MPS, the two-dimensional space that can be mapped with each pixel can be contracted with the weight tensor <inline-formula id="inf23">
<mml:math id="m29">
<mml:mrow>
<mml:msup>
<mml:mi mathvariant="italic">&#x3c8;</mml:mi>
<mml:mi>m</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>. In this way, the calculation can only be carried out between the small tensors, without directly calculating the weight tensors with high dimensionality. <xref ref-type="fig" rid="F3">Figure 3</xref> is a linear model of the decision module in audio tagging represented by the tensor network notation. For details on the tensor network notation, reference in [<xref ref-type="bibr" rid="B27">27</xref>]can be referred. As shown by the small green tensor in <xref ref-type="fig" rid="F3">Figure 3</xref>, <inline-formula id="inf24">
<mml:math id="m30">
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the form in which the two-dimensional space mapped by each pixel is connected to the weight tensor <inline-formula id="inf25">
<mml:math id="m31">
<mml:mrow>
<mml:msup>
<mml:mi mathvariant="italic">&#x3c8;</mml:mi>
<mml:mi>m</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>. The nodes in the first column are the pixels of each Mel spectrogram after being mapped to the two-dimensional space. They are connected to the weight tensor obtained after the training. There is an index <inline-formula id="inf26">
<mml:math id="m32">
<mml:mi>m</mml:mi>
</mml:math>
</inline-formula> on the right side of <inline-formula id="inf27">
<mml:math id="m33">
<mml:mrow>
<mml:msup>
<mml:mi mathvariant="italic">&#x3c8;</mml:mi>
<mml:mi>m</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>, whose dimension is the number of the final tagging classes.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Linear model of the decision module in audio tagging.</p>
</caption>
<graphic xlink:href="fphy-10-863291-g003.tif"/>
</fig>
<p>This mapping method will result in a huge number of parameters in the weight tensor. The matrix product state is the name for tensor train decomposition in physics. It approximates a large tensor to the product form of several second-order and third-order tensors. In this way, the contraction can be performed in the way on the right side of <xref ref-type="fig" rid="F3">Figure 3</xref>, to avoid the direct calculation of the ultra-high dimensional tensor, and the calculation amount will be greatly reduced. A high-dimensional tensor <inline-formula id="inf28">
<mml:math id="m34">
<mml:mi>T</mml:mi>
</mml:math>
</inline-formula> is decomposed into an approximate tensor <inline-formula id="inf29">
<mml:math id="m35">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>T</mml:mi>
<mml:mo>&#x223c;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> by the tensor train [<xref ref-type="bibr" rid="B28">28</xref>], as shown in <xref ref-type="disp-formula" rid="e7">Eq. 7</xref>.<disp-formula id="e7">
<mml:math id="m36">
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>T</mml:mi>
<mml:mo>&#x223c;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mn>...</mml:mn>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:mrow>
<mml:msubsup>
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mstyle>
<mml:msubsup>
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mn>...</mml:mn>
<mml:msubsup>
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>
</p>
<p>The weight tensor <inline-formula id="inf30">
<mml:math id="m37">
<mml:mrow>
<mml:msup>
<mml:mi mathvariant="italic">&#x3c8;</mml:mi>
<mml:mi>m</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is approximated by the product form of some two-dimensional and three-dimensional tensors according to <xref ref-type="disp-formula" rid="e7">Eq. 7</xref>. The approximated weight tensor is shown in <xref ref-type="disp-formula" rid="e8">Eq. 8</xref> and on the right of <xref ref-type="fig" rid="F3">Figure 3</xref>.<disp-formula id="e8">
<mml:math id="m38">
<mml:mrow>
<mml:msup>
<mml:mi mathvariant="italic">&#x3c8;</mml:mi>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mn>...</mml:mn>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mn>...</mml:mn>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:mrow>
<mml:msubsup>
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mstyle>
<mml:msubsup>
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msubsup>
<mml:mn>...</mml:mn>
<mml:msubsup>
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msubsup>
<mml:mn>...</mml:mn>
<mml:msubsup>
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(8)</label>
</disp-formula>where<inline-formula id="inf31">
<mml:math id="m39">
<mml:mi>A</mml:mi>
</mml:math>
</inline-formula> is the decomposed second-order and third-order tensors. The subscript <inline-formula id="inf32">
<mml:math id="m40">
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is called the free index, and the free index <inline-formula id="inf33">
<mml:math id="m41">
<mml:mi>m</mml:mi>
</mml:math>
</inline-formula> corresponds to the right side of <xref ref-type="fig" rid="F3">Figure 3</xref>, and its dimension is the number of tagging classes. The subscript <inline-formula id="inf34">
<mml:math id="m42">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is an auxiliary indicator, and its dimension is called the bond dimension, which controls the quality of the approximation. The size of the bond dimension determines the size of the tensor. The components of the tensor <inline-formula id="inf35">
<mml:math id="m43">
<mml:mi>A</mml:mi>
</mml:math>
</inline-formula> are the variational parameters determined through the training.</p>
</sec>
<sec id="s2-2-3">
<title>2.2.3 Local Orderless Operation</title>
<p>Since MPS is a one-dimensional tensor network, the neighboring pixels in the spectrogram are usually highly correlated. Therefore, directly flattening and inputting the Mel spectral feature into the MPS layer will cause the loss of spatial information. Spatial information includes the information of a single frame in the vertical direction, as well as the information between the frames in the horizontal direction, which is very important for audio tagging. In order to solve this problem, the local orderless operation according to the local orderless theory is used in the tensor network [<xref ref-type="bibr" rid="B29">29</xref>, <xref ref-type="bibr" rid="B30">30</xref>]. The local orderless operation divides a large patch into many small patches. After the small patches are contracted, the dimension of the output vector is <inline-formula id="inf36">
<mml:math id="m44">
<mml:mi mathvariant="bold-italic">v</mml:mi>
</mml:math>
</inline-formula> , and <inline-formula id="inf37">
<mml:math id="m45">
<mml:mi mathvariant="bold-italic">v</mml:mi>
</mml:math>
</inline-formula> is set to the same size as the bond dimension. This step can be interpreted as using a vector of dimension <inline-formula id="inf38">
<mml:math id="m46">
<mml:mi mathvariant="bold-italic">v</mml:mi>
</mml:math>
</inline-formula> to represent small patches of information, similar to feature extraction. Each small patch contains global features, which can better preserve the spatial information.</p>
<p>First, the Mel spectrogram is divided into four parts, as shown in <xref ref-type="fig" rid="F4">Figure 4</xref>. The first pixel of each part is taken out and combined into a <inline-formula id="inf39">
<mml:math id="m47">
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> local orderless small patch, as shown in the red box in <xref ref-type="fig" rid="F4">Figure 4</xref>. Then, the pixels of each part are combined according to this step, until the last pixel in the black box, as shown in <xref ref-type="fig" rid="F4">Figure 4</xref>. The pixel order in the patch is shown in <xref ref-type="disp-formula" rid="e9">Eq. 9</xref>.<disp-formula id="e9">
<mml:math id="m48">
<mml:mrow>
<mml:msup>
<mml:mi>&#x3a1;</mml:mi>
<mml:mi>K</mml:mi>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="italic">K</mml:mi>
<mml:mtext>&#xa0;&#xa0;&#xa0;&#xa0;&#xa0;&#xa0;</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mtext>&#xa0;&#xa0;&#xa0;&#xa0;&#xa0;&#xa0;</mml:mtext>
<mml:mi mathvariant="italic">K</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mi mathvariant="normal">W</mml:mi>
<mml:mn>2</mml:mn>
</mml:mfrac>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mi>K</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="normal">H</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi mathvariant="normal">W</mml:mi>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">H</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mi mathvariant="normal">W</mml:mi>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:mfrac>
</mml:mtd>
</mml:mtr>
</mml:mtable>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mtext>&#xa0;&#xa0;&#xa0;&#xa0;&#xa0;</mml:mtext>
<mml:mrow>
<mml:mrow>
<mml:mo>&#x2200;</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>...</mml:mn>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">H</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi mathvariant="normal">W</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>/</mml:mo>
<mml:mn>4</mml:mn>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(9)</label>
</disp-formula>where <inline-formula id="inf40">
<mml:math id="m49">
<mml:mrow>
<mml:msup>
<mml:mi>P</mml:mi>
<mml:mi>K</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> represents the local orderless small patch, the superscript K represents the sequential number of small patches, and H and W represent the height and width of the Mel spectrogram, respectively.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Local orderless operation and contraction.</p>
</caption>
<graphic xlink:href="fphy-10-863291-g004.tif"/>
</fig>
<p>Then the small patches are flattened and input into the MPS layer to contract. Then all the output vectors <inline-formula id="inf41">
<mml:math id="m50">
<mml:mi mathvariant="bold-italic">v</mml:mi>
</mml:math>
</inline-formula> are reshaped into images. The converted graph has a smaller resolution than the previous Mel spectrogram, but the important information will be preserved. This operation is repeated on the converted image. After the three MPS layers, the resolution of the generated image is already very small, but the features and spatial information of the original Mel spectrogram are well-preserved.</p>
</sec>
<sec id="s2-2-4">
<title>2.2.4 Contraction and Optimization</title>
<p>After the three MPS layers of contraction, a small size image has been generated. It has spatial structure information and important features of the Mel spectrogram. It is flattened into the last MPS layer, as shown in <xref ref-type="fig" rid="F4">Figure 4</xref>. In line with the implementation method from the MPS in Miller [<xref ref-type="bibr" rid="B31">31</xref>], the horizontal edges are first contracted in parallel to get the contracted tensors, and then, these tensors are contracted vertically. The output is generated by connecting the free indicators of the tensor. A recent work has proposed a more effective calculation method [<xref ref-type="bibr" rid="B32">32</xref>, <xref ref-type="bibr" rid="B33">33</xref>], which is expected to further accelerate the calculation speed.</p>
</sec>
</sec>
</sec>
<sec id="s3">
<title>3 Experiments</title>
<sec id="s3-1">
<title>3.1 Datasets</title>
<p>In the experiment, development and validation datasets of DCASE 2018 challenge task 5 are used to evaluate the audio tagging for domestic activities. DCASE 2018 challenge task 5 is a derivative of the SINS dataset. It contains a continuous recording of one person living in a holiday home over a period of 1&#xa0;week. It was collected using a network of 13 microphone arrays distributed over the entire home. The microphone array consisted of four linearly arranged microphones. For this task, seven microphone arrays are used in the living room and kitchen area combined. The continuous recordings are split into audio segments of 10&#xa0;s. These audio segments are provided as individual files along with the ground truth. The dataset contains 72,984 audio files. Each audio segment contains four channels. It is organized with nine class labels consisting of absence, cooking, dishwashing, eating, social activity, vacuum cleaning, watching TV, and working. The audio files are recorded with 16&#xa0;kHz sampled frequency, and the number of files in each class are not the same.</p>
</sec>
<sec id="s3-2">
<title>3.2 Evaluation Method</title>
<p>In this experiment, the development dataset and validation dataset are divided into the training set, validation set, and test set with a ratio of 8:1:1, respectively. The evaluation criteria include the precision rate, recall rate, and F1-score. Precision is the ratio of real positive samples to samples that are predicted to be positive, which is specific to the predicted samples. Recall is the ratio of the correct predictions to the positive cases in the sample, which is specific to the actual samples. The F1 score is calculated based on recall and precision. The experimental results in this article are the results of the development and the validation datasets in the divided test set, respectively. These criteria are obtained by calculating the confusion matrix given by <xref ref-type="disp-formula" rid="e10">Eqs 10</xref>&#x2013;<xref ref-type="disp-formula" rid="e12">12</xref>.<disp-formula id="e10">
<mml:math id="m51">
<mml:mrow>
<mml:mi mathvariant="italic">Pr</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(10)</label>
</disp-formula>
<disp-formula id="e11">
<mml:math id="m52">
<mml:mrow>
<mml:mi mathvariant="italic">Re</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>l</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(11)</label>
</disp-formula>
<disp-formula id="e12">
<mml:math id="m53">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>S</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mi mathvariant="italic">Pr</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi mathvariant="italic">Re</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">Pr</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>si</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="italic">Re</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(12)</label>
</disp-formula>where TP is the number of true positive results, TN is the number of true negative results, FP is the number of false positive results, and FN is the number of false negative results.</p>
</sec>
<sec id="s3-3">
<title>3.3 Experimental Setup and Result</title>
<p>There are four channels <inline-formula id="inf42">
<mml:math id="m54">
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mn>4</mml:mn>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> within one audio signal. The four channels are manually averaged [<xref ref-type="bibr" rid="B21">21</xref>] to yield <inline-formula id="inf43">
<mml:math id="m55">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mn>5</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, where <inline-formula id="inf44">
<mml:math id="m56">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mn>5</mml:mn>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mn>4</mml:mn>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>/</mml:mo>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> so as to better fuse the four channels and augment the dataset. The audio signal is converted to a Mel spectrogram, as described in <xref ref-type="sec" rid="s2">Section 2</xref>. The window type, window size, overlap, and FFT size parameters are set to Hamming, 480, 240, and 480, respectively. The Hamming window is adopted for signal framing as it can effectively overcome the leakage phenomenon [<xref ref-type="bibr" rid="B34">34</xref>]. The dimension of the Mel spectrogram is <inline-formula id="inf45">
<mml:math id="m57">
<mml:mrow>
<mml:mn>336</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>336</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> as the input of the neural network model based on the tensor network, which is composed of the two convolutional layers and four MPS layers. It is then the horizontal flip, vertical flip, and random rotation that enlarge the training data, avoid overfitting, and enhance the robustness of the model. The batch size is set to 256, bond dimension is set to 5, and the initial learning rate is 0.001. The optimizer and loss function used in the training are Adam and cross-entropy loss function. The structure of the neural network model based on the tensor network is shown in <xref ref-type="fig" rid="F2">Figure 2</xref>, and the states of TP, TN, FP, and FN in the test set of the development dataset are shown for each class on the confusion matrix in <xref ref-type="fig" rid="F5">Figure 5</xref>.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Confusion matrix for the test set of the development dataset.</p>
</caption>
<graphic xlink:href="fphy-10-863291-g005.tif"/>
</fig>
<p>As can be seen from the confusion matrix in <xref ref-type="fig" rid="F5">Figure 5</xref>, the abscissa is the true class, and the ordinate is the predicted class. The blue square indicates that the predicted class is the same as the true class. The color intensity corresponds to the number of audio tagging. It can be found from <xref ref-type="fig" rid="F5">Figure 5</xref> that the proposed model judges 165 working audios as absence and 134 absence audios as working. Because people may make very small noises at work, it is easy to confuse it with the absence class. In addition, the model judges 39 and 33 other class audios as absence and working class, and many labels for other class audios cannot be distinguished since the other class is not a class of specific activities. There are many types of features extracted from the other class audio, and the common features of other class are difficult to learn. As a result, a lot of audio signals are near the decision boundary, and it is easy to be misjudged as absence, working, and eating. But for the prediction results, it can be found from <xref ref-type="fig" rid="F5">Figure 5</xref> that the prediction results for the other class are more accurate, proving the better learning ability of the tensor network model.</p>
<p>In order to compare with other models more intuitively, other performance criteria including precision, recall, and F1-score are separately given in <xref ref-type="table" rid="T1">Table 1</xref> for each class of DCASE 2018 challenge task 5. It can be seen from the <xref ref-type="table" rid="T1">Table 1</xref> that the precision rate of the other class is much higher than the recall rate. This shows that the prediction of the other class is more accurate, but many other class audios are prone to judgment errors. Since the other class is not a class of specific activities, the tensor network can better learn the common features of the class. But for the other class audio with less commonality, it is less possible to identify the deeper rules.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Neural network model based on the tensor network performance criteria in the test set of the development dataset.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Class</th>
<th align="center">Precision/%</th>
<th align="center">Recall/%</th>
<th align="center">F1-score/%</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Absence</td>
<td align="char" char=".">89.00</td>
<td align="char" char=".">92.63</td>
<td align="char" char=".">90.78</td>
</tr>
<tr>
<td align="left">Cooking</td>
<td align="char" char=".">95.69</td>
<td align="char" char=".">95.30</td>
<td align="char" char=".">95.49</td>
</tr>
<tr>
<td align="left">Dishwashing</td>
<td align="char" char=".">82.61</td>
<td align="char" char=".">79.72</td>
<td align="char" char=".">81.14</td>
</tr>
<tr>
<td align="left">Eating</td>
<td align="char" char=".">82.61</td>
<td align="char" char=".">74.03</td>
<td align="char" char=".">78.09</td>
</tr>
<tr>
<td align="left">Other</td>
<td align="char" char=".">79.84</td>
<td align="char" char=".">50.00</td>
<td align="char" char=".">61.49</td>
</tr>
<tr>
<td align="left">Social activity</td>
<td align="char" char=".">96.11</td>
<td align="char" char=".">94.95</td>
<td align="char" char=".">95.53</td>
</tr>
<tr>
<td align="left">Vacuum cleaning</td>
<td align="char" char=".">98.00</td>
<td align="char" char=".">100.00</td>
<td align="char" char=".">98.99</td>
</tr>
<tr>
<td align="left">Watching TV</td>
<td align="char" char=".">99.79</td>
<td align="char" char=".">99.57</td>
<td align="char" char=".">99.68</td>
</tr>
<tr>
<td align="left">Working</td>
<td align="char" char=".">87.18</td>
<td align="char" char=".">89.14</td>
<td align="char" char=".">88.15</td>
</tr>
<tr>
<td align="left">Average value</td>
<td align="char" char=".">90.09</td>
<td align="char" char=".">86.15</td>
<td align="char" char=".">87.70</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>
<xref ref-type="table" rid="T2">Table 2</xref> is a comparison between the proposed model and other models, which represent several typical and commonly used networks, including the CNN, RESNET, and LSTM. This experiment selected the three models to compare with the neural network model based on tensor network (NNMBTN) model, namely, the baseline system [<xref ref-type="bibr" rid="B9">9</xref>], TC2DCNN [<xref ref-type="bibr" rid="B12">12</xref>], and INRC_2D [<xref ref-type="bibr" rid="B13">13</xref>]. The baseline system uses a neural network architecture based on the convolutional layers and dense layers. TC2DCNN is extended by operating the convolutions along the two dimensions of time and channel. INRC_2D is processed in parallel by RESNET and long short-term memory (LSTM) network, with a fully joint connected layer.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Comparison of the neural network model based on the tensor network with other models in the test set of the development dataset.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Class</th>
<th colspan="4" align="center">Detecting F1-score (%) for the used methods</th>
</tr>
<tr>
<th align="center">Baseline system [<xref ref-type="bibr" rid="B9">9</xref>]</th>
<th align="center">TC2DCNN [<xref ref-type="bibr" rid="B12">12</xref>]</th>
<th align="center">INRC_2D [<xref ref-type="bibr" rid="B13">13</xref>]</th>
<th align="center">NNMBTN</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Absence</td>
<td align="char" char=".">85.41</td>
<td align="char" char=".">86.62</td>
<td align="char" char=".">83.95</td>
<td align="char" char=".">90.78</td>
</tr>
<tr>
<td align="left">Cooking</td>
<td align="char" char=".">95.14</td>
<td align="char" char=".">93.34</td>
<td align="char" char=".">95.47</td>
<td align="char" char=".">95.49</td>
</tr>
<tr>
<td align="left">Dishwashing</td>
<td align="char" char=".">76.73</td>
<td align="char" char=".">72.68</td>
<td align="char" char=".">78.00</td>
<td align="char" char=".">81.14</td>
</tr>
<tr>
<td align="left">Eating</td>
<td align="char" char=".">83.64</td>
<td align="char" char=".">87.03</td>
<td align="char" char=".">89.68</td>
<td align="char" char=".">78.09</td>
</tr>
<tr>
<td align="left">Other</td>
<td align="char" char=".">44.76</td>
<td align="char" char=".">53.81</td>
<td align="char" char=".">55.88</td>
<td align="char" char=".">61.49</td>
</tr>
<tr>
<td align="left">Social activity</td>
<td align="char" char=".">93.92</td>
<td align="char" char=".">93.94</td>
<td align="char" char=".">93.97</td>
<td align="char" char=".">95.53</td>
</tr>
<tr>
<td align="left">Vacuum cleaning</td>
<td align="char" char=".">99.31</td>
<td align="char" char=".">99.79</td>
<td align="char" char=".">100.00</td>
<td align="char" char=".">98.99</td>
</tr>
<tr>
<td align="left">Watching TV</td>
<td align="char" char=".">99.59</td>
<td align="char" char=".">99.38</td>
<td align="char" char=".">99.40</td>
<td align="char" char=".">99.68</td>
</tr>
<tr>
<td align="left">Working</td>
<td align="char" char=".">82.03</td>
<td align="char" char=".">85.14</td>
<td align="char" char=".">85.22</td>
<td align="char" char=".">88.15</td>
</tr>
<tr>
<td align="left">Average value</td>
<td align="char" char=".">84.50</td>
<td align="char" char=".">85.75</td>
<td align="char" char=".">86.84</td>
<td align="char" char=".">87.70</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>It can be seen from <xref ref-type="table" rid="T2">Table 2</xref> that the F1-score of the proposed method on the test set of the DCASE 2018 challenge task 5 development set is 87.70%, which is 3.2% higher than the baseline system, 1.95% higher than the TC2DCNN system, and 0.86% higher than INRC_2D system. The proposed method has five classes that are higher than the baseline, TC2DCNN, and INRC_2D. This shows that the tensor network model can identify the important features well after obtaining the features extracted by the convolutional layer. At the same time, the spatial information of the audio is well-preserved. Compared with the other models, the tensor network has powerful representation ability in the high-dimensional space and can separate the different classes of audio with hyperplane. There is little difference in the F1-score performance on cooking, vacuum cleaning, and working. The score advantage of other classes is more obvious, 5.61% higher than the INRC_2D system, which shows that for classes of not specific activities, the tensor network can also learn the features better.</p>
<p>The data provided in the evaluation set are based on the sensor nodes that do not exist in the development set and can provide data from the same nodes in the development set. The F1-scores of each model in the test set of the validation set are shown in <xref ref-type="table" rid="T3">Table 3</xref>.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Comparison of the neural network model based on the tensor network with other models in the test set of the validation dataset.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Class</th>
<th colspan="4" align="center">Detecting F1-score (%) for the used methods</th>
</tr>
<tr>
<th align="center">Baseline system [<xref ref-type="bibr" rid="B9">9</xref>]</th>
<th align="center">TC2DCNN [<xref ref-type="bibr" rid="B12">12</xref>]</th>
<th align="center">INRC_2D [<xref ref-type="bibr" rid="B13">13</xref>]</th>
<th align="center">NNMBTN</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Absence</td>
<td align="char" char=".">87.7</td>
<td align="char" char=".">79.8</td>
<td align="char" char=".">79.7</td>
<td align="char" char=".">90.2</td>
</tr>
<tr>
<td align="left">Cooking</td>
<td align="char" char=".">93.0</td>
<td align="char" char=".">88.7</td>
<td align="char" char=".">86.9</td>
<td align="char" char=".">95.0</td>
</tr>
<tr>
<td align="left">Dishwashing</td>
<td align="char" char=".">77.2</td>
<td align="char" char=".">71.8</td>
<td align="char" char=".">73.8</td>
<td align="char" char=".">82.3</td>
</tr>
<tr>
<td align="left">Eating</td>
<td align="char" char=".">81.2</td>
<td align="char" char=".">78.9</td>
<td align="char" char=".">82.2</td>
<td align="char" char=".">77.0</td>
</tr>
<tr>
<td align="left">Other</td>
<td align="char" char=".">35.0</td>
<td align="char" char=".">17.6</td>
<td align="char" char=".">42.7</td>
<td align="char" char=".">55.5</td>
</tr>
<tr>
<td align="left">Social activity</td>
<td align="char" char=".">96.6</td>
<td align="char" char=".">96.2</td>
<td align="char" char=".">97.1</td>
<td align="char" char=".">93.4</td>
</tr>
<tr>
<td align="left">Vacuum cleaning</td>
<td align="char" char=".">95.8</td>
<td align="char" char=".">94.4</td>
<td align="char" char=".">97.4</td>
<td align="char" char=".">98.2</td>
</tr>
<tr>
<td align="left">Watching TV</td>
<td align="char" char=".">99.9</td>
<td align="char" char=".">99.7</td>
<td align="char" char=".">99.9</td>
<td align="char" char=".">99.5</td>
</tr>
<tr>
<td align="left">Working</td>
<td align="char" char=".">81.4</td>
<td align="char" char=".">64.6</td>
<td align="char" char=".">75.5</td>
<td align="char" char=".">82.3</td>
</tr>
<tr>
<td align="left">Average value</td>
<td align="char" char=".">83.1</td>
<td align="char" char=".">76.9</td>
<td align="char" char=".">81.7</td>
<td align="char" char=".">85.9</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Compared with the results on the development set, it can be seen from <xref ref-type="table" rid="T3">Table 3</xref> that the proposed model in the test set of the validation set is lower than the other models in the two categories of eating and social activities and higher than the other models in both categories of dishwashing and vacuuming. The advantages of the other categories are still obvious. The average F1-score reaches 85.9%, which is 9.0% higher than TC2DCNN, 4.2% higher than INRC2D, and 2.8% higher than the baseline. The F1-scores of the neural network model based on the tensor network are relatively stable, which proves that the proposed network has a good generalization ability. On the whole, the proposed model has better ability to extract and learn the important features of the data.</p>
<p>In order to better demonstrate the compression ability of the MPS to the network, the MPS layer in the proposed model is replaced by the convolutional layer, max pooling layer, and fully connected layer. We compared the proposed model (NNMBTN: 2CNN&#x2b;4MPS) with the traditional CNN-based model which is composed of four CNNs, Maxpool, and a fully connected layer. The model comparison results are shown in <xref ref-type="table" rid="T4">Table 4</xref>.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Performance and parameter comparison between the proposed model and the traditional neural network.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Model</th>
<th align="center">Precision/%</th>
<th align="center">Recall/%</th>
<th align="center">F1-score/%</th>
<th align="center">Parameter quantity (M)</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">4CNN &#x2b; Maxpool &#x2b; fully connected</td>
<td align="char" char=".">74.11</td>
<td align="char" char=".">64.08</td>
<td align="char" char=".">65.8</td>
<td align="char" char=".">23.70</td>
</tr>
<tr>
<td align="left">NNMBTN (2CNN&#x2b;4MPS)</td>
<td align="char" char=".">88.08</td>
<td align="char" char=".">84.13</td>
<td align="char" char=".">85.9</td>
<td align="char" char=".">17.74</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>It can be seen from <xref ref-type="table" rid="T4">Table 4</xref> that the parameters of the proposed model are one quarter smaller than that of the traditional neural network after replacement, and the effect is also better than that of the traditional neural network, which shows that the MPS layer has better compression ability for the network.</p>
<p>To further investigate the effect of combining the proposed model with the state-of-the-art model, separable convolutions network [<xref ref-type="bibr" rid="B5">5</xref>] are used to verify the feasibility of the proposed model. Separable convolutions network consists of four convolutional layers using <inline-formula id="inf46">
<mml:math id="m58">
<mml:mrow>
<mml:mn>5</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>5</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> filters, followed by a global pooling layer and a final MLP (Multilayer Perceptron). The separate convolution network structure is improved to be combined with the MPS tensor network in which only two layers of separate convolutions network are retained. In the comparison experiments, only the network structure was changed, and the rest remained unchanged. The experimental results are shown in <xref ref-type="table" rid="T5">Table 5</xref>.</p>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>Performance comparison between the separable convolution model and the separable convolution model combined with tensor networks.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Model</th>
<th align="center">F1-score/%</th>
<th align="center">Parameter quantity (M)</th>
<th align="center">GPU(GB)</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Separable Convolutions [<xref ref-type="bibr" rid="B5">5</xref>] (batch size &#x3d; 128)</td>
<td align="char" char=".">90.78</td>
<td align="char" char=".">4.20</td>
<td align="char" char=".">3.85</td>
</tr>
<tr>
<td align="left">(2SepConv&#x2b;3MPS) (batch size &#x3d; 128)</td>
<td align="char" char=".">89.52</td>
<td align="char" char=".">4.16</td>
<td align="char" char=".">1.72</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>It can be seen from <xref ref-type="table" rid="T5">Table 5</xref> that the GPU occupancy in the training procedure is reduced by 55% after combining with the tensor network under the same conditions except for the network structure. This shows that the tensor network can better reduce the redundancy of the network during the training. In terms of parameter quantity, the parameter quantity of the separate convolution is slightly smaller than that of ordinary convolution. The F1-score is slightly lower than the split convolutional network. Compared with the state-of-the-art model, the combination of the tensor network can reduce the redundancy of the network to achieve a balance between efficiency and accuracy. In the future, more research practices could be carried out to find a better way when combining the tensor network with the new neural network approaches.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Conclusion</title>
<p>In this article, the neural network model based on the tensor network is proposed for audio tagging of domestic activities, which takes the advantage of the CNN in extracting spatial features and the MPS tensor network for better interpretability and the ability to compress the network with tensor train decomposition. The MPS is one-dimensional tensor network structure, which is based on tensor train decomposition. It uses the chain-connected small tensors to represent the high-dimensional tensors. The proposed model is composed of two convolutional layers and four MPS layers. The function of the first three MPS layers is to extract the features, and the last MPS layer is used as a classifier. The DCASE 2018 challenge task 5 datasets are considered in the experiment, and the F1-score is calculated for performance evaluation. The experimental results show that the neural network model based on the tensor network proposed in this article has a good learning ability. The results show that the average F1-Score of the proposed neural network model based on the tensor network in the test set of the development dataset and validation dataset of DCASE 2018 challenge task 5 reached 87.7 and 85.9%, which were 3.2 and 2.8% higher than the baseline system, respectively. When compared with the state-of-the-art model, the combination of the tensor network can reduce the redundancy of the network to achieve a balance between the efficiency and accuracy. It is verified that the proposed model can function better for the task of audio tagging of domestic activities.</p>
<p>In the future, it is necessary to extract more representative audio features in the face of a huge database. There are some other structures of tensor networks, such as PEPS and MERA, and the combination of these models with the neural networks deserves a further in-depth study. In addition, the classes of the sound events in household activities are more complex, so expanding the dataset and improving the audio tagging accuracy are also necessary.</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>The raw data supporting the conclusion of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>L-DY and R-BY mainly carried out topic selection, experiments, and completion of the first draft. JW mainly checked the logical structure and writing of the manuscript. ML was mainly responsible for the revision and polishing of the manuscript.</p>
</sec>
<sec id="s7">
<title>Funding</title>
<p>The Central Government Guides Local Science and Technology Development Fund Project of China (grant number: 2021ZY0004). This work was supported by the National Natural Science Foundation of China (Grant Nos. 62071039 and 62161040), the Natural Science Foundation of Inner Mongolia (Grant No. 2021MS06030), and the Science and Technology Major Project of Inner Mongolia (Grant No. 2021GG0023).</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors, and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<label>1.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rafferty</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Nugent</surname>
<given-names>CD</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>L</given-names>
</name>
</person-group>. <article-title>From Activity Recognition to Intention Recognition for Assisted Living within Smart Homes</article-title>. <source>IEEE Trans Human-mach Syst</source> (<year>2017</year>) <volume>47</volume>(<issue>3</issue>):<fpage>368</fpage>&#x2013;<lpage>79</lpage>. <pub-id pub-id-type="doi">10.1109/thms.2016.2641388</pub-id> </citation>
</ref>
<ref id="B2">
<label>2.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Erden</surname>
<given-names>F</given-names>
</name>
<name>
<surname>Velipasalar</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Alkar</surname>
<given-names>AZ</given-names>
</name>
<name>
<surname>Cetin</surname>
<given-names>AE</given-names>
</name>
</person-group>. <article-title>Sensors in Assisted Living: A Survey of Signal and Image Processing Methods</article-title>. <source>IEEE Signal Process Mag</source> (<year>2016</year>) <volume>33</volume>(<issue>2</issue>):<fpage>36</fpage>&#x2013;<lpage>44</lpage>. <pub-id pub-id-type="doi">10.1109/msp.2015.2489978</pub-id> </citation>
</ref>
<ref id="B3">
<label>3.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Phan</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Hertel</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Maass</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Koch</surname>
<given-names>P</given-names>
</name>
<name>
<surname>Mazur</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Mertins</surname>
<given-names>A</given-names>
</name>
</person-group>. <article-title>Improved Audio Scene Classification Based on Label-Tree Embeddings and Convolutional Neural Networks</article-title>. <source>Ieee/acm Trans Audio Speech Lang Process</source> (<year>2017</year>) <volume>25</volume>(<issue>6</issue>):<fpage>1278</fpage>&#x2013;<lpage>90</lpage>. <pub-id pub-id-type="doi">10.1109/taslp.2017.2690564</pub-id> </citation>
</ref>
<ref id="B4">
<label>4.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gong</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Chung</surname>
<given-names>Y-A</given-names>
</name>
<name>
<surname>Glass</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Psla: Improving Audio Tagging with Pretraining, Sampling, Labeling, and Aggregation</article-title>. <source>Ieee/acm Trans Audio Speech Lang Process</source> (<year>2021</year>) <volume>29</volume>:<fpage>3292</fpage>&#x2013;<lpage>306</lpage>. <pub-id pub-id-type="doi">10.1109/taslp.2021.3120633</pub-id> </citation>
</ref>
<ref id="B5">
<label>5.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Bursuc</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Puy</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Jain</surname>
<given-names>H</given-names>
</name>
</person-group>. <source>Separable Convolutions and Test-Time Augmentations for Low-Complexity and Calibrated Acoustic Scene Classification</source>. <publisher-loc>Barcelona, Spain</publisher-loc>: <publisher-name>Detection and Classification of Acoustic Scenes and Events 2021</publisher-name> (<year>2021</year>). </citation>
</ref>
<ref id="B6">
<label>6.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Dekkers</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Lauwereins</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Thoen</surname>
<given-names>B</given-names>
</name>
<name>
<surname>Adhana</surname>
<given-names>MW</given-names>
</name>
<name>
<surname>Brouckxon</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Van den Bergh</surname>
<given-names>B</given-names>
</name>
<etal/>
</person-group> <source>The Sins Database for Detection of Daily Activities in a Home Environment Using an Acoustic Sensor Network</source>. <publisher-loc>Munich, Germany</publisher-loc>: <publisher-name>Detection and Classification of Acoustic Scenes and Events 2017</publisher-name> (<year>2017</year>). p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>. </citation>
</ref>
<ref id="B7">
<label>7.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Kuang</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Microphone Array Speech Enhancement Based on Tensor Filtering Methods</article-title>. <source>China Commun</source> (<year>2018</year>) <volume>15</volume>(<issue>4</issue>):<fpage>141</fpage>&#x2013;<lpage>52</lpage>. <pub-id pub-id-type="doi">10.1109/cc.2018.8357692</pub-id> </citation>
</ref>
<ref id="B8">
<label>8.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Kuang</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Tensor Completion for Recovering Multichannel Audio Signal with Missing Data</article-title>. <source>China Commun</source> (<year>2019</year>) <volume>16</volume>(<issue>4</issue>):<fpage>186</fpage>&#x2013;<lpage>95</lpage>. <pub-id pub-id-type="doi">10.12676/j.cc.2019.04.014</pub-id> </citation>
</ref>
<ref id="B9">
<label>9.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Dekkers</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Vuegen</surname>
<given-names>L</given-names>
</name>
<name>
<surname>van Waterschoot</surname>
<given-names>T</given-names>
</name>
<name>
<surname>Vanrumste</surname>
<given-names>B</given-names>
</name>
<name>
<surname>Karsmakers</surname>
<given-names>P</given-names>
</name>
</person-group>. <source>Dcase 2018 Challenge-Task 5: Monitoring of Domestic Activities Based on Multi-Channel Acoustics</source>. <publisher-loc>Surrey, United Kingdom</publisher-loc>: <publisher-name>arXiv preprint arXiv:180711246</publisher-name> (<year>2018</year>). </citation>
</ref>
<ref id="B10">
<label>10.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Inoue</surname>
<given-names>T</given-names>
</name>
<name>
<surname>Vinayavekhin</surname>
<given-names>P</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Wood</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Greco</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Tachibana</surname>
<given-names>R</given-names>
</name>
</person-group>. <source>Domestic Activities Classification Based on Cnn Using Shuffling and Mixing Data Augmentation</source>. <publisher-loc>Surrey, United Kingdom</publisher-loc>: <publisher-name>DCASE 2018 Challenge</publisher-name> (<year>2018</year>). </citation>
</ref>
<ref id="B11">
<label>11.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Tanabe</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Endo</surname>
<given-names>T</given-names>
</name>
<name>
<surname>Nikaido</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Ichige</surname>
<given-names>T</given-names>
</name>
<name>
<surname>Nguyen</surname>
<given-names>P</given-names>
</name>
<name>
<surname>Kawaguchi</surname>
<given-names>Y</given-names>
</name>
<etal/>
</person-group> <source>Multichannel Acoustic Scene Classification by Blind Dereverberation, Blind Source Separation, Data Augmentation, and Model Ensembling</source>. <publisher-loc>Surrey, United Kingdom</publisher-loc>: <publisher-name>DCASE 2018 Challenge</publisher-name> (<year>2018</year>). </citation>
</ref>
<ref id="B12">
<label>12.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Tiraboschi</surname>
<given-names>M</given-names>
</name>
</person-group>. <source>Monitoring of Domestic Activities Based on Multi-Channel Acoustics: A Time-Channel {2d}-Convolutional Approach</source>. <publisher-loc>Surrey, United Kingdom</publisher-loc>: <publisher-name>DCASE 2018 Challenge</publisher-name> (<year>2018</year>). </citation>
</ref>
<ref id="B13">
<label>13.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Raveh</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Amar</surname>
<given-names>A</given-names>
</name>
</person-group>. <source>Multi-Channel Audio Classification with Neural Network Using Scattering Transform</source>. <publisher-loc>Surrey, United Kingdom</publisher-loc>: <publisher-name>Tech. Rep. DCASE</publisher-name> (<year>2018</year>). </citation>
</ref>
<ref id="B14">
<label>14.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kong</surname>
<given-names>Q</given-names>
</name>
<name>
<surname>Iqbal</surname>
<given-names>T</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Plumbley</surname>
<given-names>MD</given-names>
</name>
</person-group>. <source>Dcase 2018 Challenge Surrey Cross-Task Convolutional Neural Network Baseline</source>. <publisher-loc>Surrey, United Kingdom</publisher-loc>: <publisher-name>arXiv preprint arXiv:180800773</publisher-name> (<year>2018</year>). </citation>
</ref>
<ref id="B15">
<label>15.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hofmann</surname>
<given-names>T</given-names>
</name>
<name>
<surname>Sch&#xf6;lkopf</surname>
<given-names>B</given-names>
</name>
<name>
<surname>Smola</surname>
<given-names>AJ</given-names>
</name>
</person-group>. <article-title>Kernel Methods in Machine Learning</article-title>. <source>Ann Stat</source> (<year>2008</year>) <volume>36</volume>(<issue>3</issue>):<fpage>1171</fpage>&#x2013;<lpage>220</lpage>. <pub-id pub-id-type="doi">10.1214/009053607000000677</pub-id> </citation>
</ref>
<ref id="B16">
<label>16.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Kuang</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Compression of Head-Related Transfer Function Based on Tucker and Tensor Train Decomposition</article-title>. <source>IEEE Access</source> (<year>2019</year>) <volume>7</volume>:<fpage>39639</fpage>&#x2013;<lpage>51</lpage>. <pub-id pub-id-type="doi">10.1109/access.2019.2906364</pub-id> </citation>
</ref>
<ref id="B17">
<label>17.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Stoudenmire</surname>
<given-names>EM</given-names>
</name>
<name>
<surname>Schwab</surname>
<given-names>DJ</given-names>
</name>
</person-group>. <source>Supervised Learning with Quantum-Inspired Tensor Networks</source>. <publisher-loc>Barcelona, Spain</publisher-loc>: <publisher-name>arXiv preprint arXiv:160505775</publisher-name> (<year>2016</year>). </citation>
</ref>
<ref id="B18">
<label>18.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Efthymiou</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Hidary</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Leichenauer</surname>
<given-names>S</given-names>
</name>
</person-group>. <source>Tensornetwork for Machine Learning</source>. <publisher-loc>Ithaca, New York</publisher-loc>: <publisher-name>arXiv preprint arXiv:190606329</publisher-name> (<year>2019</year>). </citation>
</ref>
<ref id="B19">
<label>19.</label>
<citation citation-type="book">
<person-group person-group-type="editor">
<name>
<surname>Selvan</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Dam</surname>
<given-names>EB</given-names>
</name>
</person-group>, editors. <source>Tensor Networks for Medical Image Classification</source>. <publisher-loc>Montreal, QC</publisher-loc>: <publisher-name>Medical Imaging with Deep Learning</publisher-name> (<year>2020</year>). </citation>
</ref>
<ref id="B20">
<label>20.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Evenbly</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Vidal</surname>
<given-names>G</given-names>
</name>
</person-group>. <article-title>Tensor Network States and Geometry</article-title>. <source>J Stat Phys</source> (<year>2011</year>) <volume>145</volume>(<issue>4</issue>):<fpage>891</fpage>&#x2013;<lpage>918</lpage>. <pub-id pub-id-type="doi">10.1007/s10955-011-0237-4</pub-id> </citation>
</ref>
<ref id="B21">
<label>21.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>F</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>D</given-names>
</name>
</person-group>. <source>An Ensemble System for Domestic Activity Recognition</source>. <publisher-loc>Surrey, United Kingdom</publisher-loc>: <publisher-name>DCASE2018 Challenge, Tech Rep</publisher-name> (<year>2018</year>). </citation>
</ref>
<ref id="B22">
<label>22.</label>
<citation citation-type="book">
<person-group person-group-type="editor">
<name>
<surname>Kopparapu</surname>
<given-names>SK</given-names>
</name>
<name>
<surname>Laxminarayana</surname>
<given-names>M</given-names>
</name>
</person-group>, editors. <article-title>Choice of Mel Filter Bank in Computing Mfcc of a Resampled Speech</article-title>. In: <conf-name>Proceedings of the 10th International Conference on Information Science, Signal Processing and their Applications (ISSPA 2010)</conf-name>; <conf-date>2010 May 10</conf-date>; <conf-loc>Kuala Lumpur, Malaysia</conf-loc>. <publisher-name>IEEE</publisher-name> (<year>2010</year>). </citation>
</ref>
<ref id="B23">
<label>23.</label>
<citation citation-type="book">
<person-group person-group-type="editor">
<name>
<surname>Yanai</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Kawano</surname>
<given-names>Y</given-names>
</name>
</person-group>, editors. <article-title>Food Image Recognition Using Deep Convolutional Network with Pre-training and Fine-Tuning</article-title>. In: <conf-name>Proceedings of the 2015 IEEE International Conference on Multimedia &#x26; Expo Workshops (ICMEW)</conf-name>; <conf-date>2015 June 29</conf-date>; <conf-loc>Turin, Italy</conf-loc>. <publisher-name>IEEE</publisher-name> (<year>2015</year>). </citation>
</ref>
<ref id="B24">
<label>24.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kalchbrenner</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Grefenstette</surname>
<given-names>E</given-names>
</name>
<name>
<surname>Blunsom</surname>
<given-names>P</given-names>
</name>
</person-group>. <source>A Convolutional Neural Network for Modelling Sentences</source>. <publisher-loc>Ithaca, New York</publisher-loc>: <publisher-name>arXiv preprint arXiv:14042188</publisher-name> (<year>2014</year>). </citation>
</ref>
<ref id="B25">
<label>25.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schmidt-Hieber</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Nonparametric Regression Using Deep Neural Networks with Relu Activation Function</article-title>. <source>Ann Stat</source> (<year>2020</year>) <volume>48</volume>(<issue>4</issue>):<fpage>1875</fpage>&#x2013;<lpage>97</lpage>. <pub-id pub-id-type="doi">10.1214/19-aos1875</pub-id> </citation>
</ref>
<ref id="B26">
<label>26.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sigtia</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Benetos</surname>
<given-names>E</given-names>
</name>
<name>
<surname>Dixon</surname>
<given-names>S</given-names>
</name>
</person-group>. <article-title>An End-To-End Neural Network for Polyphonic Piano Music Transcription</article-title>. <source>Ieee/acm Trans Audio Speech Lang Process</source> (<year>2016</year>) <volume>24</volume>(<issue>5</issue>):<fpage>927</fpage>&#x2013;<lpage>39</lpage>. <pub-id pub-id-type="doi">10.1109/taslp.2016.2533858</pub-id> </citation>
</ref>
<ref id="B27">
<label>27.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bridgeman</surname>
<given-names>JC</given-names>
</name>
<name>
<surname>Chubb</surname>
<given-names>CT</given-names>
</name>
</person-group>. <article-title>Hand-Waving and Interpretive Dance: An Introductory Course on Tensor Networks</article-title>. <source>J Phys A: Math Theor</source> (<year>2017</year>) <volume>50</volume>(<issue>22</issue>):<fpage>223001</fpage>. <pub-id pub-id-type="doi">10.1088/1751-8121/aa6dc3</pub-id> </citation>
</ref>
<ref id="B28">
<label>28.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Oseledets</surname>
<given-names>IV</given-names>
</name>
</person-group>. <article-title>Tensor-Train Decomposition</article-title>. <source>SIAM J Sci Comput</source> (<year>2011</year>) <volume>33</volume>(<issue>5</issue>):<fpage>2295</fpage>&#x2013;<lpage>317</lpage>. <pub-id pub-id-type="doi">10.1137/090752286</pub-id> </citation>
</ref>
<ref id="B29">
<label>29.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Koenderink</surname>
<given-names>JJ</given-names>
</name>
<name>
<surname>Van Doorn</surname>
<given-names>AJ</given-names>
</name>
</person-group>. <article-title>The Structure of Locally Orderless Images</article-title>. <source>Int J Comput Vis</source> (<year>1999</year>) <volume>31</volume>(<issue>2</issue>):<fpage>159</fpage>&#x2013;<lpage>68</lpage>. <pub-id pub-id-type="doi">10.1023/a:1008065931878</pub-id> </citation>
</ref>
<ref id="B30">
<label>30.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Oron</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Bar-Hillel</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Levi</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Avidan</surname>
<given-names>S</given-names>
</name>
</person-group>. <article-title>Locally Orderless Tracking</article-title>. <source>Int J Comput Vis</source> (<year>2015</year>) <volume>111</volume>(<issue>2</issue>):<fpage>213</fpage>&#x2013;<lpage>28</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-014-0740-6</pub-id> </citation>
</ref>
<ref id="B31">
<label>31.</label>
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Miller</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Torchmps</article-title> (<year>2019</year>). <comment>Available from: <ext-link ext-link-type="uri" xlink:href="https://githubcom/jemisjoky/torchmps">https://githubcom/jemisjoky/torchmps</ext-link> (Accessed March 1, 2019)</comment>. </citation>
</ref>
<ref id="B32">
<label>32.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Fishman</surname>
<given-names>M</given-names>
</name>
<name>
<surname>White</surname>
<given-names>SR</given-names>
</name>
<name>
<surname>Stoudenmire</surname>
<given-names>EM</given-names>
</name>
</person-group>. <source>The Itensor Software Library for Tensor Network Calculations</source>. <publisher-loc>Ithaca, New York</publisher-loc>: <publisher-name>arXiv preprint arXiv:200714822</publisher-name> (<year>2020</year>). </citation>
</ref>
<ref id="B33">
<label>33.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Novikov</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Izmailov</surname>
<given-names>P</given-names>
</name>
<name>
<surname>Khrulkov</surname>
<given-names>V</given-names>
</name>
<name>
<surname>Figurnov</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Oseledets</surname>
<given-names>IV</given-names>
</name>
</person-group>. <article-title>Tensor Train Decomposition on Tensorflow (T3f)</article-title>. <source>J Mach Learn Res</source> (<year>2020</year>) <volume>21</volume>(<issue>30</issue>):<fpage>1</fpage>&#x2013;<lpage>7</lpage>. </citation>
</ref>
<ref id="B34">
<label>34.</label>
<citation citation-type="book">
<person-group person-group-type="editor">
<name>
<surname>Astuti</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Sediono</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Aibinu</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Akmeliawati</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Salami</surname>
<given-names>M-JE</given-names>
</name>
</person-group>, editors. <article-title>Adaptive Short Time Fourier Transform (Stft) Analysis of Seismic Electric Signal (Ses): A Comparison of Hamming and Rectangular Window</article-title>. In: <conf-name>Proceedings of the 2012 IEEE Symposium on Industrial Electronics and Applications</conf-name>; <conf-loc>Bandung, Indonesia</conf-loc>; <conf-date>2012 September 23</conf-date>. <publisher-name>IEEE</publisher-name> (<year>2012</year>). </citation>
</ref>
</ref-list>
</back>
</article>