<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Phys.</journal-id>
<journal-title>Frontiers in Physics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Phys.</abbrev-journal-title>
<issn pub-type="epub">2296-424X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">858388</article-id>
<article-id pub-id-type="doi">10.3389/fphy.2022.858388</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Physics</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Number-state preserving tensor networks as classifiers for supervised learning</article-title>
<alt-title alt-title-type="left-running-head">Evenbly</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fphy.2022.858388">10.3389/fphy.2022.858388</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Evenbly</surname>
<given-names>Glen</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1539642/overview"/>
</contrib>
</contrib-group>
<aff>
<institution>School of Physics</institution>, <institution>Georgia Institute of Technology</institution>, <addr-line>Atlanta</addr-line>, <addr-line>GA</addr-line>, <country>United States</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/570113/overview">Peng Zhang</ext-link>, Tianjin University, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/299880/overview">Sayantan Choudhury</ext-link>, Shree Guru Gobind Singh Tricentenary (SGT) University, India</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1698601/overview">Ding Liu</ext-link>, Tianjin Polytechnic University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1242409/overview">Shi-Ju Ran</ext-link>, Capital Normal University, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Glen Evenbly, <email>glen.evenbly@gmail.com</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Statistical and Computational Physics, a section of the journal Frontiers in Physics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>29</day>
<month>11</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>10</volume>
<elocation-id>858388</elocation-id>
<history>
<date date-type="received">
<day>19</day>
<month>01</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>25</day>
<month>10</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Evenbly.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Evenbly</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>We propose a restricted class of tensor network state, built from number-state preserving tensors, for supervised learning tasks. This class of tensor network is argued to be a natural choice for classifiers as 1) they map classical data to classical data, and thus preserve the interpretability of data under tensor transformations, 2) they can be efficiently trained to maximize their scalar product against classical data sets, and 3) they seem to be as powerful as generic (unrestricted) tensor networks in this task. Our proposal is demonstrated using a variety of benchmark classification problems, where number-state preserving versions of commonly used networks (including MPS, TTN and MERA) are trained as effective classifiers. This work opens the path for powerful tensor network methods such as MERA, which were previously computationally intractable as classifiers, to be employed for difficult tasks such as image recognition.</p>
</abstract>
<kwd-group>
<kwd>tensor network</kwd>
<kwd>machine learning</kwd>
<kwd>matrix product ansatz</kwd>
<kwd>quantum many body theory</kwd>
<kwd>numerical optimization</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Ideas and methods from the field of machine learning are currently having a significant impact in many areas of physics research [<xref ref-type="bibr" rid="B1">1</xref>]. Machine learning offers powerful new tools for classifying phases of matter [<xref ref-type="bibr" rid="B2">2</xref>&#x2013;<xref ref-type="bibr" rid="B7">7</xref>], for processing experimental results [<xref ref-type="bibr" rid="B8">8</xref>, <xref ref-type="bibr" rid="B9">9</xref>], and for modeling quantum many-body systems [<xref ref-type="bibr" rid="B10">10</xref>&#x2013;<xref ref-type="bibr" rid="B12">12</xref>], to name but a few of the plethora of applications. With this crossing of fields has come the intriguing realization that the neural networks [<xref ref-type="bibr" rid="B13">13</xref>, <xref ref-type="bibr" rid="B14">14</xref>] used in machine learning share extensive similarities with the tensor networks [<xref ref-type="bibr" rid="B15">15</xref>] used in modeling quantum many-body systems [<xref ref-type="bibr" rid="B16">16</xref>]. These connections are perhaps not so surprising since both types of network have the primary function of encoding large sets of correlated data: neural networks encode ensembles of training data, while tensor networks encode superpositions of quantum states. Currently there is great interest in exploring the potential applications of this relation, both from the directions of 1) using ideas from neural networks and machine learning to improve methods for modeling quantum wave-functions [<xref ref-type="bibr" rid="B17">17</xref>&#x2013;<xref ref-type="bibr" rid="B20">20</xref>] and 2) examining tensor networks as a new approach for tasks in machine learning [<xref ref-type="bibr" rid="B21">21</xref>&#x2013;<xref ref-type="bibr" rid="B31">31</xref>].</p>
<p>In this manuscript we focus on the second direction (ii), and explore the use of tensor networks as classifiers for supervised learning problems. Research in this area has already produced encouraging early results, with examples where tensor networks have been trained to produce relatively competitive classifiers in both supervised and unsupervised learning tasks [<xref ref-type="bibr" rid="B21">21</xref>, <xref ref-type="bibr" rid="B25">25</xref>&#x2013;<xref ref-type="bibr" rid="B27">27</xref>, <xref ref-type="bibr" rid="B30">30</xref>, <xref ref-type="bibr" rid="B31">31</xref>]. However there are some significant issues with respect to the use of tensor networks as classifiers. One such issue is that of <italic>interpretability</italic>. Usually, when applying a tensor network as a classifier, each sample from the (classical) dataset is associated to a product state. However, under generic tensor transformations, product states can be mapped to entangled quantum states, which can no longer be re-interpreted classically. One can understand this as a problem of generic tensor networks being overly-broad when used as classifiers: they are designed to carry information about phases and/or signs between superposition states, which are <italic>necessary</italic> for describing wave-functions but seem to be <italic>extraneous</italic> from the perspective of characterizing classical datasets. A second issue is that of computational efficiency. Most previous studies have utilized only relatively simple classes of tensor networks, such as matrix product states [<xref ref-type="bibr" rid="B32">32</xref>, <xref ref-type="bibr" rid="B33">33</xref>] (MPS) and tree tensor networks [<xref ref-type="bibr" rid="B34">34</xref>, <xref ref-type="bibr" rid="B35">35</xref>] (TTN), as classifiers. The more formidable weapons in the arsenal of tensor networks, such as the multi-scale entanglement renormalization ansatz [<xref ref-type="bibr" rid="B36">36</xref>&#x2013;<xref ref-type="bibr" rid="B39">39</xref>] (MERA), which are seen as the direct analogues to the high successful convolutional neural networks [<xref ref-type="bibr" rid="B40">40</xref>&#x2013;<xref ref-type="bibr" rid="B42">42</xref>] (CNNs), have yet to be deployed in earnest for challenging problems. The primary reason being that, in order for a tensor network to be of use as a classifier, ones needs to be able to compute scalar products between the network and product states (representing the training data); this can be done efficiently for simple networks such as MPS and TTN, but is generally computationally intractable for more sophisticated networks like MERA.</p>
<p>The main motivation for this manuscript is to help resolve the two issues discussed above. In particular, we propose to use networks built from a restricted class of tensor, those which act to preserve number-states, as classifiers for supervised learning tasks. Such number-state preserving networks automatically resolve the issue of interpretability, provided that each sample of the training data is encoded as a number state. Moreover, the restriction to number-state preserving tensors endows networks with a causal cone structure when contracted against number states, similar to the causal cone structure present in isometric networks when contracted against themselves. This property allows for a broad class of number-state preserving networks, including versions of MERA, to be efficiently trained as classifiers for supervised learning problems. Furthermore, we demonstrate numerically that networks built from this restricted class of number-state preserving tensor perform well for several example classification problems. The above considerations indicate that number-state preserving tensors are a natural restriction to impose when applying tensor methods to learn from sets of classical data.</p>
<p>This manuscript is organized as follows. Firstly in <xref ref-type="sec" rid="s2">Section 2</xref>, we characterize number-state preserving tensors and some of their properties, then in <xref ref-type="sec" rid="s3">Section 3</xref> we formulate how problems in supervised learning can be approached using tensor networks. In <xref ref-type="sec" rid="s4">Section 4</xref> we propose an algorithm for training number-state preserving tensor networks to correctly classify a labeled dataset, while <xref ref-type="sec" rid="s5">Section 5</xref> we describe how single tensor environments can be efficiently evaluated, a key ingredient in the proposed training algorithm. Benchmark numerical results for number-state preserving versions of MPS, TTN and MERA applied to example classification problems are presented in <xref ref-type="sec" rid="s6">Section 6</xref>, and conclusions are presented in <xref ref-type="sec" rid="s7">Section 7</xref>.</p>
</sec>
<sec id="s2">
<title>2 Number-state preserving networks</title>
<p>Let <inline-formula id="inf1">
<mml:math id="m1">
<mml:mi mathvariant="script">L</mml:mi>
</mml:math>
</inline-formula> be a lattice of sites, with each site described by a local Hilbert space of some dimension <italic>d</italic>. We label the basis states for each site by integers, &#x7c;<italic>z</italic>&#x27e9; &#x2208; {&#x7c;0&#x27e9;, &#x7c;1&#x27e9;, &#x2026;, &#x7c;<italic>d</italic> &#x2212; 1&#x27e9;}, which are interpreted as particle number and are represented as unit vectors,<disp-formula id="e1">
<mml:math id="m2">
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mtable class="matrix">
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>1</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mo>&#x22ee;</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
<mml:mspace width="0.28em"/>
<mml:mspace width="0.28em"/>
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mtable class="matrix">
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>1</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mo>&#x22ee;</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
<mml:mspace width="0.28em"/>
<mml:mspace width="0.28em"/>
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mtable class="matrix">
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>1</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mo>&#x22ee;</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
</mml:math>
<label>(1)</label>
</disp-formula>
</p>
<p>A number state <inline-formula id="inf2">
<mml:math id="m3">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="script">L</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
</inline-formula> (or, equivalently, a Fock state) on lattice <inline-formula id="inf3">
<mml:math id="m4">
<mml:mi mathvariant="script">L</mml:mi>
</mml:math>
</inline-formula> is a product state with well-defined particle number,<disp-formula id="e2">
<mml:math id="m5">
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="script">L</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x232a;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x232a;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x232a;</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
</mml:math>
<label>(2)</label>
</disp-formula>where superscripts are here used to denote lattice position. Alternatively, if one is thinking in terms of spin degrees of freedom, a number state can be defined as a product state with a well-defined <italic>z</italic>-component of spin.</p>
<p>We now turn our considerations to transformations of number-states implemented by certain types of <italic>oriented</italic> tensor: these are tensors where each index has been fixed as either <italic>incoming</italic> or <italic>outgoing</italic>. Any oriented tensor can be interpreted as a mapping between states defined on an input lattice <inline-formula id="inf4">
<mml:math id="m6">
<mml:mi mathvariant="script">L</mml:mi>
</mml:math>
</inline-formula>, whose sites match the incoming tensor indices, to states on an output lattice <inline-formula id="inf5">
<mml:math id="m7">
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula>, whose sites match the outgoing tensor indices. We define an oriented tensor as <italic>number-state preserving</italic> if it maps any number state defined on <inline-formula id="inf6">
<mml:math id="m8">
<mml:mi mathvariant="script">L</mml:mi>
</mml:math>
</inline-formula> to another number state on <inline-formula id="inf7">
<mml:math id="m9">
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula>. Several examples of number-state preserving tensors are given in <xref ref-type="fig" rid="F1">Figure 1</xref>. Let <inline-formula id="inf8">
<mml:math id="m10">
<mml:msubsup>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> be a four index tensor, with subscripts denoting incoming indices and superscripts denoting outgoing indices, as depicted in <xref ref-type="fig" rid="F1">Figure 1b</xref>. Consider the reshape of <italic>u</italic> into an input-output matrix, i.e., where the rows of the matrix enumerate over the tensor product (<italic>i</italic> &#x2297; <italic>j</italic>) of incoming indices and columns enumerate over the tensor product of the outgoing indices (<italic>k</italic> &#x2297; <italic>l</italic>). It is easily understood that the property of <italic>u</italic> being number-state preserving is equivalent to the property that each row of the corresponding input-output matrix must have <italic>at most</italic> a single non-zero entry. Note that we also include in the definition of number-state preserving tensors those where the input-output matrix has rows with only zero entries; equivalently these are tensors which can map some number states to the null (or norm-zero) state. An important property of number-state preserving tensors is that networks formed from their composition, where outputs from one tensor are properly matched with inputs to other tensors, are also number-state preserving, as depicted in <xref ref-type="fig" rid="F2">Figure 2A</xref>. This allows us to form number-state preserving versions of commonly used tensor networks, such as MERA, as shown in <xref ref-type="fig" rid="F2">Figure 2B</xref>. However, it is vital to realize that number-state preserving tensors do not necessarily remain number-state preserving if the orientation of their indices is reversed (i.e., the incoming and outgoing indices are switched); thus number-state preserving networks can still generate interesting superpositions and entangled states when &#x201c;run&#x201d; in reverse.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>(a) An example of a number-state preserving tensor that <italic>w</italic> maps a number state &#x27e8;<italic>z</italic>
<sup>0</sup>&#x7c;&#x27e8;<italic>z</italic>
<sup>1</sup>&#x7c; on its input indices to a number state <inline-formula id="inf9">
<mml:math id="m11">
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula> on its output index. The tensor <italic>w</italic> can be equivalently represented as (ii) an explicit mapping between number states or (iii) as a matrix (after forming the product of input indices). (b) An example of a number-state preserving tensor <italic>u</italic> between two input and two output indices. (c) An example of a number-state preserving tensor <italic>v</italic> between one input and two output indices. Note that the three examples of number-state preserving tensors from (a-c) are also <italic>unital</italic>, in that all of their non-zero entries are the unit element.</p>
</caption>
<graphic xlink:href="fphy-10-858388-g001.tif"/>
</fig>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>
<bold>(A)</bold> A number-state preserving network is formed through composition of number-state preserving tensors <italic>u</italic> and <italic>w</italic>, which maps input number state <inline-formula id="inf10">
<mml:math id="m12">
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:mi mathvariant="script">Z</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula> to output <inline-formula id="inf11">
<mml:math id="m13">
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula>. <bold>(B)</bold> A binary MERA tensor network <inline-formula id="inf12">
<mml:math id="m14">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula>, assumed to be composed of number-state preserving tensors, maps an input number state <inline-formula id="inf13">
<mml:math id="m15">
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>in</mml:mtext>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula> on a lattice of 24 sites to an output number state, <inline-formula id="inf14">
<mml:math id="m16">
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>in</mml:mtext>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="script">T</mml:mi>
<mml:mo>&#x21a6;</mml:mo>
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>out</mml:mtext>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula>, on a single site.</p>
</caption>
<graphic xlink:href="fphy-10-858388-g002.tif"/>
</fig>
<p>For the main text of this paper we shall further restrict our consideration to <italic>unital</italic> number-state preserving tensors, where each tensor entry must be either a zero or a one, and each row of the corresponding input-output matrix is required to have a single non-zero entry. Note that this class of tensor maps incoming number-states to outgoing number-states of the <italic>same</italic> normalization and phase. The restriction to unital tensors will be useful in simplifying their application to supervised learning problems, although the formalism and optimization algorithms that we present are still general for all number-state preserving networks. There are many reasons why one may also wish to consider networks comprised of non-unital number-state preserving tensors, where entries can take any real or complex value, and thus change the normalization of states and introduce phases; the interested reader is directed to <xref ref-type="sec" rid="s13">Supplementary Appendix SA</xref> for further discussion.</p>
<p>Given that number-state preserving networks represent a severely restricted class of tensor network states it may be interesting to consider how much of their power has been lost, for instance, in describing ground states of quantum many-body systems. Although this remains to be explored, it seems likely that majority of many-body systems will not have ground-states that can be well-approximated by number-state preserving tensor networks. However, there does exist several examples of non-trivial quantum many-body systems related to Motzkin paths [<xref ref-type="bibr" rid="B43">43</xref>], whose ground states possess interesting entanglement and yet can be exactly represented by number-state preserving networks [<xref ref-type="bibr" rid="B44">44</xref>, <xref ref-type="bibr" rid="B45">45</xref>]. Investigation of the ability of number-state preserving networks to describe general quantum ground states remains an intriguing direction for future research.</p>
</sec>
<sec id="s3">
<title>3 Supervised learning in a tensor product space</title>
<p>In this section we discuss how the task of supervised learning can be formulated in terms of tensor networks. We consider problems where each training sample <inline-formula id="inf15">
<mml:math id="m17">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> is represented as a length <italic>N</italic> vector, with the <italic>i</italic>th component <italic>z</italic>
<sup>
<italic>i</italic>
</sup> an element of <inline-formula id="inf16">
<mml:math id="m18">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> (the set of integers modulo <italic>d</italic>), i.e., such that<disp-formula id="e3">
<mml:math id="m19">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>Z</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(3)</label>
</disp-formula>where <italic>k</italic> is a label over the set of training samples. Every training sample is assumed to be paired with a corresponding label <inline-formula id="inf17">
<mml:math id="m20">
<mml:mi>y</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, where <italic>c</italic> represents the number of distinct categories for the classification problem. The goal of the supervised learning problem is to construct a function <italic>f</italic> that maps each sample of the training set to its correct label,<disp-formula id="e4">
<mml:math id="m21">
<mml:mi>f</mml:mi>
<mml:mo>:</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>Z</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x21a6;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>.</mml:mo>
</mml:math>
<label>(4)</label>
</disp-formula>
</p>
<p>Although classifiers based on linear functions <italic>f</italic> have some considerable utility [<xref ref-type="bibr" rid="B46">46</xref>], many non-trivial classification problems require non-linear functions <italic>f</italic> in order to achieve good accuracy.</p>
<p>We now describe how a tensor network can be implemented as the classifying function in <xref ref-type="disp-formula" rid="e4">Eq. 4</xref>. At this point, one could be tempted to believe that tensor networks would have limited utility as classifiers as, given that tensors simply are extensions of matrices to higher dimensions, they are inherently <italic>linear</italic> constructs. However, in order to recast the supervised learning problem into a problem amenable to tensor networks, we first (non-linearly) embed the training data into a higher dimensional space, similar to a kernel method [<xref ref-type="bibr" rid="B47">47</xref>]. By using an appropriate non-linear embedding, a linear classifier acting the higher dimension space can reproduce the classifying power of non-linear functions in the original space; thus it remains possible that tensor network approaches could be competitive with classifiers based on (non-linear) neural networks. Indeed, as will be argued later in this manuscript, it can be understood that a tensor network of sufficiently large bond dimension <italic>&#x3c7;</italic> can, in principle, obtain perfect accuracy for any training set of a supervised learning problem as formulated above.</p>
<p>Let us recast each training sample <inline-formula id="inf18">
<mml:math id="m22">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> as a number state, denoted <inline-formula id="inf19">
<mml:math id="m23">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
</inline-formula>, defined in a vector space of total dimension <italic>d</italic>
<sup>
<italic>N</italic>
</sup>. Specifically, we associate each integer <inline-formula id="inf20">
<mml:math id="m24">
<mml:mi>z</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="double-struck">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> with a number state &#x7c;<italic>z</italic>&#x27e9; in a <italic>d</italic>-dimensional Hilbert space, represented as per <xref ref-type="disp-formula" rid="e1">Eq. 1</xref>, such that the full state vector <inline-formula id="inf21">
<mml:math id="m25">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
</inline-formula> is given as the tensor product of the single site states,<disp-formula id="e5">
<mml:math id="m26">
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2026;</mml:mo>
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(5)</label>
</disp-formula>
</p>
<p>Similarly the data labels <italic>y</italic>
<sub>
<italic>k</italic>
</sub> are recast as number states &#x7c;<italic>y</italic>
<sub>
<italic>k</italic>
</sub>&#x27e9; in a <italic>c</italic>-dimensional space. The diagrammatic tensor notation for these states is presented in <xref ref-type="fig" rid="F3">Figure 3</xref>. Given this embedding of our training data, a classifier can be represented as tensor network <inline-formula id="inf22">
<mml:math id="m27">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> that maps states <inline-formula id="inf23">
<mml:math id="m28">
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>in</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula> from the lattice of <italic>N</italic> sites of dimension <italic>d</italic> to states <inline-formula id="inf24">
<mml:math id="m29">
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>out</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula> on a single site of dimension <italic>c</italic>,<disp-formula id="e6">
<mml:math id="m30">
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>in</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="script">T</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>out</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>,</mml:mo>
</mml:math>
<label>(6)</label>
</disp-formula>see also <xref ref-type="fig" rid="F2">Figure 2B</xref> for an explicit example.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>
<bold>(A)</bold> The <italic>k</italic>th training sample <inline-formula id="inf25">
<mml:math id="m31">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> is given as a length <italic>N</italic> vector of integers <italic>z</italic>
<sub>
<italic>k</italic>
</sub> (modulo some specified base <italic>d</italic>), and is accompanied by label <italic>y</italic>
<sub>
<italic>k</italic>
</sub>. <bold>(B)</bold> The training sample <inline-formula id="inf26">
<mml:math id="m32">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> can alternatively be expressed as a unit vector <inline-formula id="inf27">
<mml:math id="m33">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
</inline-formula> in the tensor product space of dimension <italic>d</italic>
<sup>
<italic>N</italic>
</sup> formed from mapping each base-<italic>d</italic> integer to a number-state &#x7c;<italic>z</italic>
<sub>
<italic>k</italic>
</sub>&#x27e9;, see <xref ref-type="disp-formula" rid="e1">Eq. 1</xref>. <bold>(C)</bold> Diagrammatic tensor representation of training sample <inline-formula id="inf28">
<mml:math id="m34">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
</inline-formula>.</p>
</caption>
<graphic xlink:href="fphy-10-858388-g003.tif"/>
</fig>
<p>In general, the accuracy of <inline-formula id="inf29">
<mml:math id="m35">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> as a classifier could be quantified by evaluating the scalar products of the output states with the label states, <inline-formula id="inf30">
<mml:math id="m36">
<mml:mrow>
<mml:mo stretchy="false">&#x27e8;</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>out</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">&#x27e9;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, where a large scalar product would indicate good classification. However, in the particular case of <italic>unital</italic> number-preserving networks <inline-formula id="inf31">
<mml:math id="m37">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula>, the norm of states is preserved such that all scalar products <inline-formula id="inf32">
<mml:math id="m38">
<mml:mrow>
<mml:mo stretchy="false">&#x27e8;</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>out</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">&#x27e9;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> either evaluate to unity (indicating correct classification of the data sample with label <italic>y</italic>
<sub>
<italic>k</italic>
</sub>) or to zero (indicating incorrect classification of the data sample). Thus, the number of correctly classified samples <italic>N</italic>
<sub>correct</sub> simply evaluates as the sum over all the scalar products,<disp-formula id="e7">
<mml:math id="m39">
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>correct</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mfenced open="&#x27e8;" close="|">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>in</mml:mtext>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mi mathvariant="script">T</mml:mi>
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(7)</label>
</disp-formula>
</p>
<p>The diagrammatic tensor notation for <xref ref-type="disp-formula" rid="e7">Eq. 7</xref>, in the particular case that <inline-formula id="inf33">
<mml:math id="m40">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> is a binary MERA, is presented in <xref ref-type="fig" rid="F4">Figure 4</xref>. It follows we should use <xref ref-type="disp-formula" rid="e7">Eq. 7</xref> as the <italic>cost function</italic> for training the tensor network <inline-formula id="inf34">
<mml:math id="m41">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> for the supervised learning problem: the tensors contained within <inline-formula id="inf35">
<mml:math id="m42">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> should be optimized as to maximize <italic>N</italic>
<sub>correct</sub>. Methods for achieving this are discussed in the following section of this manuscript.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>
<bold>(A)</bold> The total number correctly classified samples <italic>N</italic>
<sub>correct</sub> is given as the inner product of the labels &#x7c;<italic>y</italic>
<sub>
<italic>k</italic>
</sub>&#x27e9; against the network <inline-formula id="inf36">
<mml:math id="m43">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> applied to the training data <inline-formula id="inf37">
<mml:math id="m44">
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>in</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula>, summing over all training samples <italic>k</italic>. <bold>(B)</bold> Diagrammatic representation of the equation from <bold>(A)</bold> which evaluates to <italic>N</italic>
<sub>correct</sub>. <bold>(C)</bold> For any chosen tensor, such as the shaded tensor <italic>u</italic> in <bold>(B)</bold>, the network for <italic>N</italic>
<sub>correct</sub> can be factorized into a product of the tensor with its environment &#x393;<sub>
<italic>u</italic>
</sub>, formed from contracting the entirety of the network sans <italic>u</italic>. The environment &#x393;<sub>
<italic>u</italic>
</sub> allows the optimal tensor <italic>u</italic> that maximizes <italic>N</italic>
<sub>correct</sub> (with the other tensors in <inline-formula id="inf38">
<mml:math id="m45">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> held fixed) to be identified.</p>
</caption>
<graphic xlink:href="fphy-10-858388-g004.tif"/>
</fig>
<p>Before moving on, we remark that the formalism we described (or similar formalisms consider previously [<xref ref-type="bibr" rid="B21">21</xref>, <xref ref-type="bibr" rid="B23">23</xref>, <xref ref-type="bibr" rid="B26">26</xref>&#x2013;<xref ref-type="bibr" rid="B28">28</xref>, <xref ref-type="bibr" rid="B31">31</xref>]) for addressing supervised learning problems using tensor networks could, in principle, employ arbitrary tensor networks <inline-formula id="inf39">
<mml:math id="m46">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> as classifiers (not only those built from number-state preserving tensor networks). However, it is only for certain types of network, such as MPS and TTN, that scalar products of the form <inline-formula id="inf40">
<mml:math id="m47">
<mml:mfenced open="&#x27e8;" close="|">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>in</mml:mtext>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mi mathvariant="script">T</mml:mi>
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula> can be efficiently evaluated. The cost of (exactly) evaluating the overlap of a product state with a more sophisticated tensor network state, such as a MERA, typically does not scale efficiently with system size. Thus, one would expect that a general MERA network would only be computationally feasible as a classifier for problems with a small number of sites (or variables). In contrast the output state <inline-formula id="inf41">
<mml:math id="m48">
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>out</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula> of <xref ref-type="disp-formula" rid="e6">Eq. 6</xref> can be efficiently evaluated for <italic>any</italic> number-state preserving tensor network, with cost that scales only linearly in the number of tensors in <inline-formula id="inf42">
<mml:math id="m49">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula>. Nonetheless, the result that a scalar product <inline-formula id="inf43">
<mml:math id="m50">
<mml:mfenced open="&#x27e8;" close="|">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>in</mml:mtext>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mi mathvariant="script">T</mml:mi>
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula> is efficient to evaluate does not in itself imply that the network <inline-formula id="inf44">
<mml:math id="m51">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> can be efficiently trained. In <xref ref-type="sec" rid="s5">Section 5</xref> we formulate additional requirements for network <inline-formula id="inf45">
<mml:math id="m52">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> that are sufficient to allow for efficient training.</p>
</sec>
<sec id="s4">
<title>4 Single tensor updates</title>
<p>In this section we propose a method to optimize the tensors of a network <inline-formula id="inf46">
<mml:math id="m53">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> to maximize the number <italic>N</italic>
<sub>correct</sub> of correctly identified training samples in a supervised learning problem, as formulated in <xref ref-type="disp-formula" rid="e7">Eq. 7</xref>. We follow the same strategy of <italic>single tensor updates</italic> developed in the context optimizing MERA [<xref ref-type="bibr" rid="B48">48</xref>], where only a single tensor in the network is changed at any time while all other tensors in the network are held fixed. These single tensor updates can then be organized into &#x201c;sweeps,&#x201d; in which all tensors in the network are optimized in turn, and the sweeps iterated until the entire network is sufficiently converged.</p>
<p>Key to this optimization strategy is the notion of a <italic>tensor environment</italic>, which can be understood as the derivative of the network with respect to a single tensor. Specifically, given a network that evaluates to a scalar such as that from <xref ref-type="fig" rid="F4">Figure 4B</xref>, the environment &#x393;<sub>
<italic>u</italic>
</sub> of a tensor <italic>u</italic> results from contracting the entire network sans the particular tensor <italic>u</italic> under consideration. It follows that the number of correctly classified samples <italic>N</italic>
<sub>correct</sub> from <xref ref-type="disp-formula" rid="e7">Eq. 7</xref> can always be expressed as the scalar product of a tensor <inline-formula id="inf47">
<mml:math id="m54">
<mml:mi>u</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> with its environment &#x393;<sub>
<italic>u</italic>
</sub>,<disp-formula id="e8">
<mml:math id="m55">
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>correct</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x22c5;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">&#x393;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2020;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(8)</label>
</disp-formula>where, for notational simplicity, we have recast <italic>u</italic> and &#x393;<sub>
<italic>u</italic>
</sub> into input-output matrices, see <xref ref-type="fig" rid="F4">Figure 4C</xref>. We relegate a description of the general method for computing environments &#x393;<sub>
<italic>u</italic>
</sub> to <xref ref-type="sec" rid="s5">Section 5</xref> of the manuscript, and proceed here assuming &#x393;<sub>
<italic>u</italic>
</sub> is already known.</p>
<p>Let us now turn to the problem of finding the optimal number-state preserving tensor <italic>u</italic>
<sub>opt.</sub>,<disp-formula id="e9">
<mml:math id="m56">
<mml:msub>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>opt.</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2261;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mi>argmax</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mi mathvariant="normal">r</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x22c5;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">&#x393;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2020;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(9)</label>
</disp-formula>which maximizes the number of correctly identified samples <italic>N</italic>
<sub>correct</sub> of <xref ref-type="disp-formula" rid="e8">Eq. 8</xref>, given a known environment &#x393;<sub>
<italic>u</italic>
</sub>. Here it is easy to see that <italic>u</italic>
<sub>opt.</sub> can be built by simply identifying the location of the maximal element in each row of &#x393;<sub>
<italic>u</italic>
</sub> and then placing the unit element at the corresponding location in each row of <italic>u</italic>
<sub>opt.</sub>, with all other entries zero. Note that if the maximal element in a row of &#x393;<sub>
<italic>u</italic>
</sub> is degenerate then <italic>u</italic>
<sub>opt.</sub> is not uniquely defined; one can still obtain <italic>an</italic> optimal solution by simply selecting one of the maximal elements in that row of &#x393;<sub>
<italic>u</italic>
</sub>. Let us consider a concrete example: imagine we are updating a tensor <italic>u</italic> with a 4 &#xd7; 4 input-output matrix of the form given in <xref ref-type="fig" rid="F1">Figure 1b-iii</xref>, and assume that the environment has been evaluated as<disp-formula id="e10">
<mml:math id="m57">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">&#x393;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mtable class="matrix">
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>10</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>12</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>9</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>8</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>5</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>6</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>9</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>2</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>21</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>18</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>7</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>22</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>12</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>15</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>13</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>14</mml:mn>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(10)</label>
</disp-formula>
</p>
<p>Then the (unital and number-state preserving) 4 &#xd7; 4 matrix <italic>u</italic>
<sub>opt.</sub> that maximizes <xref ref-type="disp-formula" rid="e8">Eq. 8</xref> is given as<disp-formula id="e11">
<mml:math id="m58">
<mml:msub>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>opt.</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mtable class="matrix">
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>1</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>1</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>1</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>1</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(11)</label>
</disp-formula>and the number of correctly classified training samples after this optimal update is given as <italic>N</italic>
<sub>correct</sub> &#x3d; (12 &#x2b; 9 &#x2b; 22 &#x2b; 15) &#x3d; 58. Some remarks are in order regarding this optimization strategy. Firstly, we notice that unlike many commonly used algorithms for training neural networks, our approach is not based upon a gradient descent. Instead we can directly &#x201c;hop&#x201d; to the true maximum for any single tensor (given that the other tensors in the network are held remain fixed), provided the environment is exactly known. While this strategy has some advantages over gradient based methods with respect to avoiding local maxima, getting stuck in a solution that is not globally optimal can still remain a possibility depending on the problem until consideration.</p>
<p>We now discuss methods to introduce some randomness into the optimization, in order to reduce the possibility of getting trapped in a local maxima. One approach could be to employ a similar strategy as used in the stochastic gradient descent methods [<xref ref-type="bibr" rid="B49">49</xref>], where randomness is introduced by using only select &#x201c;batch&#x201d; of training samples for each update. Instead, here we advocate a different strategy inspired by Monte Carlo methods [<xref ref-type="bibr" rid="B50">50</xref>] used in sampling many-body systems. Rather than updating to the optimal tensor <italic>u</italic>
<sub>opt.</sub> at each step, we propose to allow updates to sub-optimal solutions of <xref ref-type="disp-formula" rid="e8">Eq. 8</xref>, with a probability diminishes exponentially in relation to how far the solution is from the optimal solution. For this purpose we first introduce the difference matrix &#x3a9;, given by subtracting from each row of &#x393; the maximal element within the row,<disp-formula id="e12">
<mml:math id="m59">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a9;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">&#x393;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mi>max</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">&#x393;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(12)</label>
</disp-formula>
</p>
<p>For the example environment &#x393;<sub>
<italic>u</italic>
</sub> given in <xref ref-type="disp-formula" rid="e10">Eq. 10</xref> the corresponding difference matrix is<disp-formula id="e13">
<mml:math id="m60">
<mml:mi mathvariant="normal">&#x3a9;</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mtable class="matrix">
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>2</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>3</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>4</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>4</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>3</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>7</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>1</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>4</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>15</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>3</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>2</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>1</mml:mn>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(13)</label>
</disp-formula>
</p>
<p>We then use the difference matrix to generate a matrix <italic>p</italic>
<sup>trans.</sup> of transition probabilities, defined element-wise as<disp-formula id="e14">
<mml:math id="m61">
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>trans.</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a9;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a9;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfrac>
</mml:math>
<label>(14)</label>
</disp-formula>where <italic>&#x3b1;</italic> is a tunable parameter that sets the amount of randomness. For the example difference matrix &#x3a9; of <xref ref-type="disp-formula" rid="e13">Eq. 13</xref> and setting <italic>&#x3b1;</italic> &#x3d; 2 we get the transition matrix<disp-formula id="e15">
<mml:math id="m62">
<mml:msup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>trans.</mml:mtext>
</mml:mrow>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mtable class="matrix">
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0.21</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0.58</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0.13</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0.08</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0.10</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0.16</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0.72</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0.02</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0.35</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0.08</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0.00</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0.57</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0.10</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0.45</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0.17</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mn>0.28</mml:mn>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(15)</label>
</disp-formula>
</p>
<p>The transition matrix is then used to perform a stochastic update of the tensor <italic>u</italic> under consideration: values in each row of <italic>p</italic>
<sup>trans.</sup> set the probability for the unit element in the equivalent row of the updated <italic>u</italic> to be placed at that particular location (note that <xref ref-type="disp-formula" rid="e14">Eq. 14</xref> has been defined such that each row of <italic>p</italic>
<sup>trans</sup> sums to unit probability). Notice that in the limit <italic>&#x3b1;</italic> &#x2192; 0 the matrix <italic>p</italic>
<sup>trans.</sup> tends to <italic>u</italic>
<sub>opt.</sub> (provided &#x393; had no degeneracies in its maximal row values), since all non-optimal transitions are fully suppressed. Conversely, in limit <italic>&#x3b1;</italic> &#x2192; <italic>&#x221e;</italic> all probabilities in <italic>p</italic>
<sup>trans.</sup> tend to the same value, representing completely random transition probabilities.</p>
</sec>
<sec id="s5">
<title>5 Evaluation of tensor environments</title>
<p>Here we describe evaluation of tensor environments, crucial to the optimization algorithm discussed in the previous section. For simplicity, we describe this evaluation assuming the tensor network <inline-formula id="inf48">
<mml:math id="m63">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> under consideration is a binary MERA, although the same methodology can be employed for arbitrary (number-state preserving) tensor networks.</p>
<p>Rather than tackling the problem of computing tensor environments &#x393; directly, we first introduce the concept of configuration spaces &#x7c;<italic>&#x3d5;</italic>&#x27e9;. Proper use of configuration spaces &#x7c;<italic>&#x3d5;</italic>&#x27e9;, which play an analogous role to the local reduced density matrices <italic>&#x3c1;</italic> used to optimize tensor networks in the context of quantum many-body systems, will greatly simplify the subsequent evaluation of environments. Let us assume that the output index of the tensor network <inline-formula id="inf49">
<mml:math id="m64">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> under consideration has been fixed in some specified label state &#x7c;<italic>y</italic>&#x27e9;, and that the lattice on which it is defined has been partitioned into a region <italic>A</italic> and its compliment <italic>B</italic>. Then, given a number state <inline-formula id="inf50">
<mml:math id="m65">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
</inline-formula> on region <italic>A</italic>, we define the configuration space &#x7c;<italic>&#x3d5;</italic>
<sup>
<italic>B</italic>
</sup>&#x27e9; as<disp-formula id="e16">
<mml:math id="m66">
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">c</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mi mathvariant="normal">fi</mml:mi>
<mml:mi mathvariant="normal">g</mml:mi>
<mml:mi mathvariant="normal">s</mml:mi>
<mml:mo>:</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(16)</label>
</disp-formula>where the sum runs over all valid configurations <italic>&#x3c3;</italic> of number states <inline-formula id="inf51">
<mml:math id="m67">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
</inline-formula> defined on region <italic>B</italic> such that the combined number state <inline-formula id="inf52">
<mml:math id="m68">
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula> is classified by <inline-formula id="inf53">
<mml:math id="m69">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> into the correct category &#x7c;<italic>y</italic>&#x27e9;, i.e., such that<disp-formula id="e17">
<mml:math id="m70">
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mfenced open="&#x27e8;" close="|">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mfenced open="&#x27e8;" close="|">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mi mathvariant="script">T</mml:mi>
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>.</mml:mo>
</mml:math>
<label>(17)</label>
</disp-formula>
</p>
<p>An example of a network that could be contracted to evaluate a configuration space &#x7c;<italic>&#x3d5;</italic>
<sup>
<italic>B</italic>
</sup>&#x27e9; is depicted in <xref ref-type="fig" rid="F5">Figure 5A</xref>. It is seen that this network can be simplified, as shown <xref ref-type="fig" rid="F5">Figure 5B</xref>, by lifting the input number state <inline-formula id="inf54">
<mml:math id="m71">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
</inline-formula> through tensors in <inline-formula id="inf55">
<mml:math id="m72">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> where-ever possible (i.e., where-ever a tensor has a number state available on all of its incoming indices), using the number-state preserving tensor properties as outlined in <xref ref-type="fig" rid="F1">Figure 1</xref>. It is convenient to define the <italic>configuration</italic> causal cone <inline-formula id="inf56">
<mml:math id="m73">
<mml:mi mathvariant="script">C</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> associated to region <italic>B</italic> as the set of tensors remaining in the network <inline-formula id="inf57">
<mml:math id="m74">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> after this simplification; equivalently <inline-formula id="inf58">
<mml:math id="m75">
<mml:mi mathvariant="script">C</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> can be defined as the set of tensors <inline-formula id="inf59">
<mml:math id="m76">
<mml:mi mathvariant="script">C</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> whose output state can be affected by the choice of input state on region <italic>B</italic>.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>
<bold>(A)</bold> The network <inline-formula id="inf60">
<mml:math id="m77">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> with fixed output label &#x7c;<italic>y</italic>&#x27e9; is applied to a number state <inline-formula id="inf61">
<mml:math id="m78">
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula> defined only on a sub-region <italic>A</italic> of the initial lattice, with the state on the complimentary region <italic>B</italic> left open. <bold>(B)</bold> The input number-state <inline-formula id="inf62">
<mml:math id="m79">
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula> is lifted through <inline-formula id="inf63">
<mml:math id="m80">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> as much as is possible by using the number-state mapping properties depicted in <xref ref-type="fig" rid="F1">Figure 1</xref>. The (configuration) causal cone <inline-formula id="inf64">
<mml:math id="m81">
<mml:mi mathvariant="script">C</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> associated to region <italic>B</italic> describes the remaining set of tensors <inline-formula id="inf65">
<mml:math id="m82">
<mml:mi mathvariant="script">C</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> after this lifting; this is equivalently the set of tensors whose output states can be affected by the choice of input state on region <italic>B</italic>.</p>
</caption>
<graphic xlink:href="fphy-10-858388-g005.tif"/>
</fig>
<p>Notice that this configuration causal cone <inline-formula id="inf66">
<mml:math id="m83">
<mml:mi mathvariant="script">C</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> is precisely equivalent to the (standard) causal cone [<xref ref-type="bibr" rid="B36">36</xref>, <xref ref-type="bibr" rid="B51">51</xref>] that would emerge from an isometric MERA for the same region <italic>B</italic>, defined as the set of tensors that can affect the local reduced density matrix <italic>&#x3c1;</italic>
<sub>
<italic>B</italic>
</sub>. However, the origins of these causal cones are drastically different: the causal cones in isometric MERA result arise due to the isometric constraints imposed on tensors, whereas the number-state preserving tensors proposed in this manuscript are not required to be isometric. Similarly, configuration causal cones arise only in networks that preserve number states, and are thus ill-defined for generic MERA [Note that it is, however, possible to have networks with tensors that are both simultaneously isometric and number-state preserving, see <xref ref-type="sec" rid="s13">Supplementary Appendix SA</xref> for further discussion]. Despite the difference in the origins of these two forms of causal cone, it is not a fluke that they were exactly equivalent in the previous example. It can be understood that the configuration causal cones in any number-state preserving tensor network are always equivalent to the causal cones found in an isometric tensor network of the same geometry, provided that the index orientations (specifying incoming and outgoing indices) match between the networks. Given this equivalence, we will henceforth drop the distinction between the two definitions, such that the term &#x201c;causal cone&#x201d; can refer to either definition.</p>
<p>The process of evaluating the configuration space for a region <italic>B</italic> of three sites from a binary MERA is depicted in <xref ref-type="fig" rid="F6">Figure 6A</xref>. This evaluation can be formulated as a sequence of contractions that each &#x201c;lower&#x201d; the configuration space through the causal cone,<disp-formula id="e18">
<mml:math id="m84">
<mml:mo>&#x2026;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x232a;</mml:mo>
<mml:mo>&#x2192;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x232a;</mml:mo>
<mml:mo>&#x2192;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
<label>(18)</label>
</disp-formula>where bracketed subscripts denote configuration spaces at different depths within the network. Each of the lowering contractions is implemented by one of two geometrically different lowering operators, depicted in <xref ref-type="fig" rid="F6">Figure 6B</xref>, which are the direct analogues to the descending superoperators [<xref ref-type="bibr" rid="B48">48</xref>] used in the evaluation of density matrices from isometric MERA.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>
<bold>(A)</bold> Sequence of contractions used to evaluate the configuration space <inline-formula id="inf67">
<mml:math id="m85">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
</inline-formula> associated to region <italic>B</italic>, starting from the causal cone <inline-formula id="inf68">
<mml:math id="m86">
<mml:mi mathvariant="script">C</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> as depicted in <xref ref-type="fig" rid="F5">Figure 5B</xref>. At each step in the evaluation the tensors in shaded region are contracted into a single tensor. <bold>(B)</bold> For any region <italic>B</italic> of three contiguous sites on the initial lattice, the configuration space <inline-formula id="inf69">
<mml:math id="m87">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
</inline-formula> can be evaluated using a composition of the left/right lowering operators.</p>
</caption>
<graphic xlink:href="fphy-10-858388-g006.tif"/>
</fig>
<p>In our example using a binary MERA, the cost of evaluating <inline-formula id="inf70">
<mml:math id="m88">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>B</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
</inline-formula> for a region <italic>B</italic> of three contiguous sites scales at most linearly with the network depth, since the form of the lowering operators are self-similar at all depths. In a general (number-state) preserving network the computational cost of evaluating configuration spaces will be related to the causal structure of the network: the leading order cost will scale exponentially with maximum width of the causal cones. Thus it is apparent that not all number-state preserving tensor networks can be efficiently evaluated for local information (characterized by the configuration space &#x7c;<italic>&#x3d5;</italic>
<sub>
<italic>B</italic>
</sub>&#x27e9;); only those for which the maximum causal width is not too large. However, since MERA are precisely designed to have bounded causal width (i.e., the causal width never spreads beyond some small number of sites), it follows that number-state preserving versions of MERA networks precisely fall within the class of networks that can be efficiently evaluated.</p>
<p>Given that the evaluation of configuration spaces has been understood, we now turn to the task of building the environment &#x393;<sub>
<italic>u</italic>
</sub> associated to tensor <italic>u</italic>, as depicted in <xref ref-type="fig" rid="F7">Figure 7</xref>, which is accomplished as follows. First we lift the initial number state <inline-formula id="inf71">
<mml:math id="m89">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>in</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
</inline-formula> to a new number state <inline-formula id="inf72">
<mml:math id="m90">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
</inline-formula> that lives on the boundary of causal cone <inline-formula id="inf73">
<mml:math id="m91">
<mml:mi mathvariant="script">C</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> associated to tensor <italic>u</italic>, as depicted in <xref ref-type="fig" rid="F7">Figure 7A</xref>. Then we compute the configuration space <inline-formula id="inf74">
<mml:math id="m92">
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula> defined on the output indices of tensor <italic>u</italic>, as depicted in <xref ref-type="fig" rid="F7">Figure 7B</xref>. Then the environment &#x393;<sub>
<italic>u</italic>
</sub> is given by taking the outer product of the configuration space <inline-formula id="inf75">
<mml:math id="m93">
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula> with the piece of the state <inline-formula id="inf76">
<mml:math id="m94">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
</inline-formula> supported on the input indices of <italic>u</italic>, denoted <inline-formula id="inf77">
<mml:math id="m95">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
</inline-formula>, while summing over all training samples <italic>k</italic>,<disp-formula id="e19">
<mml:math id="m96">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">&#x393;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mfenced open="&#x27e8;" close="|">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(19)</label>
</disp-formula>see also <xref ref-type="fig" rid="F7">Figure 7C</xref>.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>The sequence of steps used to evaluate the environment &#x393;<sub>
<italic>u</italic>
</sub> of the shaded tensor <italic>u</italic>. <bold>(A)</bold> The initial state <inline-formula id="inf78">
<mml:math id="m97">
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>in</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula> is transformed through the network to form a new number state on the boundary of the causal cone <inline-formula id="inf79">
<mml:math id="m98">
<mml:mi mathvariant="script">C</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> associated to <italic>u</italic>. <bold>(B)</bold> The configuration space <inline-formula id="inf80">
<mml:math id="m99">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
</inline-formula>, defined on the output indices of <italic>u</italic>, is computed through use of the left/right lowering operators, as in <xref ref-type="fig" rid="F6">Figure 6</xref>. <bold>(C)</bold> The environment &#x393; is formed by taking the outer product of the configuration space <inline-formula id="inf81">
<mml:math id="m100">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>&#x3d5;</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x232a;</mml:mo>
</mml:math>
</inline-formula> with the state <inline-formula id="inf82">
<mml:math id="m101">
<mml:mo stretchy="false">&#x2329;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:math>
</inline-formula> defined the input of <italic>u</italic>, summing over all training samples <italic>k</italic>, see also <xref ref-type="disp-formula" rid="e19">Eq. 19</xref>.</p>
</caption>
<graphic xlink:href="fphy-10-858388-g007.tif"/>
</fig>
</sec>
<sec id="s6">
<title>6 Benchmark results</title>
<p>In this section we present benchmark results for how number-state preserving tensor networks perform as classifiers in some simple problems. The goal here is to establish the feasibility of our proposal, rather than to establish performance for challenging real-world tasks, which will be considered in future work. In particular we demonstrate 1) that the proposed optimization algorithms can efficiently and reliably train the networks under consideration, and 2) that number-preserving networks perform comparably well to unrestricted networks for classification tasks.</p>
<sec id="s6-1">
<title>6.1 Parity classification</title>
<p>For this first test, we benchmark the performance of a number-state preserving MPS for classifying the parity of binary strings. Here each test sample is a length-<italic>N</italic> binary vector <inline-formula id="inf83">
<mml:math id="m102">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>0,0,1,0,1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, which is labeled <italic>y</italic>
<sub>
<italic>k</italic>
</sub> &#x2208; {0, 1} according to its parity. The MPS that we use is depicted in <xref ref-type="fig" rid="F8">Figure 8</xref>, and is built from tensors that are number-state preserving only when acting from left-to-right. In this problem, we are free to choose the length <italic>N</italic> of the binary strings as well as the number <italic>n</italic>
<sub>samp.</sub> of training samples to use (as these can be randomly generated). We also have two hyper-parameters associated to our method: the maximal bond dimension <italic>&#x3c7;</italic>
<sub>max</sub> of the MPS and the parameter <italic>&#x3b1;</italic> from <xref ref-type="disp-formula" rid="e14">Eq. 14</xref> that controls the amount of randomness in the optimization. For each set of parameters investigated we performed 100 trial runs, each run starting with a randomly generated training set and a randomly initialized MPS, and then performed at no more than 100 optimization sweeps in each trial. The most computationally demanding trials (which consisted of: a length <italic>N</italic> &#x3d; 20 chain, <italic>n</italic>
<sub>samp.</sub> &#x3d; 20,000 training samples, a bond dimension of <italic>&#x3c7;</italic>
<sub>max</sub> &#x3d; 10, and 100 optimization sweeps) each took about 5&#xa0;s to run on a single 3&#xa0;GHz desktop CPU. At the end of each trial we also test the generalization error of the MPS classifier by evaluating its accuracy in classifying the parity of all possible 2<sup>
<italic>N</italic>
</sup> binary strings.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>
<bold>(A)</bold> Tensor <italic>v</italic> is a number-state preserving tensor mapping from two indices to a single index. <bold>(B)</bold> An MPS network <inline-formula id="inf84">
<mml:math id="m103">
<mml:mi mathvariant="script">T</mml:mi>
</mml:math>
</inline-formula> is built from tensors <italic>v</italic> that preserve number-states when mapping from left-to-right. The MPS is trained as a classifier by maximizing the scalar product <inline-formula id="inf85">
<mml:math id="m104">
<mml:msub>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="&#x27e8;" close="|">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>in</mml:mtext>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mi mathvariant="script">T</mml:mi>
<mml:mfenced open="|" close="&#x27e9;">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula>.</p>
</caption>
<graphic xlink:href="fphy-10-858388-g008.tif"/>
</fig>
<p>A summary of the results from a large number of trials is presented in <xref ref-type="table" rid="T1">Table 1</xref>. For binary strings of length <italic>N</italic> &#x3d; 16 and <italic>N</italic> &#x3d; 20 we used 1,300 and 20,000 training samples respectively; these numbers were chosen as they represent about 2% of all possible binary strings in each case (of which there are 2<sup>
<italic>N</italic>
</sup> in total). The randomness parameter was fixed at <italic>&#x3b1;</italic> &#x3d; 1 for <italic>N</italic> &#x3d; 16 and <italic>&#x3b1;</italic> &#x3d; 5 for <italic>N</italic> &#x3d; 20 length chains; these values were determined as adequate through small amount of experimentation (and are probably not those which would give optimal performance). Somewhat surprisingly, we found that each trial would produce only one of two outcomes: 1) the optimization would fail completely, achieving only slightly over 50% classification accuracy on the set of all binary strings, or 2) would converge to a perfect parity classifier, with 100% classification accuracy for all length-<italic>N</italic> binary strings. From <xref ref-type="table" rid="T1">Table 1</xref> we see the proportion <italic>n</italic>
<sub>perfect</sub> of perfect classifiers obtained increases dramatically as the bond dimension <italic>&#x3c7;</italic>
<sub>max</sub> was increased, reaching 96/100 for <italic>N</italic> &#x3d; 20 and <italic>&#x3c7;</italic>
<sub>max</sub> &#x3d; 10. This is expected, as networks with more degrees of freedom are less likely to be trapped in local minima. We found that the likelihood of obtaining a perfect classifier was also greatly improved when using a larger number of training samples, although do not provide this data here. In a recent work by Stokes and Terilla [<xref ref-type="bibr" rid="B52">52</xref>] standard (unrestricted) MPS were also trained to classify the parity of binary strings, and produced comparable results for similar strings lengths and training set sizes. This is a good indication that, for this classification problem, number-state preserving MPS are as powerful as unrestricted MPS.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Summary of results for MPS applied to the parity classification (above) and division-by-7 classification (below). Parameters are as follows: <italic>N</italic> is the length of binary strings classified, <italic>n</italic>
<sub>samp</sub> is the number of samples in the training set, <italic>&#x3c7;</italic>
<sub>max</sub> is the maximal MPS bond dimension, parameter <italic>&#x3b1;</italic> controls the randomness in the optimization as per <xref ref-type="disp-formula" rid="e14">Eq. 14</xref>, <italic>n</italic>
<sub>perfect</sub> is the proportion of trial runs that yielded perfect (100% accuracy) classifiers, <italic>n</italic>
<sub>sweeps</sub> is the average number of variational sweeps required to reach convergence.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th colspan="6" align="left">Parity classification</th>
</tr>
<tr>
<th align="left">
<italic>N</italic>
</th>
<th align="left">
<italic>n</italic>
<sub>samp</sub>
</th>
<th align="left">
<italic>&#x3c7;</italic>
<sub>max</sub>
</th>
<th align="left">
<italic>&#x3b1;</italic>
</th>
<th align="left">
<italic>n</italic>
<sub>perfect</sub>
</th>
<th align="left">
<italic>n</italic>
<sub>sweeps</sub>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">16</td>
<td align="char" char=".">1,300</td>
<td align="char" char=".">4</td>
<td align="char" char=".">1</td>
<td align="char" char="/">38/100</td>
<td align="char" char=".">31</td>
</tr>
<tr>
<td align="left">16</td>
<td align="char" char=".">1,300</td>
<td align="char" char=".">6</td>
<td align="char" char=".">1</td>
<td align="char" char="/">63/100</td>
<td align="char" char=".">28</td>
</tr>
<tr>
<td align="left">16</td>
<td align="char" char=".">1,300</td>
<td align="char" char=".">10</td>
<td align="char" char=".">1</td>
<td align="char" char="/">93/100</td>
<td align="char" char=".">25</td>
</tr>
<tr>
<td align="left">20</td>
<td align="char" char=".">20,000</td>
<td align="char" char=".">4</td>
<td align="char" char=".">5</td>
<td align="char" char="/">34/100</td>
<td align="char" char=".">26</td>
</tr>
<tr>
<td align="left">20</td>
<td align="char" char=".">20,000</td>
<td align="char" char=".">6</td>
<td align="char" char=".">5</td>
<td align="char" char="/">63/100</td>
<td align="char" char=".">21</td>
</tr>
<tr>
<td align="left">20</td>
<td align="char" char=".">20,000</td>
<td align="char" char=".">10</td>
<td align="char" char=".">5</td>
<td align="char" char="/">96/100</td>
<td align="char" char=".">27</td>
</tr>
<tr>
<td colspan="6" align="left">Division-by-7 Classification</td>
</tr>
<tr>
<td align="left">16</td>
<td align="char" char=".">3,000</td>
<td align="char" char=".">9</td>
<td align="char" char=".">1</td>
<td align="char" char="/">92/100</td>
<td align="char" char=".">43</td>
</tr>
<tr>
<td align="left">16</td>
<td align="char" char=".">3,000</td>
<td align="char" char=".">12</td>
<td align="char" char=".">1</td>
<td align="char" char="/">100/100</td>
<td align="char" char=".">36</td>
</tr>
<tr>
<td align="left">16</td>
<td align="char" char=".">3,000</td>
<td align="char" char=".">16</td>
<td align="char" char=".">1</td>
<td align="char" char="/">98/100</td>
<td align="char" char=".">29</td>
</tr>
<tr>
<td align="left">20</td>
<td align="char" char=".">30,000</td>
<td align="char" char=".">9</td>
<td align="char" char=".">5</td>
<td align="char" char="/">75/100</td>
<td align="char" char=".">56</td>
</tr>
<tr>
<td align="left">20</td>
<td align="char" char=".">30,000</td>
<td align="char" char=".">12</td>
<td align="char" char=".">5</td>
<td align="char" char="/">88/100</td>
<td align="char" char=".">44</td>
</tr>
<tr>
<td align="left">20</td>
<td align="char" char=".">30,000</td>
<td align="char" char=".">16</td>
<td align="char" char=".">5</td>
<td align="char" char="/">96/100</td>
<td align="char" char=".">26</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s6-2">
<title>6.2 Division-by-7 classification</title>
<p>For the second test we classify binary strings, interpreted as a base-2 representation of an integer, by their remainder under division by 7. We again use a number-state preserving MPS, employing the same set-up as used for the parity classification considered previously. A key difference here is that the samples now take one of seven different labels, <italic>y</italic>
<sub>
<italic>k</italic>
</sub> &#x2208; {0, 1, 2, 3, 4, 5, 6}.</p>
<p>A summary of the results from these trials is presented in <xref ref-type="table" rid="T1">Table 1</xref>. For binary strings of length <italic>N</italic> &#x3d; 16 and <italic>N</italic> &#x3d; 20 we used 3,000 and 30,000 training samples, respectively; although this was more than was used for the parity classification it is still less than 5% of the possible binary strings. Similar to the parity benchmark, we here found that each trial would either fail completely, producing no better than a random results, or would converge to a perfect division classifier, with 100% classification accuracy for all length-<italic>N</italic> binary strings. As with the parity benchmark, it is seen that the proportion of perfect classifiers obtained increases steadily with the bond dimension <italic>&#x3c7;</italic>
<sub>max</sub>. However, this problem required larger dimensions <italic>&#x3c7;</italic>
<sub>max</sub> than used for the parity benchmark, which is expected since here we have many more classification categories.</p>
</sec>
<sec id="s6-3">
<title>6.3 Height classification</title>
<p>The final test problem that we consider, which we refer to as height classification, takes length-<italic>N</italic> strings of integers from the set <italic>z</italic> &#x2208; { &#x2212; 1, 0, 1} and classifies them with labels <italic>y</italic>
<sub>
<italic>k</italic>
</sub> &#x2208; {0, 1, 2} depending on whether the sum (under regular addition) of the integers is positive, zero or negative, respectively. We test the effectiveness of both number-state preserving binary TTN and binary MERA as classifiers for this problem, working with strings of length <italic>N</italic> &#x3d; 24. A binary MERA of the form depicted in <xref ref-type="fig" rid="F2">Figure 2B</xref> is used, and is compared with the binary TTN that would result from restricting to trivial disentanglers <italic>u</italic> throughout the MERA network. Given that the problem is translation-invariant, we imposed that all tensors within a network layer are identical. In terms of the optimization, this is achieved by updating using the average single-tensor environment from all equivalent tensors within a network layer. We found that the injection of randomness into the optimization was unnecessary, possibly due to the imposition of translational invariance, such that the randomness parameter <italic>&#x3b1;</italic> from <xref ref-type="disp-formula" rid="e14">Eq. 14</xref> could be set at <italic>&#x3b1;</italic> &#x3d; 0. This left the bond dimension of the networks as the only hyper-parameter in the calculation, which was fixed at maximum dimension <italic>&#x3c7;</italic>
<sub>max</sub> &#x3d; 9.</p>
<p>The benchmark results are displayed in <xref ref-type="fig" rid="F9">Figure 9</xref>, and consisted of 100 trials, each trial starting from 12,000 randomly generated training samples (with 4,000 samples from each label category) and a randomly initialized network. Rather than running separate TTN and MERA trials they were instead combined: the first 20 sweeps were performed with trivial disentanglers <italic>u</italic>, such that underlying the network was a TTN, the <italic>u</italic> were then &#x201c;switched on&#x201d; for the remaining 40 sweeps such that the network became a MERA. At the conclusion of each trial, the generalization error was estimated by applying the trained classifiers to a randomly generated test set of the same size as the training set. Most of the trials converged smoothly, with the proportion of wrongly identified testing samples decreasing monotonically with optimization, although about 5 trials failed to properly converge (yielding classifiers with greater than 30% error). Discarding the worst 10 trials from consideration, of the 90 remaining trials the TTN gave average training/test errors of 14.15% and 14.91%, while MERA gave substantially reduced average training/test errors of 1.13% and 1.86%. These results clearly demonstrate the extra representation power endowed through use of the disentanglers <italic>u</italic> in MERA. Impressive is that both networks generalized well, with only relatively small differences between test and training accuracies, despite being trained on less than 5 &#xd7; 10<sup>&#x2013;6</sup> percent of the possible 3<sup>24</sup> training samples.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>(left) Results of training TTN and MERA for the height classification problem, displaying how much of training set is wrongly classified as a function of the number of optimization sweeps performed. The first 20 sweeps are performed while keeping trivial disentanglers <italic>u</italic>, such that underlying the network is a TTN, while the <italic>u</italic> are then &#x201c;switched on&#x201d; for the remaining sweeps such that the network becomes a MERA. The figure displays results from 10 different trials, where each trial starts with a randomly generated training set and randomly initialized network. (right) Average results of the training data from 100 trial runs (after discarding the 10 worst trials). Dashed lines show the average generalization error computed from applying the trained TTN and MERA applied to a randomly generated test set. For TTN we get average training/test errors of 14.15% and 14.91%, while for MERA we get average training/test errors of 1.13% and 1.86%.</p>
</caption>
<graphic xlink:href="fphy-10-858388-g009.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="conclusion" id="s7">
<title>7 Conclusion</title>
<p>We have proposed the class of number-state preserving tensor networks for use as classifiers in supervised learning tasks and have shown that a large class of these networks, specifically those with bounded causal structure, are efficiently trainable for large problems. In particular we have described a training algorithm that, for any chosen tensor in the network under consideration, exactly identifies the optimal tensor for that location (i.e., that which maximizes the number of correctly classified training samples), all with cost that scales only linearly in number of training samples. Importantly, the class of efficiently trainable number-state preserving networks includes realizations of sophisticated networks such as MERA, which would otherwise be computationally intractable. As such, we believe this could be the first computationally viable proposal which would allow MERA, close tensor network analogues to convolutional neural networks, to be applied as classifiers for challenging tasks such as image recognition. This remains an interesting direction for future research.</p>
<p>Although number-state preserving tensors represent a highly restricted class of tensor, the preliminary results of <xref ref-type="sec" rid="s6">Section 6</xref> are encouraging that this class is sufficient when applying tensor networks as classifiers for learning problems as outlined in <xref ref-type="sec" rid="s3">Section 3</xref>. It still remains to be seen whether number-state preserving tensor networks are as powerful as generic tensors networks for these tasks; this question requires further theoretical and numerical investigation. However it is relatively easy to understand that, in the limit of large bond dimension, a number-state preserving tensor network could in principle achieve 100% accuracy on any training problem outlined in <xref ref-type="sec" rid="s3">Section 3</xref>. The reasoning follows similarly to the argument that a generic tensor network can represent an arbitrary quantum state in the limit of large bond dimension. Consider, for instance, the MERA depicted in <xref ref-type="fig" rid="F2">Figure 2B</xref>. One could increase the bond dimension of indices within the network until the output index of each <italic>w</italic> tensor matches the product of its input dimensions, in which case each <italic>w</italic> could be fixed as a trivial identity tensor when viewed as an input-output matrix. In this scenario, the top tensor <italic>w</italic>
<sub>top</sub> could implement an arbitrary classifier that would perfectly map every training sample to its designated label, regardless of the training data given.</p>
<p>A major difficulty with the use of MERA in <italic>D</italic> &#x3d; 2 or higher spatial dimensions [<xref ref-type="bibr" rid="B37">37</xref>, <xref ref-type="bibr" rid="B38">38</xref>] is their high scaling of computational cost with bond dimension <italic>&#x3c7;</italic>. However, there is reason to be more optimistic for their application as classifiers. The cost of contracting an isometric metric MERA for a density matrix, necessary for its optimization towards the ground state of a local Hamiltonian, is related to the size of the maximum causal width of the network. For instance, the most efficient known 2<italic>D</italic> isometric MERA [<xref ref-type="bibr" rid="B38">38</xref>] has a causal width of 2 &#xd7; 2 sites, such that the density matrices within the causal cone have 8 indices. The cost of computing these density matrices can be shown to scale at most as <italic>O</italic> (<italic>&#x3c7;</italic>
<sup>16</sup>). However, while a number-state preserving version of this 2<italic>D</italic> MERA would also have a causal width of 2 &#xd7; 2 sites, the relevant configuration space &#x7c;<italic>&#x3c8;</italic>&#x27e9; within the causal cone would only have 4 indices (which follows as the density matrix involves both the <italic>bra</italic> and the <italic>ket</italic> state, whereas the configuration space only involves the <italic>ket</italic>). Thus the cost of optimizing a number-state preserving version of this 2<italic>D</italic> MERA, where the key step is the evaluation of configuration spaces, will scale roughly as <italic>O</italic> (<italic>&#x3c7;</italic>
<sup>8</sup>) (i.e., the square-root of the cost of optimizing an isometric MERA for a quantum ground state). This square-root reduction in cost scaling as a function of bond dimension <italic>&#x3c7;</italic> from isometric to number-state preserving networks will hold in general, such that number-state preserving networks could realize much larger bond dimensions given a fixed computational budget. This advantage is somewhat mitigated by the fact that the cost of optimizing a number-state preserving network comes with a factor <italic>n</italic>
<sub>
<italic>samp</italic>
</sub> related to the size of the training set, which could be very large. However, it would also be straight-forward to parallelize the evaluation of environments over the samples.</p>
<p>Although the main text of this manuscript focused on number-state preserving versions of MERA, many other forms of hierarchical network could also be of useful as classifiers as discussed further in <xref ref-type="sec" rid="s13">Supplementary Appendix SB1</xref>. In particular the network of <xref ref-type="sec" rid="s13">Supplementary Figure SB2</xref>, which does not have an isometric counterpart, seems to be the closest tensor network analogue to a convolutional neural network. Rather than disentanglers, this network uses <italic>&#x3b4;</italic>-function tensors to effectively allow neighboring <italic>w</italic> tensors to &#x201c;read&#x201d; from the same boundary sites, mirroring the overlap of feature maps arising in a convolution (and similar to the generalized networks recently proposed in Ref. [<xref ref-type="bibr" rid="B31">31</xref>]). It would be interesting to compare the effectiveness of this structure <italic>versus</italic> a traditional MERA, which will be considered in future work.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s8">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/Supplementary Material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s9">
<title>Author contributions</title>
<p>Manuscript is the singular effort of GE.</p>
</sec>
<sec id="s10">
<title>Funding</title>
<p>This research was supported in part by the National Science Foundation under Grant No. NSF PHY-1748958.</p>
</sec>
<ack>
<p>The author thanks Miles Stoudenmire and John Terilla for useful discussions and comments.</p>
</ack>
<sec sec-type="COI-statement" id="s11">
<title>Conflict of interest</title>
<p>The author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s12">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s13">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fphy.2022.858388/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fphy.2022.858388/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="Presentation1.pdf" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<label>1.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Carleo</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Cirac</surname>
<given-names>JI</given-names>
</name>
<name>
<surname>Cranmer</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Daudet</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Schuld</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Tishby</surname>
<given-names>N</given-names>
</name>
<etal/>
</person-group> <article-title>Machine learning and the physical sciences</article-title>. <source>Rev Mod Phys</source> (<year>2019</year>) <volume>91</volume>:<fpage>045002</fpage>. <pub-id pub-id-type="doi">10.1103/RevModPhys.91.045002</pub-id>
</citation>
</ref>
<ref id="B2">
<label>2.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Carrasquilla</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Melko</surname>
<given-names>RG</given-names>
</name>
</person-group>. <article-title>Machine learning phases of matter</article-title>. <source>Nat Phys</source> (<year>2017</year>) <volume>13</volume>:<fpage>431</fpage>&#x2013;<lpage>4</lpage>. <pub-id pub-id-type="doi">10.1038/nphys4035</pub-id>
</citation>
</ref>
<ref id="B3">
<label>3.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Broecker</surname>
<given-names>P</given-names>
</name>
<name>
<surname>Carrasquilla</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Melko</surname>
<given-names>RG</given-names>
</name>
<name>
<surname>Trebst</surname>
<given-names>S</given-names>
</name>
</person-group>. <article-title>Machine learning quantum phases of matter beyond the fermion sign problem</article-title>. <source>Sci Rep</source> (<year>2017</year>) <volume>7</volume>:<fpage>8823</fpage>. <pub-id pub-id-type="doi">10.1038/s41598-017-09098-0</pub-id>
</citation>
</ref>
<ref id="B4">
<label>4.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ch&#x2019;ng</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Carrasquilla</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Melko</surname>
<given-names>RG</given-names>
</name>
<name>
<surname>Khatami</surname>
<given-names>E</given-names>
</name>
</person-group>. <article-title>Machine learning phases of strongly correlated fermions</article-title>. <source>Phys Rev X</source> (<year>2017</year>) <volume>7</volume>:<fpage>031038</fpage>. <pub-id pub-id-type="doi">10.1103/physrevx.7.031038</pub-id>
</citation>
</ref>
<ref id="B5">
<label>5.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huembeli</surname>
<given-names>P</given-names>
</name>
<name>
<surname>Dauphin</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Wittek</surname>
<given-names>P</given-names>
</name>
</person-group>. <article-title>Identifying quantum phase transitions with adversarial neural networks</article-title>. <source>Phys Rev B</source> (<year>2018</year>) <volume>97</volume>:<fpage>134109</fpage>. <pub-id pub-id-type="doi">10.1103/physrevb.97.134109</pub-id>
</citation>
</ref>
<ref id="B6">
<label>6.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>Y-H</given-names>
</name>
<name>
<surname>van Nieuwenburg</surname>
<given-names>EPL</given-names>
</name>
</person-group>. <article-title>Discriminative cooperative networks for detecting phase transitions</article-title>. <source>Phys Rev Lett</source> (<year>2018</year>) <volume>120</volume>:<fpage>176401</fpage>. <pub-id pub-id-type="doi">10.1103/physrevlett.120.176401</pub-id>
</citation>
</ref>
<ref id="B7">
<label>7.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Canabarro</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Fanchini</surname>
<given-names>FF</given-names>
</name>
<name>
<surname>Malvezzi</surname>
<given-names>AL</given-names>
</name>
<name>
<surname>Pereira</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Chaves</surname>
<given-names>R</given-names>
</name>
</person-group>. <article-title>Unveiling phase transitions with machine learning</article-title>. <source>Phys Rev B</source> (<year>2019</year>) <volume>100</volume>:<fpage>045129</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevB.100.045129</pub-id>
</citation>
</ref>
<ref id="B8">
<label>8.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Torlai</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Mazzola</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Carrasquilla</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Troyer</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Melko</surname>
<given-names>RG</given-names>
</name>
<name>
<surname>Carleo</surname>
<given-names>G</given-names>
</name>
</person-group>. <article-title>Neural-network quantum state tomography</article-title>. <source>Nat Phys</source> (<year>2018</year>) <volume>14</volume>:<fpage>447</fpage>&#x2013;<lpage>50</lpage>. <pub-id pub-id-type="doi">10.1038/s41567-018-0048-5</pub-id>
</citation>
</ref>
<ref id="B9">
<label>9.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Carrasquilla</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Torlai</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Melko</surname>
<given-names>RG</given-names>
</name>
<name>
<surname>Aolita</surname>
<given-names>L</given-names>
</name>
</person-group>. <article-title>Reconstructing quantum states with generative models</article-title>. <source>Nat Mach Intell</source> (<year>2019</year>) <volume>1</volume>:<fpage>155</fpage>&#x2013;<lpage>61</lpage>. <pub-id pub-id-type="doi">10.1038/s42256-019-0028-1</pub-id>
</citation>
</ref>
<ref id="B10">
<label>10.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Torlai</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Melko</surname>
<given-names>RG</given-names>
</name>
</person-group>. <article-title>Learning thermodynamics with Boltzmann machines</article-title>. <source>Phys Rev B</source> (<year>2016</year>) <volume>94</volume>:<fpage>165134</fpage>. <pub-id pub-id-type="doi">10.1103/physrevb.94.165134</pub-id>
</citation>
</ref>
<ref id="B11">
<label>11.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Carleo</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Troyer</surname>
<given-names>M</given-names>
</name>
</person-group>. <article-title>Solving the quantum many-body problem with artificial neural networks</article-title>. <source>Science</source> (<year>2017</year>) <volume>355</volume>:<fpage>602</fpage>&#x2013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1126/science.aag2302</pub-id>
</citation>
</ref>
<ref id="B12">
<label>12.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Choo</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Neupert</surname>
<given-names>T</given-names>
</name>
<name>
<surname>Carleo</surname>
<given-names>G</given-names>
</name>
</person-group>. <article-title>Study of the two-dimensional frustrated J1-J2 model with neural network quantum states</article-title>. <source>Phys Rev B</source> (<year>2019</year>) <volume>100</volume>:<fpage>125124</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevB.100.125124</pub-id>
</citation>
</ref>
<ref id="B13">
<label>13.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Hassoun</surname>
<given-names>M</given-names>
</name>
</person-group>. <source>Fundamentals of artificial neural networks</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>MIT Press</publisher-name> (<year>1995</year>).</citation>
</ref>
<ref id="B14">
<label>14.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Schalkoff</surname>
<given-names>RJ</given-names>
</name>
</person-group>. <source>Artificial neural networks</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>McGraw-Hill</publisher-name> (<year>1997</year>).</citation>
</ref>
<ref id="B15">
<label>15.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Or&#xfa;s</surname>
<given-names>R</given-names>
</name>
</person-group>. <article-title>A practical introduction to tensor networks: Matrix product states and projected entangled pair states</article-title>. <source>Ann Phys (N Y)</source> (<year>2014</year>) <volume>349</volume>:<fpage>117</fpage>&#x2013;<lpage>58</lpage>. <pub-id pub-id-type="doi">10.1016/j.aop.2014.06.013</pub-id>
</citation>
</ref>
<ref id="B16">
<label>16.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Levine</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Sharir</surname>
<given-names>O</given-names>
</name>
<name>
<surname>Cohen</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Shashua</surname>
<given-names>A</given-names>
</name>
</person-group>. <article-title>Quantum Entanglement in Deep Learning Architectures</article-title>. <source>Phys Rev Lett</source> (<year>2019</year>) <volume>122</volume>:<fpage>065301</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevLett.122.065301</pub-id>
</citation>
</ref>
<ref id="B17">
<label>17.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Moore</surname>
<given-names>JE</given-names>
</name>
</person-group>. <article-title>Neural network representation of tensor network and chiral states</article-title>. <source>Phys Rev Lett</source> (<year>2021</year>) <volume>127</volume>:<fpage>170601</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevLett.127.170601</pub-id>
</citation>
</ref>
<ref id="B18">
<label>18.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Deng</surname>
<given-names>D-L</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Das Sarma</surname>
<given-names>S</given-names>
</name>
</person-group>. <article-title>Quantum entanglement in neural network states</article-title>. <source>Phys Rev X</source> (<year>2017</year>) <volume>7</volume>:<fpage>021021</fpage>. <pub-id pub-id-type="doi">10.1103/physrevx.7.021021</pub-id>
</citation>
</ref>
<ref id="B19">
<label>19.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Glasser</surname>
<given-names>I</given-names>
</name>
<name>
<surname>Pancotti</surname>
<given-names>N</given-names>
</name>
<name>
<surname>August</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Rodriguez</surname>
<given-names>ID</given-names>
</name>
<name>
<surname>Cirac</surname>
<given-names>JI</given-names>
</name>
</person-group>. <article-title>Neural-network quantum states, string-bond states, and chiral topological states</article-title>. <source>Phys Rev X</source> (<year>2018</year>) <volume>8</volume>:<fpage>011006</fpage>. <pub-id pub-id-type="doi">10.1103/physrevx.8.011006</pub-id>
</citation>
</ref>
<ref id="B20">
<label>20.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cai</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Approximating quantum many-body wave functions using artificial neural networks</article-title>. <source>Phys Rev B</source> (<year>2018</year>) <volume>97</volume>:<fpage>035116</fpage>. <pub-id pub-id-type="doi">10.1103/physrevb.97.035116</pub-id>
</citation>
</ref>
<ref id="B21">
<label>21.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Stoudenmire</surname>
<given-names>EM</given-names>
</name>
<name>
<surname>Schwab</surname>
<given-names>DJ</given-names>
</name>
</person-group>. <article-title>Supervised learning with quantum-inspired tensor networks</article-title>. <source>Adv Neural Inf Process Syst</source> (<year>2016</year>) <volume>29</volume>:<fpage>4799</fpage>. <pub-id pub-id-type="doi">10.48550/arXiv.1605.05775</pub-id>
</citation>
</ref>
<ref id="B22">
<label>22.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cohen</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Sharir</surname>
<given-names>O</given-names>
</name>
<name>
<surname>Levine</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Tamari</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Yakira</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Shashua</surname>
<given-names>A</given-names>
</name>
</person-group>. <article-title>Analysis and design of convolutional networks via hierarchical tensor decompositions</article-title>. <comment>arXiv:1705.02302</comment> (<year>2017</year>). <pub-id pub-id-type="doi">10.48550/arXiv.1705.02302</pub-id>
</citation>
</ref>
<ref id="B23">
<label>23.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Han</surname>
<given-names>Z-Y</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Fan</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>P</given-names>
</name>
</person-group>. <article-title>Unsupervised generative modeling using matrix product states</article-title>. <source>Phys Rev X</source> (<year>2018</year>) <volume>8</volume>:<fpage>031012</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevX.8.031012</pub-id>
</citation>
</ref>
<ref id="B24">
<label>24.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cichocki</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Phan</surname>
<given-names>A-H</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Q</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Oseledets</surname>
<given-names>I</given-names>
</name>
<name>
<surname>Sugiyama</surname>
<given-names>M</given-names>
</name>
<etal/>
</person-group> <article-title>Tensor networks for dimensionality reduction and large-scale optimization: Part 2 applications and future perspectives</article-title>. <source>FNT Machine Learn</source> (<year>2017</year>) <volume>9</volume>:<fpage>249</fpage>&#x2013;<lpage>429</lpage>. <pub-id pub-id-type="doi">10.1561/2200000067</pub-id>
</citation>
</ref>
<ref id="B25">
<label>25.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Ran</surname>
<given-names>S-J</given-names>
</name>
<name>
<surname>Wittek</surname>
<given-names>P</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Garca</surname>
<given-names>RB</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>G</given-names>
</name>
<etal/>
</person-group> <article-title>Machine learning by two-dimensional hierarchical tensor networks: A quantum information theoretic perspective on deep architectures</article-title>. <source>New J Phys</source> (<year>2019</year>) <volume>21</volume>:<fpage>073059</fpage>. <pub-id pub-id-type="doi">10.1088/1367-2630/ab31ef</pub-id>
</citation>
</ref>
<ref id="B26">
<label>26.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hallam</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Grant</surname>
<given-names>E</given-names>
</name>
<name>
<surname>Stojevic</surname>
<given-names>V</given-names>
</name>
<name>
<surname>Severini</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Green</surname>
<given-names>AG</given-names>
</name>
</person-group>. <article-title>Compact neural networks based on the multiscale entanglement renormalization ansatz</article-title>. <comment>arXiv:1711.03357</comment> (<year>2017</year>). <pub-id pub-id-type="doi">10.48550/arXiv.1711.03357</pub-id>
</citation>
</ref>
<ref id="B27">
<label>27.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Stoudenmire</surname>
<given-names>EM</given-names>
</name>
</person-group>. <article-title>Learning relevant features of data with multi-scale tensor networks</article-title>. <source>Quan Sci Technol</source> (<year>2018</year>) <volume>3</volume>:<fpage>034003</fpage>. <pub-id pub-id-type="doi">10.1088/2058-9565/aaba1a</pub-id>
</citation>
</ref>
<ref id="B28">
<label>28.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>W&#x2010;J</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Lewenstein</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Ran</surname>
<given-names>S-J</given-names>
</name>
</person-group>. <article-title>Entanglement-Based Feature Extraction by Tensor Network Machine Learning</article-title>. <source>Front Appl Math Stat</source> (<year>2021</year>). <pub-id pub-id-type="doi">10.3389/fams.2021.716044</pub-id>
</citation>
</ref>
<ref id="B29">
<label>29.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huggins</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Patel</surname>
<given-names>P</given-names>
</name>
<name>
<surname>Whaley</surname>
<given-names>KB</given-names>
</name>
<name>
<surname>Stoudenmire</surname>
<given-names>EM</given-names>
</name>
</person-group>. <article-title>Towards quantum machine learning with tensor networks</article-title>. <source>Quan Sci Technol</source> (<year>2019</year>) <volume>4</volume>:<fpage>024001</fpage>. <pub-id pub-id-type="doi">10.1088/2058-9565/aaea94</pub-id>
</citation>
</ref>
<ref id="B30">
<label>30.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Grant</surname>
<given-names>E</given-names>
</name>
<name>
<surname>Benedetti</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Hallam</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Lockhart</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Stojevic</surname>
<given-names>V</given-names>
</name>
<etal/>
</person-group> <article-title>Hierarchical quantum classifiers</article-title>. <source>npj Quan Info</source> (<year>2018</year>) <volume>4</volume>:<fpage>65</fpage>. <pub-id pub-id-type="doi">10.1038/s41534-018-0116-9</pub-id>
</citation>
</ref>
<ref id="B31">
<label>31.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Glasser</surname>
<given-names>I</given-names>
</name>
<name>
<surname>Pancotti</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Cirac</surname>
<given-names>JI</given-names>
</name>
</person-group>. <article-title>From probabilistic graphical models to generalized tensor networks for supervised learning</article-title>. <comment>arXiv:1806.05964v2</comment> (<year>2018</year>). <pub-id pub-id-type="doi">10.48550/arXiv.1806.05964</pub-id>
</citation>
</ref>
<ref id="B32">
<label>32.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fannes</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Nachtergaele</surname>
<given-names>B</given-names>
</name>
<name>
<surname>Werner</surname>
<given-names>RF</given-names>
</name>
</person-group>. <article-title>Finitely correlated states on quantum spin chains</article-title>. <source>Commun Math Phys</source> (<year>1992</year>) <volume>144</volume>:<fpage>443</fpage>&#x2013;<lpage>90</lpage>. <pub-id pub-id-type="doi">10.1007/bf02099178</pub-id>
</citation>
</ref>
<ref id="B33">
<label>33.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ostlund</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Rommer</surname>
<given-names>S</given-names>
</name>
</person-group>. <article-title>Thermodynamic limit of density matrix renormalization</article-title>. <source>Phys Rev Lett</source> (<year>1995</year>) <volume>75</volume>:<fpage>3537</fpage>&#x2013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1103/physrevlett.75.3537</pub-id>
</citation>
</ref>
<ref id="B34">
<label>34.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shi</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Duan</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Vidal</surname>
<given-names>G</given-names>
</name>
</person-group>. <article-title>Classical simulation of quantum many-body systems with a tree tensor network</article-title>. <source>Phys Rev A (Coll Park)</source> (<year>2006</year>) <volume>74</volume>:<fpage>022320</fpage>. <pub-id pub-id-type="doi">10.1103/physreva.74.022320</pub-id>
</citation>
</ref>
<ref id="B35">
<label>35.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tagliacozzo</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Evenbly</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Vidal</surname>
<given-names>G</given-names>
</name>
</person-group>. <article-title>Simulation of two-dimensional quantum systems using a tree tensor network that exploits the entropic area law</article-title>. <source>Phys Rev B</source> (<year>2009</year>) <volume>80</volume>:<fpage>235127</fpage>. <pub-id pub-id-type="doi">10.1103/physrevb.80.235127</pub-id>
</citation>
</ref>
<ref id="B36">
<label>36.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vidal</surname>
<given-names>G</given-names>
</name>
</person-group>. <article-title>Class of quantum many-body states that can Be efficiently simulated</article-title>. <source>Phys Rev Lett</source> (<year>2008</year>) <volume>101</volume>:<fpage>110501</fpage>. <pub-id pub-id-type="doi">10.1103/physrevlett.101.110501</pub-id>
</citation>
</ref>
<ref id="B37">
<label>37.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cincio</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Dziarmaga</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Rams</surname>
<given-names>MM</given-names>
</name>
</person-group>. <article-title>Multiscale entanglement renormalization ansatz in two dimensions: Quantum ising model</article-title>. <source>Phys Rev Lett</source> (<year>2008</year>) <volume>100</volume>:<fpage>240603</fpage>. <pub-id pub-id-type="doi">10.1103/physrevlett.100.240603</pub-id>
</citation>
</ref>
<ref id="B38">
<label>38.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Evenbly</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Vidal</surname>
<given-names>G</given-names>
</name>
</person-group>. <article-title>Entanglement renormalization in two spatial dimensions</article-title>. <source>Phys Rev Lett</source> (<year>2009</year>) <volume>102</volume>:<fpage>180406</fpage>. <pub-id pub-id-type="doi">10.1103/physrevlett.102.180406</pub-id>
</citation>
</ref>
<ref id="B39">
<label>39.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Evenbly</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Vidal</surname>
<given-names>G</given-names>
</name>
</person-group>. <article-title>Quantum criticality with the multi-scale entanglement renormalization ansatz, chapter 4 in strongly correlated systems: Numerical methods</article-title>. In: <person-group person-group-type="editor">
<name>
<surname>Avella,</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Mancini</surname>
<given-names>F</given-names>
</name>
</person-group>, editors. <publisher-name>Springer Series in Solid-State Sciences</publisher-name> (<year>2013</year>).</citation>
</ref>
<ref id="B40">
<label>40.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>LeCun</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Boser</surname>
<given-names>B</given-names>
</name>
<name>
<surname>Denker</surname>
<given-names>JS</given-names>
</name>
<name>
<surname>Henderson</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Howard</surname>
<given-names>RE</given-names>
</name>
<name>
<surname>Hubbard</surname>
<given-names>W</given-names>
</name>
<etal/>
</person-group> <article-title>Backpropagation applied to handwritten zip code recognition</article-title>. <source>Neural Comput</source> (<year>1989</year>) <volume>1</volume>(<issue>4</issue>):<fpage>541</fpage>&#x2013;<lpage>51</lpage>. <pub-id pub-id-type="doi">10.1162/neco.1989.1.4.541</pub-id>
</citation>
</ref>
<ref id="B41">
<label>41.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Krizhevsky</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Sutskever</surname>
<given-names>I</given-names>
</name>
<name>
<surname>Hinton</surname>
<given-names>GE</given-names>
</name>
</person-group>. <article-title>ImageNet classification with deep convolutional neural networks</article-title>. <source>Commun ACM</source> (<year>2017</year>) <volume>60</volume>(<issue>6</issue>):<fpage>84</fpage>&#x2013;<lpage>90</lpage>. <pub-id pub-id-type="doi">10.1145/3065386</pub-id>
</citation>
</ref>
<ref id="B42">
<label>42.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Simonyan</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Zisserman</surname>
<given-names>A</given-names>
</name>
</person-group>. <article-title>Very deep convolutional networks for large-scale image recognition</article-title>. <comment>arXiv:1409.1556</comment> (<year>2014</year>). <pub-id pub-id-type="doi">10.48550/arXiv.1409.1556</pub-id>
</citation>
</ref>
<ref id="B43">
<label>43.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bravyi</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Caha</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Movassagh</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Nagaj</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Shor</surname>
<given-names>P</given-names>
</name>
</person-group>. <article-title>Criticality without frustration for quantum spin-1 chains</article-title>. <source>Phys Rev Lett</source> (<year>2012</year>) <volume>109</volume>:<fpage>207202</fpage>. <pub-id pub-id-type="doi">10.1103/physrevlett.109.207202</pub-id>
</citation>
</ref>
<ref id="B44">
<label>44.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Alexander</surname>
<given-names>RN</given-names>
</name>
<name>
<surname>Evenbly</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Klich</surname>
<given-names>I</given-names>
</name>
</person-group>. <article-title>Exact holographic tensor networks for the Motzkin spin chain</article-title>. <source>Quan</source> (<year>2021</year>) <volume>5</volume>:<fpage>546</fpage>. <pub-id pub-id-type="doi">10.22331/q-2021-09-21-546</pub-id>
</citation>
</ref>
<ref id="B45">
<label>45.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Alexander</surname>
<given-names>RN</given-names>
</name>
<name>
<surname>Ahmadain</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Klich</surname>
<given-names>I</given-names>
</name>
</person-group>. <article-title>Holographic rainbow networks for colorful Motzkin and Fredkin spin chains</article-title>. <source>Phys Rev B</source> (<year>2019</year>) <volume>100</volume>:<fpage>214430</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevB.100.214430</pub-id>
</citation>
</ref>
<ref id="B46">
<label>46.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname>
<given-names>G-X</given-names>
</name>
<name>
<surname>Ho</surname>
<given-names>C-H</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>C-J</given-names>
</name>
</person-group>. <article-title>Recent advances of large-scale linear classification</article-title>. <source>Proc IEEE</source> (<year>2012</year>) <volume>100</volume>(<issue>9</issue>):<fpage>2584</fpage>&#x2013;<lpage>603</lpage>. <pub-id pub-id-type="doi">10.1109/jproc.2012.2188013</pub-id>
</citation>
</ref>
<ref id="B47">
<label>47.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hofmann</surname>
<given-names>T</given-names>
</name>
<name>
<surname>Sch&#xf6;lkopf</surname>
<given-names>B</given-names>
</name>
<name>
<surname>Smola</surname>
<given-names>AJ</given-names>
</name>
</person-group>. <article-title>Kernel methods in machine learning</article-title>. <source>Ann Statist</source> (<year>2008</year>) <volume>36</volume>:<fpage>1171</fpage>. <pub-id pub-id-type="doi">10.1214/009053607000000677</pub-id>
</citation>
</ref>
<ref id="B48">
<label>48.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Evenbly</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Vidal</surname>
<given-names>G</given-names>
</name>
</person-group>. <article-title>Algorithms for entanglement renormalization</article-title>. <source>Phys Rev B</source> (<year>2009</year>) <volume>79</volume>:<fpage>144108</fpage>. <pub-id pub-id-type="doi">10.1103/physrevb.79.144108</pub-id>
</citation>
</ref>
<ref id="B49">
<label>49.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>T</given-names>
</name>
</person-group>. <article-title>Solving large-scale linear prediction problems with stochastic gradient descent</article-title>. In: <conf-name>Proceedings of the international conference on machine learning</conf-name> (<year>2004</year>).</citation>
</ref>
<ref id="B50">
<label>50.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Metropolis</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Rosenbluth</surname>
<given-names>AW</given-names>
</name>
<name>
<surname>Rosenbluth</surname>
<given-names>MN</given-names>
</name>
<name>
<surname>Teller</surname>
<given-names>AH</given-names>
</name>
<name>
<surname>Teller</surname>
<given-names>E</given-names>
</name>
</person-group>. <article-title>Equation of state calculations by fast computing machines</article-title>. <source>J Chem Phys</source> (<year>1953</year>) <volume>21</volume>:<fpage>1087</fpage>&#x2013;<lpage>92</lpage>. <pub-id pub-id-type="doi">10.1063/1.1699114</pub-id>
</citation>
</ref>
<ref id="B51">
<label>51.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Evenbly</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Vidal</surname>
<given-names>G</given-names>
</name>
</person-group>. <article-title>Scaling of entanglement entropy in the (branching) multi-scale entanglement renormalization ansatz</article-title>. <source>Phys Rev B</source> (<year>2014</year>) <volume>89</volume>:<fpage>235113</fpage>. <pub-id pub-id-type="doi">10.1103/physrevb.89.235113</pub-id>
</citation>
</ref>
<ref id="B52">
<label>52.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Stokes</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Terilla</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Probabilistic modeling with matrix product states</article-title>. <comment>arXiv:1902.06888</comment> (<year>2019</year>). <pub-id pub-id-type="doi">10.3390/e21121236</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>