<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="discussion" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Appl. Math. Stat.</journal-id>
<journal-title>Frontiers in Applied Mathematics and Statistics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Appl. Math. Stat.</abbrev-journal-title>
<issn pub-type="epub">2297-4687</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">641239</article-id>
<article-id pub-id-type="doi">10.3389/fams.2021.641239</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Applied Mathematics and Statistics</subject>
<subj-group>
<subject>Opinion</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Information Theory and Consciousness</article-title>
<alt-title alt-title-type="left-running-head">Jost</alt-title>
<alt-title alt-title-type="right-running-head">Information Theory and Consciousness</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Jost</surname>
<given-names>J&#xfc;rgen</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/751097/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>Max Planck Institute for Mathematics in the Sciences, <addr-line>Leipzig</addr-line>, <country>Germany</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>Santa Fe Institute, <addr-line>Santa Fe</addr-line>, <addr-line>NM</addr-line>, <country>United States</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/771472/overview">Johannes Kleiner</ext-link>, Ludwig Maximilian University of Munich, Germany</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/821342/overview">Miguel Pineda</ext-link>, University College London, United&#x20;Kingdom</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: J&#xfc;rgen Jost, <email>jost@mis.mpg.de</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Dynamical Systems, a section of the journal Frontiers in Applied Mathematics and Statistics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>10</day>
<month>08</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>7</volume>
<elocation-id>641239</elocation-id>
<history>
<date date-type="received">
<day>04</day>
<month>06</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>29</day>
<month>06</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Jost.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Jost</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<kwd-group>
<kwd>consciousness</kwd>
<kwd>information</kwd>
<kwd>integration</kwd>
<kwd>compression</kwd>
<kwd>evolutionary aspects</kwd>
<kwd>social conditions</kwd>
<kwd>reflexivity</kwd>
</kwd-group>
<contract-sponsor id="cn001">Max-Planck-Institut f&#xfc;r Mathematik in den Naturwissenschaften10.13039/501100013296</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Perhaps everything is just an illusion. But even an illusion might be useful. Therefore, in any case, we should try to understand it. And scientific understanding of consciousness, like any other phenomenon, should be based on the available facts. Let us recall some of them.<list list-type="simple">
<list-item>
<p>&#x2022; Consciousness is a subjectively experienced phenomenon that cannot be doubted, as Descartes famously observed. As such, while it may admit gradations, it is experienced as an all-or-none phenomenon, that is, one is either conscious or&#x20;not.</p>
</list-item>
<list-item>
<p>&#x2022; It has different aspects, from a sense of awareness to the qualitative aspects of feelings and of the sensations that give rise to feelings and to the sense of selfhood, identifying the self as an integrated system distinct from the rest of the&#x20;world.</p>
</list-item>
<list-item>
<p>&#x2022; It is correlated with neurophysiological dynamics and depends on neuroanatomical structures (see, for instance, [<xref ref-type="bibr" rid="B1">1</xref>]), although much has been argued about the underlying causality. Neuroscience has identified some of these dynamics and structures, although the answers at present are not yet conclusive. It is not simply a question of numbers of neurons, as some parts of the brain, such as the cerebellum, apparently do not play an essential role for consciousness. Some curious phenomena have been observed in split-brain patients where for medical reasons, the corpus callosum that connects the two hemispheres of the brain has been cut. It seems that in such patients, the two halves of the brain are separately conscious, even though only the left hemisphere can express itself verbally.</p>
</list-item>
<list-item>
<p>&#x2022; Human consciousness seems to depend on language and on a social context. Humans who have been brought up in complete isolation cannot learn to speak normally and may only have some rudimentary form of consciousness, if at&#x20;all.</p>
</list-item>
</list>
</p>
<p>These facts raise many questions, among them.<list list-type="simple">
<list-item>
<p>&#x2022; To what extent can animals be conscious? Some higher animals seem to possess a sense of self, and they can assess the knowledge that others possess and understand their intentions and on this basis anticipate their actions. That is, they seem to have some rudimentary form of a theory of mind. They cannot, however, communicate that to us, and we can only reconstruct their minds from indirect evidence. Also, human infants seem to be very clever in their own way (see, for instance, [<xref ref-type="bibr" rid="B2">2</xref>]) [<xref ref-type="bibr" rid="B3">3</xref>], but they strikingly lack any episodic memory from their first years of life. We nevertheless grant them consciousness, even though it is a robust criterion for the conscious thinking and decisions of adults that it can be recalled and remembered (and of course be forgotten later). What we consider as significant conscious acts, we can remember for the rest of our&#x20;lives.</p>
</list-item>
<list-item>
<p>&#x2022; To what extent does consciousness depend on neural wetware? That is, could computers or distributed programs possibly become conscious? Functionalism, as first advocated [<xref ref-type="bibr" rid="B4">4</xref>] and then rejected [<xref ref-type="bibr" rid="B5">5</xref>] by Putnam, would admit that possibility. Other scientists, such as Koch [<xref ref-type="bibr" rid="B6">6</xref>], maintain that consciousness can only emerge in human or perhaps animal brains and requires particular neural and/or (thalamo)cortical structures. We shall argue below that consciousness at least does not automatically arise when a system becomes sufficiently complex. We are not conscious simply because we have a large brain, but rather humans have evolved to become conscious when exposed to other conscious humans during a critical phase of their development. That is, first, consciousness is partly a social phenomenon, even though it seems that a main aspect of consciousness is to distinguish a self from others, and second, there were evolutionary reasons for the emergence of consciousness.</p>
</list-item>
<list-item>
<p>&#x2022; Whether the evolution of consciousness was a gradual process or a sudden jump is another question to which research on nonhuman primates can provide some partial answers. More generally, one tries to associate conscious wakefulness, the orienting and focusing of attention, sensory perception, the immediacy of memory, and perhaps even the emergence of a sense of self with the evolution of the mammalian brain, in particular the neocortex or the thalamocortical&#x20;loop.</p>
</list-item>
<list-item>
<p>&#x2022; In any case, it seems that so far neuroscience has not identified any qualitative difference between human brains and those of other mammals. Human brains are larger than those of other great apes. In particular, in human brains, the prefrontal cortex is enlarged (see [<xref ref-type="bibr" rid="B7">7</xref>] for some possible evolutionary explanation that intertwines functional and structural aspects), but the prefrontal cortex, while important for action planning for instance, does not seem decisive for consciousness. Also, human brains are not the largest mammalian brains. Mammalian brains scale with body size, simply because a larger body requires more neural control, and human brains are somewhat larger than this scaling relation would suggest, but certainly not spectacularly so. Perhaps the ongoing investigation of the connectivity pattern of the human brain, the so-called connectome (see, for instance, [<xref ref-type="bibr" rid="B8">8</xref>]), may identify some crucial differences in the wiring pattern. But this remains to be&#x20;seen.</p>
</list-item>
<list-item>
<p>&#x2022; This also leads to the question of whether different brain structures, such as those of birds or cephalopods, that is, of other branches in the animal kingdom that have evolved versions of intelligent behavior, could, at least in principle, support forms of consciousness that are possibly very different from ours (Nagel [<xref ref-type="bibr" rid="B9">9</xref>] famously discussed the case of bats that have a sensory system, the sonar, that we do not possess, and therefore experience their environments very differently. But, on the other hand, the example of aircontrollers shows that we are in principle able to build up a sophisticated intuition about a novel class of inputs, radar images in their case, and perceive our environment accordingly.).</p>
</list-item>
<list-item>
<p>&#x2022; In this regard, we should also keep in mind that evolutionary origin and current function of a structure or a feature need not coincide. What had originally evolved in a specific context for a particular function, or even as the sideproduct of some other functional structure, may have subsequently acquired a very different function, as systematically argued by Gould [<xref ref-type="bibr" rid="B10">10</xref>]. Furthermore, many, if not all, of our systems and structures have a variety of functions, and also general human abilities such as language cannot be reduced to a single function. We should expect that this also applies to consciousness. In particular, we should keep this in mind in the discussion of the next&#x20;items.</p>
</list-item>
<list-item>
<p>&#x2022; The famous observations and experiments of Libet [<xref ref-type="bibr" rid="B11">11</xref>] show that the brain prepares acts and apparently decides to execute them some time before they enter consciousness.<xref ref-type="fn" rid="fn1">
<sup>1</sup>
</xref> Is consciousness therefore only an inconsequential side effect of a neural decision that occurs subconsciously in the brain? Or is the only function of consciousness to intervene when such subconsciously determined acts might have unwanted consequences? But conscious decisions take much longer than subconscious routines, and so, there seems to be some fundamental difference. Even if consciousness were only a side effect or some mechanisms that offer an opportunity for subsequent control, what decides which decisions become conscious? Why some and not others? And why does learning a new skill usually require a conscious effort? Is this simply a mechanism that evaluates whether the execution of the corresponding action has been sufficiently successful?</p>
</list-item>
<list-item>
<p>&#x2022; Or is the purpose of consciousness the rationalization of subconsciously generated actions? In addition to Libet&#x2019;s findings, also some observations from split-brain patients, where the left hemisphere invents explanations for actions triggered by stimuli presented to the right hemisphere, might indicate such a function. However, one should always be careful in drawing conclusions from the apparent misfunctioning of some system under abnormal conditions, here the severing of the connection between the two halves of the brain, about its normal functioning.</p>
</list-item>
<list-item>
<p>&#x2022; Is our version of consciousness perhaps still very imperfect? Could evolution produce superior ones?</p>
</list-item>
</list>
</p>
<p>In this contribution, the issue of consciousness is approached from the conceptual framework of information theory. This does not mean that a formal theory of consciousness will be developed, but only that information theoretical principle will guide our thinking. I believe that this is helpful for clarifying some important conceptual issues in the discussion of consciousness. Of course, information theory has been applied to the theory of cognition in general, but we do not intend to provide an overview here, as the subject is too vast, but only refer to [<xref ref-type="bibr" rid="B12">12</xref>] for a systematic approach that introduces some concepts and touches some aspects that are also relevant&#x20;here.</p>
</sec>
<sec id="s2">
<title>2 Information and Structure</title>
<p>In this section, we recall some basic principles concerning information theory and complexity measures. Let X be a random variable, for example the state of the environment as perceived in some sensory modality. Thus, X can assume several states, denoted by x, taken from some set <inline-formula id="inf1">
<mml:math id="m1">
<mml:mi mathvariant="script">X</mml:mi>
</mml:math>
</inline-formula>, with probability <inline-formula id="inf2">
<mml:math id="m2">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. Here, <inline-formula id="inf3">
<mml:math id="m3">
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2264;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>, and<disp-formula id="e1">
<mml:math id="m4">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">X</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mstyle>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>
</p>
<p>The <italic>Shannon information</italic> or <italic>entropy</italic> [<xref ref-type="bibr" rid="B13">13</xref>] of <italic>X</italic> then is<disp-formula id="e2">
<mml:math id="m5">
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>:</mml:mo>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">X</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mtext>log</mml:mtext>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>with the convention <inline-formula id="inf4">
<mml:math id="m6">
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mtext>log</mml:mtext>
<mml:mn>0</mml:mn>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>, to also include cases where <inline-formula id="inf5">
<mml:math id="m7">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> for some <italic>x</italic>. <inline-formula id="inf6">
<mml:math id="m8">
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> quantifies the <italic>reduction of uncertainty</italic> when we initially only know the probabilities <inline-formula id="inf7">
<mml:math id="m9">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> and then observe which of the possible values of <italic>x</italic> actually occurs. When we have <italic>N</italic> possible events, the entropy can be at most <inline-formula id="inf8">
<mml:math id="m10">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mtext>log</mml:mtext>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, and this value is only achieved if all states have the same probability <inline-formula id="inf9">
<mml:math id="m11">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>. In that case, we can learn most from an observation of the actual state, as our uncertainty had been highest before the observation. When the probabilities are different, then <inline-formula id="inf10">
<mml:math id="m12">
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> becomes lower. Now assume that we have two random variables <inline-formula id="inf11">
<mml:math id="m13">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, with joint probabilities <inline-formula id="inf12">
<mml:math id="m14">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> for the corresponding states, as well as marginal probabilities<disp-formula id="e3">
<mml:math id="m15">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3be;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>&#x3be;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>and</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3be;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3be;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>Again, the information<disp-formula id="e4">
<mml:math id="m16">
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:munder>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mtext>log</mml:mtext>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>is highest when all possible pairs <inline-formula id="inf13">
<mml:math id="m17">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> occur with the same probability. In particular, in that case<disp-formula id="e5">
<mml:math id="m18">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>that is, the joint probability is simply the product of the individual probabilities. This means that the two random variables <inline-formula id="inf14">
<mml:math id="m19">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are independent of each other. And when (4) is maximal, again the marginal probabilities <inline-formula id="inf15">
<mml:math id="m20">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mo>.</mml:mo>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf16">
<mml:math id="m21">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mo>.</mml:mo>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> have to be constant, that is, independent of the particular value of <inline-formula id="inf17">
<mml:math id="m22">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> or <inline-formula id="inf18">
<mml:math id="m23">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>,&#x20;resp.</p>
<p>Now, there are two possibilities to decrease the entropy <inline-formula id="inf19">
<mml:math id="m24">
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> in <xref ref-type="disp-formula" rid="e4">Eq. 4</xref>. One possibility consists in keeping the product structure (5), but varying the marginal probabilities. The other possibility consists in keeping the marginal probabilities <inline-formula id="inf20">
<mml:math id="m25">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf21">
<mml:math id="m26">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, but introducing correlations between them so that <italic>p</italic> no longer is a product. Let us consider the simplest nontrivial example. Both <inline-formula id="inf22">
<mml:math id="m27">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf23">
<mml:math id="m28">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> can only assume two possible states, denoted by 0 and 1. Thus, there are four combinations, <inline-formula id="inf24">
<mml:math id="m29">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>0,0</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1,0</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>0,1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1,1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. When each of them occurs with probability <inline-formula id="inf25">
<mml:math id="m30">
<mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, then 5) holds, and <inline-formula id="inf26">
<mml:math id="m31">
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>. Also, all marginals are <inline-formula id="inf27">
<mml:math id="m32">
<mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> in this case. We could then vary the marginals <inline-formula id="inf28">
<mml:math id="m33">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf29">
<mml:math id="m34">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. One example is where one of the two variables, say <inline-formula id="inf30">
<mml:math id="m35">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, becomes deterministic, for instance <inline-formula id="inf31">
<mml:math id="m36">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>, but the other assumes its two values with equal probability <inline-formula id="inf32">
<mml:math id="m37">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. Then <inline-formula id="inf33">
<mml:math id="m38">
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>. The uncertainty has been reduced by knowing the state of the variable <inline-formula id="inf34">
<mml:math id="m39">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> in advance. When both variables are deterministic, the entropy <inline-formula id="inf35">
<mml:math id="m40">
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> becomes 0. Now let us consider the other possibility for reducing <inline-formula id="inf36">
<mml:math id="m41">
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, introducing correlations. Again, we consider an extreme possibility, where only the two combinations <inline-formula id="inf37">
<mml:math id="m42">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>0,0</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf38">
<mml:math id="m43">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1,1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> can occur, both of them with probability <inline-formula id="inf39">
<mml:math id="m44">
<mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. Then <inline-formula id="inf40">
<mml:math id="m45">
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> again, but 5) no longer holds, even though all marginals are still <inline-formula id="inf41">
<mml:math id="m46">
<mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. In fact, this is the most extreme case for the failure of <xref ref-type="disp-formula" rid="e5">Eq. 5</xref> (of course, the correlations could also be such that only the cases <inline-formula id="inf42">
<mml:math id="m47">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>0,1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf43">
<mml:math id="m48">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1,0</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> occur, but this makes no formal difference). In a sense, we have most structure here. Ay [<xref ref-type="bibr" rid="B14">14</xref>] therefore considers this as the most complex situation in the present example, as the entropy is still relatively large, but we are as far away as possible from a product distribution (5). The distribution where only <inline-formula id="inf44">
<mml:math id="m49">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>0,0</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf45">
<mml:math id="m50">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1,1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> can occur is distinguished from a product distribution by the fact that it possesses some regularity that can be expressed by a rule. In this&#x20;example, the rule is simply that the two variables always have to assume the same state. For the distribution that only allows <inline-formula id="inf46">
<mml:math id="m51">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>0,1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf47">
<mml:math id="m52">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1,0</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, the rule would be that the two variables always assume opposite states. In general, quantified notions of&#x20;complexity measure to what extent the regularities of a structure allow for a compressed description, see for instance [<xref ref-type="bibr" rid="B15">15</xref>&#x2013;<xref ref-type="bibr" rid="B17">17</xref>].</p>
<p>Returning to our concrete setting, with more than two random variables, this construction needs to be refined. When, for instance, we have three variables <inline-formula id="inf48">
<mml:math id="m53">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, we could again have a product distribution<disp-formula id="e6">
<mml:math id="m54">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>with hopefully obvious notation. We can then again search for other probability distributions that are as far away as possible from such a product distribution. But it becomes more interesting, if we also consider the intermediate class of probability distributions where we admit only pairwise correlations between two of the three variables, but no triple ones. The most complex structures then should be those that are very different from the product ones, but also from those with only pairwise correlations. Obviously, this principle can be extended to an arbitrary number of variables. This has been systematically developed in [<xref ref-type="bibr" rid="B18">18</xref>] (see also the systematic exposition in [<xref ref-type="bibr" rid="B19">19</xref>]), and a class of complexity measures has been introduced that includes the particular measure of [<xref ref-type="bibr" rid="B20">20</xref>] and gives it a new interpretation. Since such measures are fundamental for the approach to consciousness of [<xref ref-type="bibr" rid="B6">6</xref>, <xref ref-type="bibr" rid="B21">21</xref>] and since for our purposes also a second aspect, the role of memory, will be important, we shall now sketch the approach of [<xref ref-type="bibr" rid="B18">18</xref>] (see also [<xref ref-type="bibr" rid="B17">17</xref>]) which includes and unifies both aspects. We first need the concept of the Kullback-Leibler divergence for two probability distributions <italic>p</italic> and <italic>q</italic> on <inline-formula id="inf49">
<mml:math id="m55">
<mml:mi mathvariant="script">X</mml:mi>
</mml:math>
</inline-formula>,<disp-formula id="e7">
<mml:math id="m56">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mo>&#x2225;</mml:mo>
<mml:mi>q</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>:</mml:mo>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">X</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mtext>log</mml:mtext>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mfrac>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>q</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>where we require that whenever <inline-formula id="inf50">
<mml:math id="m57">
<mml:mrow>
<mml:mi>q</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> for some <italic>x</italic>, then also <inline-formula id="inf51">
<mml:math id="m58">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>; if that requirement is not satisfied, we put <inline-formula id="inf52">
<mml:math id="m59">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mo>&#x2225;</mml:mo>
<mml:mi>q</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x221e;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>. This is positive, i.e.,<disp-formula id="e8">
<mml:math id="m60">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mo>&#x2225;</mml:mo>
<mml:mi>q</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3e;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mi>i</mml:mi>
<mml:mi>f</mml:mi>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mi>p</mml:mi>
<mml:mo>&#x2260;</mml:mo>
<mml:mi>q</mml:mi>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(8)</label>
</disp-formula>
</p>
<p>More generally, we can look at the case where we have an additional random variable <italic>Y</italic> with state space <inline-formula id="inf53">
<mml:math id="m61">
<mml:mi mathvariant="script">Y</mml:mi>
</mml:math>
</inline-formula>. We then have joint probabilities <inline-formula id="inf54">
<mml:math id="m62">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> for the simultaneous realization of the value <italic>x</italic> of <italic>X</italic> and the value <italic>y</italic> of <italic>Y</italic>, as well as the marginals <inline-formula id="inf55">
<mml:math id="m63">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> for the occurrence of <italic>x</italic> and <inline-formula id="inf56">
<mml:math id="m64">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> for that of <italic>y</italic> (from now on, in contrast to the more careful notation of <xref ref-type="disp-formula" rid="e6">Eq. 6</xref>, we use the same letter <italic>p</italic> here, although the distributions of <inline-formula id="inf57">
<mml:math id="m65">
<mml:mrow>
<mml:mi>X</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>Y</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf58">
<mml:math id="m66">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>X</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>Y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> live on different spaces). We have<disp-formula id="e9">
<mml:math id="m67">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mi>y</mml:mi>
</mml:munder>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>and</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mi>x</mml:mi>
</mml:munder>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(9)</label>
</disp-formula>
</p>
<p>Importantly, the joint distribution <inline-formula id="inf59">
<mml:math id="m68">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is in general different from the product <inline-formula id="inf60">
<mml:math id="m69">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> of the marginals, due to correlations between <italic>X</italic> and <italic>Y</italic>. This is quantified by the <italic>mutual information</italic>
<disp-formula id="e10">
<mml:math id="m70">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>I</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>X</mml:mi>
<mml:mo>:</mml:mo>
<mml:mi>Y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2225;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(10)</label>
</disp-formula>
</p>
<p>When <italic>X</italic> and <italic>Y</italic> are independent, that is, <inline-formula id="inf61">
<mml:math id="m71">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> for all <inline-formula id="inf62">
<mml:math id="m72">
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, then the mutual information <inline-formula id="inf63">
<mml:math id="m73">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>I</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>X</mml:mi>
<mml:mo>:</mml:mo>
<mml:mi>Y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> vanishes, because in that case, observing the value of <italic>X</italic> will not reduce our uncertainty about <italic>Y</italic>, and conversely. If there are dependencies, however, then <inline-formula id="inf64">
<mml:math id="m74">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>I</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>X</mml:mi>
<mml:mo>:</mml:mo>
<mml:mi>Y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3e;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>, and one variable provides information about the other. We also note that in contrast to the Kullback-Leibler divergence <inline-formula id="inf65">
<mml:math id="m75">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2225;</mml:mo>
<mml:mi>q</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> which in general is not symmetric between <italic>p</italic> and <italic>q</italic>, the mutual information <inline-formula id="inf66">
<mml:math id="m76">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>I</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>X</mml:mi>
<mml:mo>:</mml:mo>
<mml:mi>Y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is symmetric between <italic>X</italic> and <italic>Y</italic>. That is, if observing <italic>X</italic> provides information about <italic>Y</italic>, then also the converse holds, and the amount of information is the same in either direction.</p>
<p>We can also interpret the product distribution <inline-formula id="inf67">
<mml:math id="m77">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> as the result of projecting our original distribution <inline-formula id="inf68">
<mml:math id="m78">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> onto the simpler class of product distributions. That is, among all product distributions of the form <inline-formula id="inf69">
<mml:math id="m79">
<mml:mrow>
<mml:mi>q</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>q</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, the production distribution <inline-formula id="inf70">
<mml:math id="m80">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is that for which the divergence <inline-formula id="inf71">
<mml:math id="m81">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2225;</mml:mo>
<mml:mi>q</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>q</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is smallest. That is, the product distribution <inline-formula id="inf72">
<mml:math id="m82">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> preserves as much information about <inline-formula id="inf73">
<mml:math id="m83">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> as is possible for a product distribution. Also, the product distribution <inline-formula id="inf74">
<mml:math id="m84">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> has higher entropy than <inline-formula id="inf75">
<mml:math id="m85">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, unless the latter is already a product distribution, because <inline-formula id="inf76">
<mml:math id="m86">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> ignores all the information that one variable has about the other. In fact, <inline-formula id="inf77">
<mml:math id="m87">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> has the highest entropy among all distributions with the same marginals as <inline-formula id="inf78">
<mml:math id="m88">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>This principle can be iterated. When we have distribution <inline-formula id="inf79">
<mml:math id="m89">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> for the values of three random variables, we can not only look at the product distribution <inline-formula id="inf80">
<mml:math id="m90">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>z</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, but also at those of the form <inline-formula id="inf81">
<mml:math id="m91">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>z</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf82">
<mml:math id="m92">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> or <inline-formula id="inf83">
<mml:math id="m93">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, that is, where only correlations between at most two of the variables are allowed.</p>
<p>For the general principle, we assume a state set <inline-formula id="inf84">
<mml:math id="m94">
<mml:mi mathvariant="script">V</mml:mi>
</mml:math>
</inline-formula> that consists of the possible values of <italic>N</italic> variables <inline-formula id="inf85">
<mml:math id="m95">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. We let <inline-formula id="inf86">
<mml:math id="m96">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> be the family of subsets of <inline-formula id="inf87">
<mml:math id="m97">
<mml:mi mathvariant="script">V</mml:mi>
</mml:math>
</inline-formula> with <inline-formula id="inf88">
<mml:math id="m98">
<mml:mrow>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> elements, from which we get the set of probability distributions <inline-formula id="inf89">
<mml:math id="m99">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x2130;</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> with dependencies of order <inline-formula id="inf90">
<mml:math id="m100">
<mml:mrow>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>. Thus, <inline-formula id="inf91">
<mml:math id="m101">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x2130;</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the family of distributions that are simply the products <inline-formula id="inf92">
<mml:math id="m102">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x22ef;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> of their marginals. In particular, for a probability distribution in this family, there are no correlations between the probabilities of two or more of the variables. In <inline-formula id="inf93">
<mml:math id="m103">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x2130;</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, we then allow for pairwise correlations, but no triple or higher order ones. We can also consider other families of subsets of <inline-formula id="inf94">
<mml:math id="m104">
<mml:mi mathvariant="script">V</mml:mi>
</mml:math>
</inline-formula> and the corresponding probability distributions. For instance, when <inline-formula id="inf95">
<mml:math id="m105">
<mml:mi mathvariant="script">V</mml:mi>
</mml:math>
</inline-formula> is the ordered set of integers <inline-formula id="inf96">
<mml:math id="m106">
<mml:mrow>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, one could consider the family of those subsets that consist of uninterrupted strings of length <inline-formula id="inf97">
<mml:math id="m107">
<mml:mrow>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>. For example, for <inline-formula id="inf98">
<mml:math id="m108">
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>, we would consider distributions of the form <inline-formula id="inf99">
<mml:math id="m109">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> or <inline-formula id="inf100">
<mml:math id="m110">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>We consider the hierarchy<disp-formula id="e11">
<mml:math id="m111">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x2286;</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x2286;</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>&#x2286;</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2286;</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
<mml:mo>:</mml:mo>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mn>2</mml:mn>
<mml:mo>&#x0394;</mml:mo>
</mml:msup>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(11)</label>
</disp-formula>
</p>
<p>We let <inline-formula id="inf101">
<mml:math id="m112">
<mml:mrow>
<mml:msup>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> be the projection of <italic>p</italic> onto <inline-formula id="inf102">
<mml:math id="m113">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x2130;</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. This means that for a distribution <italic>p</italic>, we seek that distribution <inline-formula id="inf103">
<mml:math id="m114">
<mml:mrow>
<mml:msup>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> for which <inline-formula id="inf104">
<mml:math id="m115">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mo>&#x2225;</mml:mo>
<mml:msup>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is smallest. For instance, <inline-formula id="inf105">
<mml:math id="m116">
<mml:mrow>
<mml:msup>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is the product distribution with the same marginals as&#x20;<italic>p</italic>.</p>
<p>These projections are related by the Pythagoras relation<disp-formula id="e12">
<mml:math id="m117">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>&#x2113;</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2225;</mml:mo>
<mml:msup>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>m</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>m</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x2113;</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2225;</mml:mo>
<mml:msup>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(12)</label>
</disp-formula>for <inline-formula id="inf106">
<mml:math id="m118">
<mml:mrow>
<mml:mi>&#x2113;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>m</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf107">
<mml:math id="m119">
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>&#x2113;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>. In particular,<disp-formula id="e13">
<mml:math id="m120">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mo>&#x2225;</mml:mo>
<mml:msup>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2225;</mml:mo>
<mml:msup>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(13)</label>
</disp-formula>
</p>
<p>The Pythagoras relation <xref ref-type="disp-formula" rid="e12">Eq. 12</xref> implies that we can leave out or insert intermediate steps into our hierarchy <xref ref-type="disp-formula" rid="e11">Eq. 11</xref> of projections, without changing the final result. That is, instead of first projecting onto <inline-formula id="inf108">
<mml:math id="m121">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, then onto <inline-formula id="inf109">
<mml:math id="m122">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and so on, until we finally project onto <inline-formula id="inf110">
<mml:math id="m123">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, we could also directly project onto <inline-formula id="inf111">
<mml:math id="m124">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, and the end result will be the&#x20;same.</p>
<p>We can then introduce the <italic>complexity measure</italic> of [<xref ref-type="bibr" rid="B18">18</xref>] with weight vector <inline-formula id="inf112">
<mml:math id="m125">
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi mathvariant="normal">&#x211d;</mml:mi>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>.<disp-formula id="e14">
<mml:math id="m126">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>&#x3b1;</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>:</mml:mo>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mo>&#x2225;</mml:mo>
<mml:msup>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mi>&#x3b2;</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2225;</mml:mo>
<mml:msup>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(14)</label>
</disp-formula>where <inline-formula id="inf113">
<mml:math id="m127">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b2;</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:msubsup>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>k</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mi>&#x2113;</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
</inline-formula> because of <xref ref-type="disp-formula" rid="e12">Eq. 12</xref> or <xref ref-type="disp-formula" rid="e13">Eq.&#x20;13</xref>.</p>
<p>
<inline-formula id="inf114">
<mml:math id="m128">
<mml:mrow>
<mml:msup>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is the distribution of highest entropy among all those with the same correlations of order <inline-formula id="inf115">
<mml:math id="m129">
<mml:mrow>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> as&#x20;<italic>p</italic>.</p>
<p>
<xref ref-type="disp-formula" rid="e14">Eq. 14</xref> is a weighted sum of the higher order correlation structure. We can then choose the weights. For instance, when we choose <inline-formula id="inf116">
<mml:math id="m130">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mi>k</mml:mi>
<mml:mi>N</mml:mi>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula>, we obtain the Tononi-Sporns-Edelman (TSE) complexity&#x20;[<xref ref-type="bibr" rid="B20">20</xref>].</p>
<p>The TSE measure was introduced to capture the interplay between differentiation and integration, and it served as the basis of the consciousness measure of [<xref ref-type="bibr" rid="B21">21</xref>]. The idea is that a conscious state should be one that is capable of many distinctions between different possibilities, but at the same time integrates the different variables into a coherent&#x20;whole.</p>
<p>Thus, the preceding information theoretical considerations can capture the interplay of differentiation and integration in neural (and other) dynamics. It can also capture the temporal aspects of memory utilization, as we shall now explain. We need the notion of a <italic>Markov process</italic>. We consider a sequence <inline-formula id="inf117">
<mml:math id="m131">
<mml:mrow>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> of random variables, called a stochastic process, where the index <italic>n</italic> of <inline-formula id="inf118">
<mml:math id="m132">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> now stands for discrete time <inline-formula id="inf119">
<mml:math id="m133">
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="normal">&#x2124;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> (an analogous notion can be developed for continuous time <inline-formula id="inf120">
<mml:math id="m134">
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="normal">&#x211d;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, but for simplicity, we only discuss the case of discrete time here). We say that this stochastic process has the Markov property if for all <italic>n</italic> and all values <inline-formula id="inf121">
<mml:math id="m135">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>m</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> of the <inline-formula id="inf122">
<mml:math id="m136">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>m</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>.<disp-formula id="e15">
<mml:math id="m137">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x7c;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x7c;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(15)</label>
</disp-formula>that is, if at time <italic>n</italic>, taken as the present, the value at the future time <inline-formula id="inf123">
<mml:math id="m138">
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> depends only on the present state, but not on any further values from past times. (Here, <inline-formula id="inf124">
<mml:math id="m139">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x7c;</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is a conditional property, the probability that a random variable assumes the state <italic>x</italic> when the state <italic>y</italic> of some other variable is assumed or known.) Thus, a Markov process does not have or does not need a memory.</p>
<p>More generally, a <italic>k</italic>th order Markov process has the property<disp-formula id="e16">
<mml:math id="m140">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x7c;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x7c;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(16)</label>
</disp-formula>that is, for the best possible prediction, it may need <inline-formula id="inf125">
<mml:math id="m141">
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> memory steps in addition to the knowledge of the present state, but not more. A 0th order Markov is of course one that does not need any memory, and not even the knowledge of the current state, i.e.,<disp-formula id="e17">
<mml:math id="m142">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x7c;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(17)</label>
</disp-formula>
</p>
<p>Tossing a fair coin, for instance, is such a process.</p>
<p>For technical reasons, we assume that the process <inline-formula id="inf126">
<mml:math id="m143">
<mml:mrow>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> is stationary in the sense that the marginals are invariant under time shifts, that is, for all <inline-formula id="inf127">
<mml:math id="m144">
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>&#x2113;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>.<disp-formula id="e18">
<mml:math id="m145">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>m</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x2113;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x2113;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x2113;</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(18)</label>
</disp-formula>
</p>
<p>We then consider subprocesses of length <italic>N</italic>. By stationarity, we may take <inline-formula id="inf128">
<mml:math id="m146">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. We denote its probability distribution by <inline-formula id="inf129">
<mml:math id="m147">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> As in <xref ref-type="disp-formula" rid="e11">Eq. 11</xref>, we consider the hierarchy<disp-formula id="e19">
<mml:math id="m148">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2286;</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2286;</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>&#x2286;</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(19)</label>
</disp-formula>where <inline-formula id="inf130">
<mml:math id="m149">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> now consists of the Markov processes of order <italic>k</italic>. As before, we let <inline-formula id="inf131">
<mml:math id="m150">
<mml:mrow>
<mml:msubsup>
<mml:mi>p</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> be the projection of <inline-formula id="inf132">
<mml:math id="m151">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> onto <inline-formula id="inf133">
<mml:math id="m152">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x2130;</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x1d505;</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. And we have the Pythagoras relation <xref ref-type="disp-formula" rid="e12">Eq. 12</xref>, i.e.,<disp-formula id="e20">
<mml:math id="m153">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>p</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>&#x2113;</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2225;</mml:mo>
<mml:msubsup>
<mml:mi>p</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>m</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>m</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x2113;</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>p</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2225;</mml:mo>
<mml:msubsup>
<mml:mi>p</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(20)</label>
</disp-formula>for <inline-formula id="inf134">
<mml:math id="m154">
<mml:mrow>
<mml:mi>&#x2113;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>m</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf135">
<mml:math id="m155">
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>&#x2113;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>. And we can measure the complexity of the subprocess <inline-formula id="inf136">
<mml:math id="m156">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> by <xref ref-type="disp-formula" rid="e14">Eq. 14</xref>, i.e.,<disp-formula id="e21">
<mml:math id="m157">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>&#x3b1;</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
<mml:mo>&#x2225;</mml:mo>
<mml:msubsup>
<mml:mi>p</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mi>&#x3b2;</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>p</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2225;</mml:mo>
<mml:msubsup>
<mml:mi>p</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(21)</label>
</disp-formula>with weight vectors <inline-formula id="inf137">
<mml:math id="m158">
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> related by <inline-formula id="inf138">
<mml:math id="m159">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b2;</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:msubsup>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mi>k</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mi>&#x2113;</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>Now, for such stationary stochastic processes, there is a fundamental complexity measure, the <italic>excess entropy</italic> introduced by Han [<xref ref-type="bibr" rid="B22">22</xref>] and Grassberger [<xref ref-type="bibr" rid="B23">23</xref>]. It is given by<disp-formula id="e22">
<mml:math id="m160">
<mml:mrow>
<mml:munder>
<mml:mrow>
<mml:mtext>lim</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2192;</mml:mo>
<mml:mi>&#x221e;</mml:mi>
</mml:mrow>
</mml:munder>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:mfrac>
<mml:mi>k</mml:mi>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>p</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2225;</mml:mo>
<mml:msubsup>
<mml:mi>p</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(22)</label>
</disp-formula>
</p>
<p>Thus, again the complexity measure <xref ref-type="disp-formula" rid="e21">Eq. 21</xref> of [<xref ref-type="bibr" rid="B18">18</xref>] generalizes a fundamental quantity that evaluates the complexity of a process.</p>
<p>Thus, following [<xref ref-type="bibr" rid="B17">17</xref>&#x2013;<xref ref-type="bibr" rid="B19">19</xref>], we have developed complexity measures that can capture both families of processes running in parallel as in [<xref ref-type="bibr" rid="B20">20</xref>], as well as sequential processes as in [<xref ref-type="bibr" rid="B22">22</xref>, <xref ref-type="bibr" rid="B23">23</xref>]. For measuring neural brain complexity, both seem relevant, and also both can be estimated from brain recordings. In particular, for a sequential process, one may derive an estimate from recordings without spatial, but good temporal resolution, like EEG. If one accepts the thesis of [<xref ref-type="bibr" rid="B21">21</xref>] that such measures might even allow us to assess the level of consciousness, then what we have provided here enlarges the array and the scope of such measures.</p>
</sec>
<sec id="s3">
<title>3 Some Formal Aspects</title>
<p>The dynamics of the world might be a Markov process. That has two aspects.<list list-type="simple">
<list-item>
<p>1. The dynamics is not completely deterministic. There are only probabilities for future events, conditioned on the present state of the world. (This uncertainty about the future may ultimately arise from the quantum world, but that is not our topic here.)</p>
</list-item>
<list-item>
<p>2. When we know the present completely, we cannot improve our predictions for future events by using additional information about the&#x20;past.</p>
</list-item>
</list>
</p>
<p>In addition.<list list-type="simple">
<list-item>
<p>3. for us as finite beings, the information about the present state of the world is always incomplete.</p>
</list-item>
</list>
</p>
<p>And this has consequences.<list list-type="simple">
<list-item>
<p>1. While our sensory data may provide us only with some probability distribution over the actual state of the world and for possible future developments, we need to perform concrete actions, and not probability distributions over actions, as pointed out in [<xref ref-type="bibr" rid="B24">24</xref>]. That is, we need some mechanism that transforms a probability distribution into a single action. This need not necessarily invoke consciousness, as also actions that are subconsciously planned and executed are determinate, but consciousness may be important when prior experience and learned patterns are not able to directly select a unique action. Even though according to Libet&#x2019;s findings [<xref ref-type="bibr" rid="B11">11</xref>], the action itself may be decided before it becomes conscious, the process is different from subconscious routines and typically takes much longer. Consciousness may thus not be the ultimate cause, but only a witness of such an action selection mechanism, but making it conscious may at least help to guide future behavior in similar situations, that is, be an efficient mechanism for memorizing the process. Insight can be gained here from investigating the neural processes of long-term memory formation.</p>
</list-item>
<list-item>
<p>2. Even if the dynamics satisfy the Markov property, our partial representation of it will not. It is a general result that the projection or the coarse graining of a Markov process typically is no longer Markovian, see the analysis in [<xref ref-type="bibr" rid="B25">25</xref>]. That is, memory will help to make better predictions.</p>
</list-item>
</list>
</p>
<p>Let us explain in more detail why the coarsening of a Markov process need no longer be Markovian. For a simple example, take a random process on the integers where at each integer time <italic>t</italic> a state <inline-formula id="inf139">
<mml:math id="m161">
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="normal">&#x2124;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is selected, with the transition that at time <inline-formula id="inf140">
<mml:math id="m162">
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>, there are two possible successor states, <inline-formula id="inf141">
<mml:math id="m163">
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf142">
<mml:math id="m164">
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>, each attained with probability <inline-formula id="inf143">
<mml:math id="m165">
<mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. This is a Markov process. But when we now coarse grain the integers, and lump five consecutive numbers together, defining <inline-formula id="inf144">
<mml:math id="m166">
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>n</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mn>5</mml:mn>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>5</mml:mn>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1,5</mml:mn>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>2,5</mml:mn>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>4,5</mml:mn>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>5</mml:mn>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
<mml:mo>&#x2282;</mml:mo>
<mml:mi mathvariant="normal">&#x2124;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> for <inline-formula id="inf145">
<mml:math id="m167">
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="normal">&#x2124;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, we get an induced process where <inline-formula id="inf146">
<mml:math id="m168">
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>n</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> can take the values <inline-formula id="inf147">
<mml:math id="m169">
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>n</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>n</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>n</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>. But this is no longer a Markov process, because memory about the underlying process <inline-formula id="inf148">
<mml:math id="m170">
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> helps to improve the prediction. In fact, when <inline-formula id="inf149">
<mml:math id="m171">
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>5</mml:mn>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>, then we can conclude that <inline-formula id="inf150">
<mml:math id="m172">
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>n</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>n</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. And when <inline-formula id="inf151">
<mml:math id="m173">
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>5</mml:mn>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> or <inline-formula id="inf152">
<mml:math id="m174">
<mml:mrow>
<mml:mn>5</mml:mn>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>, then the value <inline-formula id="inf153">
<mml:math id="m175">
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>n</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>n</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> is not possible.</p>
<p>Another example (taken from [<xref ref-type="bibr" rid="B26">26</xref>]) shows the phenomenon even more drastically. We cast a fair die. The result at time <inline-formula id="inf154">
<mml:math id="m176">
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="normal">&#x2124;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is denoted by <inline-formula id="inf155">
<mml:math id="m177">
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. The values <inline-formula id="inf156">
<mml:math id="m178">
<mml:mrow>
<mml:mn>1,2,3,4,5,6</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> occur with equal probability, independently of any prior results. This is a 0th order Markov process, because even knowledge of the current state does not improve the prediction of the next one. We now consider a derived process on <inline-formula id="inf157">
<mml:math id="m179">
<mml:mrow>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mn>0,1</mml:mn>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> where we record a 1&#xa0;at time <italic>t</italic> if <inline-formula id="inf158">
<mml:math id="m180">
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3e;</mml:mo>
<mml:mi>w</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> and a 0 else. This derived process is no longer Markovian because the more 1s we have already observed in a row, the less likely the next result is a 1 again. In fact, we cannot have more than five 1s in a row. Thus, the memory of many steps back in the past improves our abilities to predict the next result. Actually, in this example, the length of memory that can be utilized for an improvement of the prediction of the next value can be arbitrarily large. Consider two sequences of length <italic>n</italic>, 121,212... and 565,656.... They both give rise to the derived sequence 101,010... but when we have seen several consecutive 1s before that sequence, then the first one is excluded, while the second is still possible.</p>
<p>It seems that memory is fundamental for consciousness, as it can be utilized to improve predictions of future states on the basis of memorized past ones. And the length of the memory span that can be relevant can be arbitrarily large. We can still make inferences and for instance avoid dangers on the basis of events that we have experienced in our childhood. Therefore, complexity measures such as <xref ref-type="disp-formula" rid="e21">Eq. 21</xref>, <xref ref-type="disp-formula" rid="e22">Eq. 22</xref> may indeed be relevant for a quantification, although of course in practice we do not possess a stream of EEG recordings since our childhood. And a more important feature of consciousness than long-term memory might be the temporal integration of the recent past and the immediate future, as will be argued in more detail below. But the principle of quantifying that via complexity measures such as those that we have derived remains valid and applicable.</p>
</sec>
<sec id="s4">
<title>4 Integration</title>
<p>From the preceding, we conclude that a decision process that has to handle ambiguity should integrate many diverse types of data and background information, that is, involve the coordination of large parts of the brain, as in the theory of Baars [<xref ref-type="bibr" rid="B27">27</xref>&#x2013;<xref ref-type="bibr" rid="B29">29</xref>]. It need not involve, however, those regions of the brain that carry out routine motor behavior, that is, in particular, the cerebellum. It should rather activate sensory areas in the cortex, and perhaps the areas in the thalamus that are incorporated in feedback loops with those cortical areas, and in particular, the premotor and the motor cortex, and perhaps also the hippocampus, where various types of memory are located. It should also crucially include recent memories and therefore integrate a certain stretch of time, what is felt as present in our consciousness, perhaps of a duration of a second or two, as in the time-on theory of Libet [<xref ref-type="bibr" rid="B30">30</xref>]. That duration is, however, necessarily limited, as otherwise reaction times would get too&#x20;long.</p>
<p>The underlying brain activity should exhibit evidence of the integration of information. It should also, however, be able to differentiate between different types of input and different consequential actions. Tononi [<xref ref-type="bibr" rid="B21">21</xref>] has developed a corresponding measure for the brain activity, to test consciousness in patients who for whatever reason are not able to communicate directly. This measure is based on the complexity measure of Tononi, Sporns, and Edelman [<xref ref-type="bibr" rid="B20">20</xref>]. This measure has been generalized and interpreted in [<xref ref-type="bibr" rid="B18">18</xref>] as measuring the amount of higher order correlations between the elements of a dynamical system that cannot be reduced to lower order&#x20;ones.</p>
<p>Thus, even though the theories of Libet [<xref ref-type="bibr" rid="B30">30</xref>], of Baars [<xref ref-type="bibr" rid="B29">29</xref>], and of Tononi [<xref ref-type="bibr" rid="B21">21</xref>] and Koch [<xref ref-type="bibr" rid="B6">6</xref>] do not agree on many aspects of consciousness, they nevertheless share the emphasis on the integrative role of consciousness.</p>
</sec>
<sec id="s5">
<title>5 Reflexivity</title>
<p>We can never be sure that we know something, but we can be sure that we believe something. This may sound paradoxical, as knowledge is usually associated with certainty and belief with uncertainty. It is even common to define knowledge as justified belief. But a categorical distinction might be important here. We are not certain about what we believe, but only about the fact that we believe something. When the latter fact corresponds to some brain state, we are sure about our internal state, but may still question to what that state refers. When we feel pain, we are sure that we have pain, but we may err about where the pain is coming from, as most clearly demonstrated by phantom pain in amputated&#x20;limbs.</p>
<p>In a different context, that of semiotics, a sign relates a signifier and a signified, to be distinguished from a referent, but we may have the signifier without its referent. A conscious, that is, reflexive brain state can then be sure about itself, independently of what it refers&#x20;to.</p>
<p>And the integrative nature of consciousness then allows the brain to access a wide range of memories, via something like neural association mechanisms. When our arm is pinched, we associate that with other similar sensations, and the qualitative feeling of pain develops. Similarly, when we see something red, our consciousness associates this with other sensations of red that we have experienced in the past. As this is, however, not explicit, but only implicit, it leads to the qualium of red. Something like that has been taken as a definition in [<xref ref-type="bibr" rid="B31">31</xref>]. The point I want to make here is that of course, our brain does not have the capacity to explicitly recollect all prior instances of red that we have seen, and our consciousness therefore needs to compress them into something implicit, a qualium, as reflexively experienced.</p>
<p>The reflexive moment is most evidently seen in selfconsciousness. Abstractly, we have a finite system trying to represent itself in itself. Since the system is finite, it cannot contain a perfect copy of itself in itself, because it would then also have to contain the copy, and so on, leading to an infinite regress. Mathematically, only infinite sets can contain isomorphic copies of themselves as subsets. This implies (unless one wants to invoke mechanisms of quantum processes) that selfconsciousness must have some blind spot, that is, it cannot have full access to itself. Leibniz was perhaps the first philosopher aware of this problem in some sense, and he therefore postulated the existence of subconscious processes (see [<xref ref-type="bibr" rid="B32">32</xref>] for an approach to Leibniz from the perspective of contemporary science), but this problem seems to plague the theories of many later philosophers, such as Fichte, one of the key representants of idealistic philosophy. While it seems evident from the perspective of modern psychology, and in particular from that of psychoanalysis, that our self-knowledge is very incomplete, there are still some theoretical loopholes that do not make this state of affairs completely inevitable.</p>
<p>First, it may be that a system admits a complete description in a condensed or compressed form. To illustrate this with a computer science example, it is possible in principle that from a zip file, the original data can be reconstructed without any loss. But this requires that the data possess regularities that allow for a compressed description (we may think here of the notion of Kolmogorov complexity [<xref ref-type="bibr" rid="B15">15</xref>]). This example is, however, not yet complete for our purposes, because we would also have to require that the compressed file is part of the original data, to make this completely reflexive. This indicates the difficulty we are facing&#x20;here.</p>
<p>Second, we may externalize some of our memory. We may think here of such mundane devices as writing things on a sheet of paper or putting them into a computer file, as some form of external memory that we can access whenever we need. More generally, and conceptually more interestingly, this brings us to the topic of embodied cognition. It had been underestimated for a long time to what extent our cognition depends on locating sensory information in our environments, instead of memorizing it internally, and in how sophisticated a manner we are capable of integrating our environments into our cognitive processes, see for instance [<xref ref-type="bibr" rid="B33">33</xref>]. This may apply in particular to consciousness, where we observe not only the material world, but also other humans that we interact with. As an extreme example, we might go to a psychiatrist to bring our subconscious desires to our conscious attention.</p>
</sec>
<sec id="s6">
<title>6 Self</title>
<p>There is evidence that some higher animals have a notion of self, but since they do not possess language and cannot communicate that notion and thereby relate it to the notions of self of others, their self-identity may be ultimately very different from ours. In particular, it may have evolutionarily developed in directions not accessible to&#x20;us.</p>
<p>Humans, however, seem to develop such a notion of self only indirectly, through their interactions with others. A human self distinguishes himself or herself from other selfs, and a child learns to say &#x201c;I&#x201d; because it hears others saying &#x201c;I.&#x201d; That is, when others act, behave, and speak as integrated selfs, the child may conclude that it also is a self itself.</p>
<p>This cannot be quite so simple, however. To develop a notion of self-identity, also resonances are needed. That means that the child has to learn that it can anticipate and control the sensory consequences of its own actions (such a principle is formalized in a different context in [<xref ref-type="bibr" rid="B34">34</xref>]). That is, it develops the ability to act in such a manner that it can produce particular consequences. When it pinches its arm, it feels pain at the pinched spot. When it pinches somebody else&#x2019;s arm or an inanimate object or whatever, no such reaction will occur. In that manner, it learns what belongs to itself. But as argued before, the integrated notion of self-identity may require the interaction with others. Here, we should also note the systematic psychological theory of Prinz [<xref ref-type="bibr" rid="B35">35</xref>] of the social construction of the notion of&#x20;self.</p>
</sec>
<sec id="s7">
<title>7 Summary and Conclusion</title>
<p>As Libet&#x2019;s experiments [<xref ref-type="bibr" rid="B11">11</xref>] show, neuronal activity needs to build up before a decision or a process or whatever can become conscious. When one assumes, as I do here, that consciousness emerges from some neurophysiological substratum, that is, there are neuronal, and therefore ultimately biophysical, processes underlying conscious experience, then this is what one should expect. But it would be rash to conclude from this that consciousness is simply an illusion that, with some temporal delay, accompanies a deterministic physical process.</p>
<p>It rather seems that the function of consciousness is to integrate on one hand synchronous and probabilistic information from various sources, both external and internal, into a coherent percept and a concrete action that is no longer probabilistic. And, on the other hand, it temporally integrates the recent past and the immediate future with the present into some extended present of a duration of perhaps a second or so. There are neuronal mechanisms for both the integration of the past and of the present. Short-term memory may be based on neuronal reverberations that are repeated in short periodic cycles, and it may be integrated with long-term memory that is perhaps stored in synaptic weights between neurons. Moreover, synaptic learning rules of STDP type can be interpreted as facilitating the anticipation of consequences of stimuli before those consequences actually occur, as analyzed in [<xref ref-type="bibr" rid="B36">36</xref>]. That is, we understand the neuronal basis of an anticipation of an immediate future. Such temporal integration seems to underlie the extended present characteristic of consciousness and the stream of consciousness. The richness of consciousness and the selection of a specific action become enhanced from an integration of a wide range of inputs.</p>
<p>Selfconsciousness, that is, the feeling of one&#x2019;s identity as distinct from others, emerges from resonances between perceptions and actions, and from mirroring oneself in others in social settings. Finally, the fact that a wealth of prior experiences is condensed into qualia is some compression mechanism. And we cannot be conscious of our full brain activity because a system cannot represent itself isomorphically in itself.</p>
<p>These are our main conclusions.<list list-type="simple">
<list-item>
<p>
<bold>Thesis 1.</bold> Consciousness integrates information that is distributed in the brain and in the immediate environment and that includes the recent past and anticipates the near future, on a timescale that is adapted to the requirements for reactions to external stimuli, in order to select a single action on the basis of a probability distribution of possible stimulus interpretations.</p>
</list-item>
<list-item>
<p>
<bold>Thesis 2.</bold> The preceding is quantifiable by complexity measures.</p>
</list-item>
<list-item>
<p>
<bold>Thesis 3.</bold> The development of consciousness depends on resonances between sensory inputs and actions, and selfconsciousness therefore can only emerge in the context of interactions with other conscious individuals.</p>
</list-item>
<list-item>
<p>
<bold>Thesis 4.</bold> The feeling of qualia is the result of an efficient compression of information about prior experiences.</p>
</list-item>
</list>
</p>
</sec>
</body>
<back>
<sec id="s8">
<title>Author Contributions</title>
<p>The author confirms being the sole contributor of this work and has approved it for publication.</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of Interest</title>
<p>The author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s10" sec-type="disclaimer">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>I thank Lukas Barth for several discussions and critical comments on my manuscript.</p>
</ack>
<fn-group>
<fn id="fn1">
<label>1</label>
<p>It could be and has been argued that one has to be careful in drawing strong implications about consciousness from Libet&#x2019;s experiments. The decision about whether or when their arm is raised made in the experiment is not really important for the persons involved. It could well be that in important and consequential decisions, the temporal order can become reversed.</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<label>1.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Koch</surname>
<given-names>C</given-names>
</name>
</person-group>. <article-title>The Quest for Consciousness</article-title>. In: <source>A Neurobiological Approach</source>. <publisher-loc>Englewood</publisher-loc>: <publisher-name>Foreword by F.Crick, Roberts and Co.</publisher-name> (<year>2004</year>). </citation>
</ref>
<ref id="B2">
<label>2.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gopnik</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Meltzoff</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Kuhl</surname>
<given-names>P</given-names>
</name>
</person-group>. <source>The Scientist in the Crib: Minds, Brains, and How Children Learn</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>William Morrow</publisher-name> (<year>1999</year>).</citation>
</ref>
<ref id="B3">
<label>3.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Dehaine</surname>
<given-names>S</given-names>
</name>
</person-group>. <source>Consciousness and the Brain</source>. <publisher-name>Penguin</publisher-name> (<year>2014</year>).</citation>
</ref>
<ref id="B4">
<label>4.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Putnam</surname>
<given-names>H</given-names>
</name>
</person-group>. <article-title>The Meaning of &#x201c;Meaning&#x201d;</article-title>. In: <person-group person-group-type="editor">
<name>
<surname>Gunderson</surname>
<given-names>K</given-names>
</name>
</person-group>, editor. <source>Language, Mind, and Knowledge, Minnesota Stud</source>. <publisher-name>Phil. Sci.</publisher-name> (<year>1975</year>). <publisher-loc>New York</publisher-loc>, p. <fpage>131</fpage>&#x2013;<lpage>93</lpage>. </citation>
</ref>
<ref id="B5">
<label>5.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Putnam</surname>
<given-names>H</given-names>
</name>
</person-group>. <source>Representation and Reality</source>. <publisher-loc>London</publisher-loc>: <publisher-name>MIT Press, Cambridge/Mass.</publisher-name> (<year>1988</year>).</citation>
</ref>
<ref id="B6">
<label>6.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Koch</surname>
<given-names>C</given-names>
</name>
</person-group>. <source>The Feeling of Life Itself</source>. <publisher-loc>Cambridge/ Mass</publisher-loc>: <publisher-name>MIT Press</publisher-name> (<year>2019</year>).</citation>
</ref>
<ref id="B7">
<label>7.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Jost</surname>
<given-names>J</given-names>
</name>
</person-group>. <source>Biologie und Mathematik</source>. <publisher-name>Berlin: Springer</publisher-name> (<year>2019</year>).</citation>
</ref>
<ref id="B8">
<label>8.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Sporns</surname>
<given-names>O</given-names>
</name>
</person-group>. <source>Discovering the Human Connectome</source>. <publisher-loc>Cambridge/Mass</publisher-loc>: <publisher-name>MIT Press</publisher-name> (<year>2012</year>).</citation>
</ref>
<ref id="B9">
<label>9.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nagel</surname>
<given-names>T</given-names>
</name>
</person-group>. <article-title>What Is it like to Be a Bat?</article-title> <source>Philos Rev</source> (<year>1974</year>) <volume>83</volume>:<fpage>435</fpage>&#x2013;<lpage>50</lpage>. <pub-id pub-id-type="doi">10.2307/2183914</pub-id> </citation>
</ref>
<ref id="B10">
<label>10.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gould</surname>
<given-names>SJ</given-names>
</name>
</person-group>. <source>The Structure of Evolutionary Theory</source>. <publisher-loc>Cambridge/Mass</publisher-loc>: <publisher-name>Belknap Press Harvard University Press</publisher-name> (<year>2002</year>).</citation>
</ref>
<ref id="B11">
<label>11.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Libet</surname>
<given-names>B</given-names>
</name>
</person-group>. <article-title>Unconscious Cerebral Initiative and the Role of Conscious Will in Voluntary Action</article-title>. <source>Behav Brain Sci</source> (<year>1985</year>) <volume>8</volume>:<fpage>529</fpage>&#x2013;<lpage>39</lpage>. <pub-id pub-id-type="doi">10.1017/s0140525x00044903</pub-id> </citation>
</ref>
<ref id="B12">
<label>12.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Jost</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Sensorimotor Contingencies and the Dynamical Creation of Structural Relations Underlying Percepts</article-title>. In: <person-group person-group-type="editor">
<name>
<surname>Engel</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Friston</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Kragic</surname>
<given-names>D</given-names>
</name>
</person-group>, editors. <source>Str&#xfc;ngmann Forum Reports 18, The Pragmatic Turn: Toward Action-Oriented Views in Cognitive Science</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>MIT Press</publisher-name> (<year>2016</year>). p. <fpage>121</fpage>&#x2013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.7551/mitpress/9780262034326.003.0008</pub-id> </citation>
</ref>
<ref id="B13">
<label>13.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Shannon</surname>
<given-names>C</given-names>
</name>
</person-group>. <article-title>The Mathematical Theory of Communication</article-title>. In: <person-group person-group-type="editor">
<name>
<surname>Blahut</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Hajek</surname>
<given-names>B</given-names>
</name>
</person-group>, editors. <source>The Mathematical Theory of Communication</source>. <publisher-name>Univ. Illinois Press</publisher-name> (<year>1948</year>). p. <fpage>29</fpage>&#x2013;<lpage>125</lpage>. </citation>
</ref>
<ref id="B14">
<label>14.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ay</surname>
<given-names>N</given-names>
</name>
</person-group>. <article-title>An Information-Geometric Approach to a Theory of Pragmatic Structuring</article-title>. <source>Ann Prob</source> (<year>2002</year>) <volume>30</volume>:<fpage>416</fpage>&#x2013;<lpage>36</lpage>. <pub-id pub-id-type="doi">10.1214/aop/1020107773</pub-id> </citation>
</ref>
<ref id="B15">
<label>15.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Vit&#xe1;nyi</surname>
<given-names>P</given-names>
</name>
</person-group>. <source>An Introduction to Kolmogorov Complexity and its Applications</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>Springer</publisher-name> (<year>1997</year>).</citation>
</ref>
<ref id="B16">
<label>16.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Moore</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Mertens</surname>
<given-names>S</given-names>
</name>
</person-group>. <source>The Nature of Computation</source>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford Univ.Press</publisher-name> (<year>2011</year>).</citation>
</ref>
<ref id="B17">
<label>17.</label>
<citation citation-type="other">
<person-group person-group-type="author">
<name>
<surname>Ay</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Bertschinger</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Jost</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Olbrich</surname>
<given-names>E</given-names>
</name>
<name>
<surname>Rauh</surname>
<given-names>J</given-names>
</name>
</person-group> (<year>2021</year>). <source>Information and Complexity, or: Where Is the Information?</source> <publisher-name>Springer Lecture Notes Math</publisher-name>. <comment>to appear</comment>.</citation>
</ref>
<ref id="B18">
<label>18.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ay</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Olbrich</surname>
<given-names>E</given-names>
</name>
<name>
<surname>Bertschinger</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Jost</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>A Geometric Approach to Complexity</article-title>. <source>Chaos</source> (<year>2011</year>) <volume>21</volume>:<fpage>037103</fpage>. <pub-id pub-id-type="doi">10.1063/1.3638446</pub-id> </citation>
</ref>
<ref id="B19">
<label>19.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ay</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Jost</surname>
<given-names>J</given-names>
</name>
<name>
<surname>L&#xea;</surname>
<given-names>HV</given-names>
</name>
<name>
<surname>Schwachh&#xf6;fer</surname>
<given-names>L</given-names>
</name>
</person-group>. <article-title>Information Geometry</article-title> In: <source>Ergebnisse der Mathematik</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name> (<year>2017</year>). </citation>
</ref>
<ref id="B20">
<label>20.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tononi</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Sporns</surname>
<given-names>O</given-names>
</name>
<name>
<surname>Edelman</surname>
<given-names>GM</given-names>
</name>
</person-group>. <article-title>A Measure for Brain Complexity: Relating Functional Segregation and Integration in the Nervous System</article-title>. <source>Proc Natl Acad Sci</source> (<year>1994</year>) <volume>91</volume>:<fpage>5033</fpage>&#x2013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.91.11.5033</pub-id> </citation>
</ref>
<ref id="B21">
<label>21.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tononi</surname>
<given-names>G</given-names>
</name>
</person-group>. <article-title>Consciousness as Integrated Information: a Provisional Manifesto</article-title>. <source>Biol Bull</source> (<year>2008</year>) <volume>215</volume>:<fpage>216</fpage>&#x2013;<lpage>42</lpage>. <pub-id pub-id-type="doi">10.2307/25470707</pub-id> </citation>
</ref>
<ref id="B22">
<label>22.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Han</surname>
<given-names>TS</given-names>
</name>
</person-group>. <article-title>Nonnegative Entropy Measures of Multivariate Symmetric Correlations</article-title>. <source>Inf Control</source> (<year>1978</year>) <volume>36</volume>:<fpage>133</fpage>&#x2013;<lpage>56</lpage>. <pub-id pub-id-type="doi">10.1016/s0019-9958(78)90275-9</pub-id> </citation>
</ref>
<ref id="B23">
<label>23.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Grassberger</surname>
<given-names>P</given-names>
</name>
</person-group>. <article-title>Toward a Quantitative Theory of Self-Generated Complexity</article-title>. <source>Int J&#x20;Theor Phys</source> (<year>1986</year>) <volume>25</volume>(<issue>9</issue>):<fpage>907</fpage>&#x2013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.1007/bf00668821</pub-id> </citation>
</ref>
<ref id="B24">
<label>24.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Morsella</surname>
<given-names>E</given-names>
</name>
<name>
<surname>Godwin</surname>
<given-names>CA</given-names>
</name>
<name>
<surname>Jantz</surname>
<given-names>TK</given-names>
</name>
<name>
<surname>Krieger</surname>
<given-names>SC</given-names>
</name>
<name>
<surname>Gazzaley</surname>
<given-names>A</given-names>
</name>
</person-group>. <article-title>Homing in on Consciousness in the Nervous System: An Action-Based Synthesis</article-title>. <source>Behav Brain Sci</source> (<year>2016</year>) <volume>39</volume>:<fpage>e168</fpage>. <pub-id pub-id-type="doi">10.1017/S0140525X15000643</pub-id> </citation>
</ref>
<ref id="B25">
<label>25.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pfante</surname>
<given-names>O</given-names>
</name>
<name>
<surname>Bertschinger</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Olbrich</surname>
<given-names>E</given-names>
</name>
<name>
<surname>Ay</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Jost</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Comparison between Different Methods of Level Identification</article-title>. <source>Advs Complex Syst</source> (<year>2014</year>) <volume>17</volume>:<fpage>1450007</fpage>. <pub-id pub-id-type="doi">10.1142/s0219525914500076</pub-id> </citation>
</ref>
<ref id="B26">
<label>26.</label>
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Pfante</surname>
<given-names>O</given-names>
</name>
<name>
<surname>Bertschinger</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Olbrich</surname>
<given-names>E</given-names>
</name>
<name>
<surname>Ay</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Jost</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Wie findet man eine geeignete Beschreibungsebene f&#xfc;r ein komplexes System?, Jahrbuch der Max-Planck-Gesellschaft, Forschungsbericht - Max-Planck-Institut f&#xfc;r Mathematik in den Naturwissenschaften</article-title> (<year>2016</year>). <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://www.mpg.de/9821421/MPI_MIS_JB_2016?c=10583665">https://www.mpg.de/9821421/MPI_MIS_JB_2016?c&#x3d;10583665</ext-link>
</comment> (<comment>Accessed August 2, 2021</comment>). </citation>
</ref>
<ref id="B27">
<label>27.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Baars</surname>
<given-names>B</given-names>
</name>
</person-group>. <source>A Cognitive Theory of Consciousness</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name> (<year>1988</year>).</citation>
</ref>
<ref id="B28">
<label>28.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Baars</surname>
<given-names>B</given-names>
</name>
</person-group>. <source>The Theater of Consciousness</source>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford University Press</publisher-name> (<year>1997</year>).</citation>
</ref>
<ref id="B29">
<label>29.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Baars</surname>
<given-names>BJ</given-names>
</name>
</person-group>. <article-title>The Conscious Access Hypothesis: Origins and Recent Evidence</article-title>. <source>Trends Cogn Sci</source> (<year>2002</year>) <volume>6</volume>(<issue>1</issue>):<fpage>47</fpage>&#x2013;<lpage>52</lpage>. <pub-id pub-id-type="doi">10.1016/s1364-6613(00)01819-2</pub-id> </citation>
</ref>
<ref id="B30">
<label>30.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Libet</surname>
<given-names>B</given-names>
</name>
</person-group>. <source>Neurophysiology of Consciousness</source>. <publisher-loc>Birkh&#x00E4;user, Boston</publisher-loc>: (<year>1993</year>).</citation>
</ref>
<ref id="B31">
<label>31.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>O&#x2019;Regan</surname>
<given-names>K</given-names>
</name>
<name>
<surname>No&#xeb;</surname>
<given-names>A</given-names>
</name>
</person-group>. <article-title>A Sensorimotor Account of Vision and Visual Consciousness</article-title>. <source>Behav Brain Sci</source> (<year>2001</year>) <volume>24</volume>:<fpage>939</fpage>&#x2013;<lpage>73</lpage>. </citation>
</ref>
<ref id="B32">
<label>32.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Jost</surname>
<given-names>J</given-names>
</name>
</person-group>. <source>Leibniz und die moderne Naturwissenschaft</source>. <publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer</publisher-name> (<year>2019</year>).</citation>
</ref>
<ref id="B33">
<label>33.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Clark</surname>
<given-names>A</given-names>
</name>
</person-group>. <source>Surfing Uncertainty</source>. <publisher-loc>Oxford</publisher-loc>: <publisher-name>Oxford Univ. Press</publisher-name> (<year>2016</year>).</citation>
</ref>
<ref id="B34">
<label>34.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Klyubin</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Polani</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Nehaniv</surname>
<given-names>C</given-names>
</name>
</person-group>. <article-title>Empowerment: A Universal Agent-Centric Measure of Control</article-title>. In: <conf-name>2005 IEEE Congress on Evolutionary Computation</conf-name>; <conf-date>2005 2&#x2013;5 Sept</conf-date> (<year>2005</year>). p. <fpage>128</fpage>&#x2013;<lpage>35</lpage>. </citation>
</ref>
<ref id="B35">
<label>35.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Prinz</surname>
<given-names>W</given-names>
</name>
</person-group>. <source>Open Mind: The Social Making of agency and Intentionality</source>. <publisher-loc>Cambridge/Mass</publisher-loc>: <publisher-name>MIT Press</publisher-name> (<year>2012</year>).</citation>
</ref>
<ref id="B36">
<label>36.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vilimelis Aceituno</surname>
<given-names>P</given-names>
</name>
<name>
<surname>Ehsani</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Jost</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Spiking Time-dependent Plasticity Leads to Efficient Coding of Predictions</article-title>. <source>Biol Cybern</source> (<year>2020</year>) <volume>114</volume>:<fpage>43</fpage>&#x2013;<lpage>61</lpage>. <pub-id pub-id-type="doi">10.1007/s00422-019-00813-w</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>