<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comput. Neurosci.</journal-id>
<journal-title>Frontiers in Computational Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comput. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-5188</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fncom.2014.00010</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research Article</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>When do microcircuits produce beyond-pairwise correlations?</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Barreiro</surname> <given-names>Andrea K.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<xref ref-type="author-notes" rid="fn003"><sup>&#x02020;</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Gjorgjieva</surname> <given-names>Julijana</given-names></name>
<xref ref-type="aff" rid="aff2 "><sup>2</sup></xref>
<xref ref-type="author-notes" rid="fn003"><sup>&#x02020;</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Rieke</surname> <given-names>Fred</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Shea-Brown</surname> <given-names>Eric</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Department of Applied Mathematics, University of Washington</institution> <country>Seattle, WA, USA</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Applied Mathematics and Theoretical Physics, University of Cambridge</institution> <country>Cambridge, UK</country></aff>
<aff id="aff3"><sup>3</sup><institution>Department of Physiology and Biophysics, University of Washington</institution> <country>Seattle, WA, USA</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Robert Rosenbaum, University of Pittsburgh, USA</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Tim Gollisch, University Medical Center G&#x000F6;ttingen, Germany; Tatjana Tchumatchenko, Max Planck Institute for Brain Research, Germany</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Andrea K. Barreiro, Department of Mathematics, Southern Methodist University, 3200 Dyer Street, PO Box 750156, Dallas, TX 75275-0156, USA e-mail: <email>abarreiro&#x00040;smu.edu</email></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to the journal Frontiers in Computational Neuroscience.</p></fn>
<fn fn-type="present-address" id="fn003"><p>&#x02020;Present address: Andrea K. Barreiro, Department of Mathematics, Southern Methodist University, Dallas, USA; Julijana Gjorgjieva, Center for Brain Science, Harvard University, Cambridge, USA</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>06</day>
<month>02</month>
<year>2014</year>
</pub-date>
<pub-date pub-type="collection">
<year>2014</year>
</pub-date>
<volume>8</volume>
<elocation-id>10</elocation-id>
<history>
<date date-type="received">
<day>01</day>
<month>07</month>
<year>2013</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>01</month>
<year>2014</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2014 Barreiro, Gjorgjieva, Rieke and Shea-Brown.</copyright-statement>
<copyright-year>2014</copyright-year>
<license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/3.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract><p>Describing the collective activity of neural populations is a daunting task. Recent empirical studies in retina, however, suggest a vast simplification in how multi-neuron spiking occurs: the activity patterns of retinal ganglion cell (RGC) populations under some conditions are nearly completely captured by pairwise interactions among neurons. In other circumstances, higher-order statistics are required and appear to be shaped by input statistics and intrinsic circuit mechanisms. Here, we study the emergence of higher-order interactions in a model of the RGC circuit in which correlations are generated by common input. We quantify the impact of higher-order interactions by comparing the responses of mechanistic circuit models vs. &#x0201C;null&#x0201D; descriptions in which all higher-than-pairwise correlations have been accounted for by lower order statistics; these are known as pairwise maximum entropy (PME) models. We find that over a broad range of stimuli, output spiking patterns are surprisingly well captured by the pairwise model. To understand this finding, we study an analytically tractable simplification of the RGC model. We find that in the simplified model, bimodal input signals produce larger deviations from pairwise predictions than unimodal inputs. The characteristic light filtering properties of the upstream RGC circuitry suppress bimodality in light stimuli, thus removing a powerful source of higher-order interactions. This provides a novel explanation for the surprising empirical success of pairwise models.</p></abstract>
<kwd-group>
<kwd>retinal ganglion cells</kwd>
<kwd>maximum entropy distribution</kwd>
<kwd>stimulus-driven</kwd>
<kwd>correlations</kwd>
<kwd>computational model</kwd>
</kwd-group>
<counts>
<fig-count count="10"/>
<table-count count="2"/>
<equation-count count="52"/>
<ref-count count="59"/>
<page-count count="25"/>
<word-count count="19022"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="introduction" id="s1">
<title>1. Introduction</title>
<p>Information in neural circuits is often encoded in the activity of large, highly interconnected neural populations. The combinatoric explosion of possible responses of such circuits poses major conceptual, experimental, and computational challenges. How much of this potential complexity is realized? What do statistical regularities in population responses tell us about circuit architecture? Can simple circuit models with limited interactions among cells capture the relevant information content? These questions are central to our understanding of neural coding and decoding.</p>
<p>Two developments have advanced studies of synchronous activity in recent years. First, new experimental techniques provide access to responses from the large groups of neurons necessary to adequately sample synchronous activity patterns (Baudry and Taketani, <xref ref-type="bibr" rid="B7">2006</xref>). Second, maximum entropy approaches from statistical physics have provided a powerful approach to distinguish genuine higher-order synchrony (correlations) from that explainable by pairwise statistical interactions among neurons (Martignon et al., <xref ref-type="bibr" rid="B30">2000</xref>; Amari, <xref ref-type="bibr" rid="B1">2001</xref>; Schneidman et al., <xref ref-type="bibr" rid="B45">2003</xref>). These approaches have produced diverse findings. In some instances, activity of neural populations is extremely well described by pairwise interactions alone, so that pairwise maximum entropy (PME) models provide a nearly complete description (Shlens et al., <xref ref-type="bibr" rid="B49">2006</xref>, <xref ref-type="bibr" rid="B48">2009</xref>). In other cases, while pairwise models bring major improvements over independent descriptions, it is not clear that they fully capture the data (Martignon et al., <xref ref-type="bibr" rid="B30">2000</xref>; Schneidman et al., <xref ref-type="bibr" rid="B44">2006</xref>; Tang et al., <xref ref-type="bibr" rid="B50">2008</xref>; Yu et al., <xref ref-type="bibr" rid="B58">2008</xref>; Montani et al., <xref ref-type="bibr" rid="B32">2009</xref>; Ohiorhenuan et al., <xref ref-type="bibr" rid="B36">2010</xref>; Santos et al., <xref ref-type="bibr" rid="B43">2010</xref>). Empirical studies indicate that pairwise models can fail to explain the responses of spatially localized triplets of cells (Ohiorhenuan et al., <xref ref-type="bibr" rid="B36">2010</xref>; Ganmor et al., <xref ref-type="bibr" rid="B18">2011</xref>), as well as the activity of populations of &#x0007E;100 cells responding to natural stimuli (Ganmor et al., <xref ref-type="bibr" rid="B18">2011</xref>). Overall, the diversity of empirical results highlights the need to understand the network and input features that control the statistical complexity of synchronous activity patterns.</p>
<p>Several themes have emerged from efforts to link the correlation structure of spiking activity to circuit mechanisms using both abstract (Amari et al., <xref ref-type="bibr" rid="B2">2003</xref>; Krumin and Shoham, <xref ref-type="bibr" rid="B23">2009</xref>; Macke et al., <xref ref-type="bibr" rid="B26">2009</xref>; Roudi et al., <xref ref-type="bibr" rid="B40">2009a</xref>) and biologically-based models (Bohte et al., <xref ref-type="bibr" rid="B9">2000</xref>; Martignon et al., <xref ref-type="bibr" rid="B30">2000</xref>; Roudi et al., <xref ref-type="bibr" rid="B41">2009b</xref>); these models, however, do not provide a full description for why the PME models succeed or fail to capture neural circuit dynamics. First, thresholding non-linearities in circuits with Gaussian input signals can generate correlations that cannot be explained by pairwise statistics (Amari et al., <xref ref-type="bibr" rid="B2">2003</xref>); the deviations from pairwise predictions are modest at moderate population sizes (Macke et al., <xref ref-type="bibr" rid="B26">2009</xref>), but may become severe as population size grows large (Amari et al., <xref ref-type="bibr" rid="B2">2003</xref>; Macke et al., <xref ref-type="bibr" rid="B27">2011</xref>). The pairwise model also fails in networks of recurrent integrate-and-fire units with adapting thresholds and refractory potassium currents (Bohte et al., <xref ref-type="bibr" rid="B9">2000</xref>). The same is true for &#x0201C;Boltzmann-type&#x0201D; networks with hidden units (Koster et al., <xref ref-type="bibr" rid="B22">2013</xref>). Finally, small groups of model neurons that perform logical operations can be shown to generate higher-order interactions by introducing noisy processes with synergistic effects (Schneidman et al., <xref ref-type="bibr" rid="B45">2003</xref>), but it is unclear what neural mechanisms might produce similar distributions. These diverse findings point to the important role that circuit features and mechanisms&#x02014;input statistics, input/output relationships, and circuit connectivity&#x02014;can play in regulating higher-order interactions. Nevertheless, we lack a systematic understanding that links these features and their combinations to the success and failure of pairwise statistical models.</p>
<p>A second theme that has emerged is the use of perturbation approaches to explain why maximum entropy models with purely pairwise interactions capture circuit behavior in the limit in which the population firing rate is very low (i.e., the total number of firing events from all cells in the same small time window is small) (Cocco et al., <xref ref-type="bibr" rid="B12">2009</xref>; Roudi et al., <xref ref-type="bibr" rid="B40">2009a</xref>; Tkacik et al., <xref ref-type="bibr" rid="B54">2009</xref>). Also in this regime, higher-order interactions cannot be introduced as an artifact of under-sampling the network (Tkacik et al., <xref ref-type="bibr" rid="B54">2009</xref>), a concern at higher population firing rates. However, the low to moderate population firing rates observed in many studies permit <italic>a priori</italic> a fairly broad range in the quality of pairwise fits. What is left to explain then is why circuits operating outside the low population firing rate regime often produce fits consistent with the PME model.</p>
<p>We approach this issue here by systematically characterizing the ability of PME models to capture the responses of a class of circuit models with the following defining features. First, we consider relatively small circuits of 3&#x02013;16 cells, each with identical intrinsic dynamics (i.e., spike-generating mechanism and level of excitability). Second, we assume a particular structure for inputs across the circuit. Each neuron receives the same global input which, for example, represents stimuli in the receptive fields of all modeled cells. Neurons also receive an independent, Gaussian-like noise term. Third, the circuit has either no reciprocal coupling, or has all-to-all excitatory or gap junction coupling. We begin with circuit models fully constrained by measured properties of primate ON parasol ganglion networks, receiving full-field and checkerboard light inputs. We then explore a simple thresholding model for which we exhaustively search over the entire parameter space.</p>
<p>We identify general principles that describe higher-order spike correlations in the circuits we study. First, in all cases we examined, the overall strength of higher-order correlations are constrained to be far lower than the statistically possible limits. Second, for the higher-order correlations that do occur, the primary factor that determines how significant they will be is the bimodal vs. unimodal profile of the common input signal. A secondary factor is the strength of recurrent coupling, which has a non-monotonic impact on higher-order correlations. Our findings provide insight into why some previously measured activity patterns are well captured by PME descriptions, and provide predictions for the mechanisms that allow for higher-order spike correlations to emerge.</p>
</sec>
<sec sec-type="results" id="s2">
<title>2. Results</title>
<sec>
<title>2.1. Quantifying higher-order correlations in neural circuits</title>
<p>One strategy to identify higher-order interactions is to compare multi-neuron spike data against a description in which any higher-order interactions have been removed in a principled way&#x02014;that is, a description in which all higher-order correlations are completely described by lower-order statistics. Such a description may be given by a maximum entropy model (Jaynes, <xref ref-type="bibr" rid="B20">1957a</xref>,<xref ref-type="bibr" rid="B21">b</xref>; Amari, <xref ref-type="bibr" rid="B1">2001</xref>), in which one identifies the most unstructured, or maximum entropy, distribution consistent with the constraints. Comparing the predicted and measured probabilities of different responses tests whether the constraints used are sufficient to explain observed network activity, or whether additional constraints need to be considered. Such constraints would produce additional structure in the predicted response distribution, and hence lower the entropy.</p>
<p>A common approach is to limit the constraints to a given statistical order&#x02014;for example, to consider only the first and second moments of the distributions, which are determined by the mean and pairwise interactions. In the context of spiking neurons, we denote &#x003BC;<sub><italic>i</italic></sub> &#x02261; <bold>E</bold>[<italic>x</italic><sub><italic>i</italic></sub>] as the firing rate of neuron <italic>i</italic> and <inline-formula><mml:math id="M1"><mml:mover accent='true'><mml:mi>&#x003C1;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover></mml:math></inline-formula><sub><italic>ij</italic></sub> &#x02261; <bold>E</bold> [<italic>x</italic><sub><italic>i</italic></sub> <italic>x</italic><sub><italic>j</italic></sub>] as the joint probability that neurons <italic>i</italic> and <italic>j</italic> will fire. The distribution with the largest entropy for a given &#x003BC;<sub><italic>i</italic></sub> and <inline-formula><mml:math id="M2"><mml:mover accent='true'><mml:mi>&#x003C1;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover></mml:math></inline-formula><sub><italic>ij</italic></sub> is referred to as the <italic>PME</italic> model.</p>
<p>We use the Kullback&#x02013;Leibler divergence, <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M3"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>), to quantify the accuracy of the PME approximation <inline-formula><mml:math id="M4"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula> to a distribution <italic>P</italic>. This measure has a natural interpretation as the contribution of higher-order interactions to the response entropy <italic>S</italic>(<italic>P</italic>) (Amari, <xref ref-type="bibr" rid="B1">2001</xref>; Schneidman et al., <xref ref-type="bibr" rid="B45">2003</xref>), and may in this context be written as the difference of entropies <italic>S</italic>(<inline-formula><mml:math id="M5"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) &#x02212; <italic>S</italic>(<italic>P</italic>). In addition, <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M6"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) is approximately &#x02212;log<sub>2</sub><italic>L</italic>, where <italic>L</italic> is the average likelihood (over different observations) that a sequence of data drawn from the distribution <italic>P</italic> was instead drawn from the model <inline-formula><mml:math id="M7"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula> (Cover and Thomas, <xref ref-type="bibr" rid="B13">1991</xref>; Shlens et al., <xref ref-type="bibr" rid="B49">2006</xref>). For example, if <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M8"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) &#x0003D; 1, the average likelihood that a single sample, i.e., a single network response, came from <inline-formula><mml:math id="M9"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula> relative to the likelihood that it came from <italic>P</italic> is 2<sup>&#x02212;1</sup> (we use the base 2 logarithm in our definition of the Kullback&#x02013;Leibler divergence, so all numerical values are in units of bits).</p>
<p>An alternative measure of the quality of the pairwise model comes from normalizing <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M10"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) by the corresponding distance of the distribution <italic>P</italic> from an <italic>independent maximum entropy</italic> fit <italic>D</italic><sub>KL</sub>(<italic>P, P</italic><sub>1</sub>), where <italic>P</italic><sub>1</sub> is the highest entropy distribution consistent with the mean firing rates of the cells (equivalently, the product of single-cell marginal firing probabilities) (Amari, <xref ref-type="bibr" rid="B1">2001</xref>). Many studies (Schneidman et al., <xref ref-type="bibr" rid="B44">2006</xref>; Shlens et al., <xref ref-type="bibr" rid="B49">2006</xref>, <xref ref-type="bibr" rid="B48">2009</xref>; Roudi et al., <xref ref-type="bibr" rid="B40">2009a</xref>) use</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M11"><mml:mrow><mml:mi>&#x00394;</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mtext>KL</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mtext>KL</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>;</mml:mo></mml:mrow></mml:math></disp-formula>
<p>a value of &#x00394; &#x0003D; 1 indicates that the pairwise model perfectly captures the additional information left out of the independent model, while a value of &#x00394; &#x0003D; 0 indicates that the pairwise model gives no improvement over the independent model. To aid comparison with other studies, we report values of &#x00394; in parallel with <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M12"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) when appropriate.</p>
<p>We next explore and interpret the achievable range of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M13"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) values. The problem is made simpler if, following previous studies (Bohte et al., <xref ref-type="bibr" rid="B9">2000</xref>; Amari, <xref ref-type="bibr" rid="B1">2001</xref>; Macke et al., <xref ref-type="bibr" rid="B26">2009</xref>; Montani et al., <xref ref-type="bibr" rid="B32">2009</xref>), we consider only permutation-symmetric spiking patterns, in which the firing rate and correlation do not depend on the identity of the cells; i.e., &#x003BC;<sub><italic>i</italic></sub> &#x0003D; &#x003BC;, <inline-formula><mml:math id="M14"><mml:mover accent='true'><mml:mi>&#x003C1;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover></mml:math></inline-formula><sub><italic>ij</italic></sub> &#x0003D; <inline-formula><mml:math id="M15"><mml:mover accent='true'><mml:mi>&#x003C1;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover></mml:math></inline-formula> for <italic>i</italic> &#x02260; <italic>j</italic>. We start with three cells having binary responses and assume that the response is stationary and uncorrelated in time. From symmetry, the possible network responses are</p>
<disp-formula id="E2"><mml:math id="M16"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>p</italic><sub><italic>i</italic></sub> denotes the probability that a particular set of <italic>i</italic> cells spike and the remaining 3 &#x02212; <italic>i</italic> do not. Possible values of (<italic>p</italic><sub>0</sub>, <italic>p</italic><sub>1</sub>, <italic>p</italic><sub>2</sub>, <italic>p</italic><sub>3</sub>) are constrained by the fact that <italic>P</italic> is a probability distribution, so that the sum of <italic>p</italic><sub><italic>i</italic></sub> over all eight states is one.</p>
<p>To assess the numerical significance of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M17"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>), we can compare it with the maximal achievable value for any symmetric distribution on three spiking cells. For three cells, the maximal value is <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M18"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) &#x0003D; 1 (or 1/3 bits per neuron), achieved by the XOR operation (Schneidman et al., <xref ref-type="bibr" rid="B45">2003</xref>). This distribution is illustrated in Figure <xref ref-type="fig" rid="F1">1A</xref> (right), together with two distributions produced by our mechanistic circuit models&#x02014;illustrating observed deviations from PME fits for unimodal (left) and bimodal (middle) distributions of inputs (see below). The <italic>KL</italic>-divergence for these two patterns is 0.0013 and 0.091, respectively. As suggested by these bar plots (and explored in detail below), the distributions produced by a wide set of mechanistic circuit models are quite well captured by the PME approximation: to use the likelihood interpretation described above, an observer would need to draw many more samples from these distributions in order to distinguish between the true and model distributions: &#x02248;1000 times and &#x02248;10 times, respectively, in comparison to the XOR operator.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>A survey of the quality of the pairwise maximum entropy (PME) model for symmetric spiking distributions on three cells</bold>. <bold>(A)</bold> Probability distribution <italic>P</italic> (dark blue) and pairwise approximation <inline-formula><mml:math id="M19"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula> (light pink) for three example distributions. From left to right: an example from the simple sum-and-threshold model receiving skewed common input; an example from the sum-and-threshold model receiving bimodal common input [specifically, the distribution with maximal <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M20"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>)]; a specific probability distribution resulting from application of the XOR operator [for illustration of a &#x0201C;worst case&#x0201D; fit of the PME model (Schneidman et al., <xref ref-type="bibr" rid="B45">2003</xref>)]. <bold>(B)</bold> <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M21"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) vs. firing rate and &#x00394; vs. firing rate, for a comprehensive survey of possible symmetric spiking distributions on three cells (see text for details). Firing rate is defined as the probability of a spike occurring per cell per random draw of the sum-and-threshold model, as defined in Equation (16). Color indicates output correlation coefficient &#x003C1; ranging from black for &#x003C1; &#x02208; (0, 0.1), to white for &#x003C1; &#x02208; (0.9, 1), as illustrated in the color bars.</p></caption>
<graphic xlink:href="fncom-08-00010-g0001.tif"/>
</fig>
<p>To further identify appropriate &#x0201C;benchmark&#x0201D; values of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M22"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) with which to compare our mechanistic circuit models, in Figure <xref ref-type="fig" rid="F1">1B</xref> we show plots of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M23"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) and &#x00394; vs. firing rate produced by an exhaustive sampling of symmetric distributions on three cells. From this picture, we can see that it is possible to find symmetric, three-cell spiking distributions that are poorly fit by the pairwise model at a range of firing rates and pairwise correlations, with the largest values of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M24"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) found at low correlations (note that the XOR distribution has an average pairwise covariance of zero (i.e., <bold>E</bold>[<italic>X</italic><sub>1</sub> <italic>X</italic><sub>2</sub>] &#x0003D; <bold>E</bold>[<italic>X</italic><sub>1</sub>] <bold>E</bold>[<italic>X</italic><sub>2</sub>])).</p>
<sec>
<title>2.1.1. A condition for higher-order correlations</title>
<p>Possible solutions to the symmetric PME problem take the form of exponential functions characterized by two parameters, &#x003BB;<sub>1</sub> and &#x003BB;<sub>2</sub>, which serve as Lagrange multipliers for the constraints:</p>
<disp-formula id="E3"><label>(2)</label><mml:math id="M25"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>Z</mml:mi></mml:mfrac><mml:mi>exp</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>&#x003BB;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo></mml:mrow> </mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mrow><mml:mrow><mml:msub><mml:mi>&#x003BB;</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The factor <italic>Z</italic> normalizes <italic>P</italic> to be a probability distribution.</p>
<p>By combining individual probabilities of events as given by Equation (2) the following relationship must be satisfied by any symmetric PME solution:</p>
<disp-formula id="E4"><label>(3)</label><mml:math id="M26"><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>3</mml:mn></mml:msup><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
<p>This is equivalent to the condition that the <italic>strain</italic> measure of Ohiorhenuan and Victor (<xref ref-type="bibr" rid="B37">2010</xref>) be zero (in particular, the strain is negative whenever <italic>p</italic><sub>3</sub>/<italic>p</italic><sub>0</sub> &#x02212; (<italic>p</italic><sub>2</sub>/<italic>p</italic><sub>1</sub>)<sup>3</sup> &#x0003C; 0, a condition identified in Ohiorhenuan and Victor (<xref ref-type="bibr" rid="B37">2010</xref>) as corresponding to sparsity in the neural code).</p>
<p>For three-cell, symmetric networks, models that exactly satisfy Equation (3) will also be exactly described via PME. Moreover, note that probability models that meet this constraint fall on a surface in the space of (normalized) histograms, given by the probabilities <italic>p</italic><sub><italic>j</italic></sub>. One can verify by straightforward calculations (see Appendix) that&#x02014;given fixed lower order moments&#x02014;<italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M27"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) is a convex function of the probabilities <italic>p</italic><sub><italic>j</italic></sub>. This has interesting consequences for predicting when large vs. small values of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M28"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) will be found (see Appendix).</p>
<p>It is not necessary to assume permutation symmetry when deriving the PME fit <inline-formula><mml:math id="M29"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula> to an observed distribution <italic>P</italic>, or in computing derived quantities such as <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M30"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>), and we do not do so in this study. However, most of the distributions we study are derived from mechanistic models that are themselves symmetric or near-symmetric. Therefore, we anticipate that the simplified calculations for permutation-symmetric distributions will yield analytical insight into our findings.</p>
</sec>
</sec>
<sec>
<title>2.2. Mechanisms that impact beyond-pairwise correlations in triplets of on-parasol retinal ganglion cells</title>
<p>Having established the range of beyond-pairwise correlations that are possible statistically, we turn our focus to coding in retinal ganglion cell (RGC) populations, an area that has received a great deal of attention empirically. Specifically, PME approaches have been effective in capturing the activity of small RGC populations (Schneidman et al., <xref ref-type="bibr" rid="B44">2006</xref>; Shlens et al., <xref ref-type="bibr" rid="B49">2006</xref>, <xref ref-type="bibr" rid="B48">2009</xref>). This success does not have an obvious anatomical correlate; there are multiple opportunities in the retinal circuitry for interactions among three or more ganglion cells. We explored circuits composed of three RGC cells with input statistics, recurrent connectivity and spike-generating mechanisms based directly on experiment. We based our model on ON parasol RGCs, one of the RGC types for which PME approaches have been applied extensively (Shlens et al., <xref ref-type="bibr" rid="B49">2006</xref>, <xref ref-type="bibr" rid="B48">2009</xref>). In addition, by examining how marginal input statistics are shaped by stimulus filtering, we also reveal the role that the specific filtering properties of ON parasol cells have in shaping higher-order interactions.</p>
<sec>
<title>2.2.1. RGC model</title>
<p>We modeled a single ON parasol RGC in two stages (for details see section 4). First, we characterized the light-dependent excitatory and inhibitory synaptic inputs to cell <italic>k</italic> (<italic>g</italic><sup>exc</sup><sub><italic>k</italic></sub>(<italic>t</italic>), <italic>g</italic><sup>inh</sup><sub><italic>k</italic></sub>(<italic>t</italic>)) in response to randomly fluctuating light inputs <italic>s</italic><sub><italic>k</italic></sub>(<italic>t</italic>) via a linear-nonlinear model, e.g.,:</p>
<disp-formula id="E5"><label>(4)</label><mml:math id="M31"><mml:mrow><mml:msubsup><mml:mi>g</mml:mi><mml:mi>k</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msubsup><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mi>N</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msup><mml:mi>L</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msup><mml:mo>&#x02217;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:msubsup><mml:mi>&#x003B7;</mml:mi><mml:mi>k</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msubsup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <italic>N</italic><sup>exc</sup> is a static non-linearity, <italic>L</italic><sup>exc</sup> is a linear filter, and &#x003B7;<sup>exc</sup><sub><italic>k</italic></sub> is an effective input noise that captures variability in the response to repetitions of the same time-varying stimulus. These parameters were determined from fits to experimental data collected under conditions similar to those in which PME models have been tested empirically (Shlens et al., <xref ref-type="bibr" rid="B49">2006</xref>, <xref ref-type="bibr" rid="B48">2009</xref>; Trong and Rieke, <xref ref-type="bibr" rid="B55">2008</xref>). The modeled excitatory and inhibitory conductances captured many of the statistical features of the real conductances, particularly the correlation time and skewness (data not shown).</p>
<p>Second, we used Equation (4) and an equivalent expression for <italic>g</italic><sup>inh</sup><sub><italic>k</italic></sub>(<italic>t</italic>) as inputs to an integrate-and-fire model incorporating a non-linear voltage and history-dependent term to account for refractory interactions between spikes (Badel et al., <xref ref-type="bibr" rid="B4">2007</xref>, <xref ref-type="bibr" rid="B3">2008</xref>). The voltage evolution equation was of the form</p>
<disp-formula id="E6"><label>(5)</label><mml:math id="M32"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>d</mml:mi><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>V</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mtext>last</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>input</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mi>C</mml:mi></mml:mfrac><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <italic>F</italic>(<italic>V</italic>, <italic>t</italic> &#x02212; <italic>t</italic><sub>last</sub>) was allowed to depend on the time of the last spike <italic>t</italic><sub>last</sub>. Briefly, we obtained data from a dynamic clamp experiment (Sharpe et al., <xref ref-type="bibr" rid="B46">1993</xref>; Murphy and Rieke, <xref ref-type="bibr" rid="B34">2006</xref>) in which currents corresponding to <italic>g</italic><sup>exc</sup>(<italic>t</italic>) and <italic>g</italic><sup>inh</sup>(<italic>t</italic>) were injected into a cell and the resulting voltage response measured. The input current <italic>I</italic><sub>input</sub> injected during one time step was determined by scaling the excitatory and inhibitory conductances by driving forces based on the measured voltage in the previous time step; that is,</p>
<disp-formula id="E7"><label>(6)</label><mml:math id="M33"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>input</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mi>g</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msup><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>V</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>E</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mi>g</mml:mi><mml:mrow><mml:mtext>inh</mml:mtext></mml:mrow></mml:msup><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>V</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>I</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>We used this data to determine <italic>F</italic> and <italic>C</italic> using the procedure described in Badel et al. (<xref ref-type="bibr" rid="B4">2007</xref>); details, including values of all fitted parameters, are described in section 4. Recurrent connections were implemented by adding an input current proportional to the voltage difference between the two coupled cells.</p>
<p>The prescription above provided a flexible model that allowed us to study the responses of three-cell RGC networks to a wide range of light inputs and circuit connectivities. Specifically, we simulated RGC responses to light stimuli that were (1) constant, (2) time-varying and spatially uniform, and (3) varying in both space and time. Correlations between cell inputs arose from shared stimuli, from shared noise originating in the retinal circuitry (Trong and Rieke, <xref ref-type="bibr" rid="B55">2008</xref>), or from recurrent connections (Dacey and Brace, <xref ref-type="bibr" rid="B14">1992</xref>; Trong and Rieke, <xref ref-type="bibr" rid="B55">2008</xref>). Shared stimuli were described by correlations among the light inputs <italic>s</italic><sub><italic>k</italic></sub>. Shared noise arose via correlations in &#x003B7;<sup>exc</sup><sub><italic>k</italic></sub> and &#x003B7;<sup>ink</sup><sub><italic>k</italic></sub> as described in section 4. The recurrent connections were chosen to be consistent with observed gap-junctional coupling between ON parasol cells. We also investigated how stimulus filtering by <italic>L</italic><sup>exc</sup> and <italic>L</italic><sup>inh</sup> influenced network statistics. To compare our results with empirical studies, constant light, and spatially and temporally fluctuating checkerboard stimuli were used as in Shlens et al. (<xref ref-type="bibr" rid="B49">2006</xref>, <xref ref-type="bibr" rid="B48">2009</xref>).</p>
</sec>
<sec>
<title>2.2.2. The feedforward RGC circuit is well-described by the PME model for full-field light stimuli</title>
<p>We start by considering networks without recurrent connectivity and with constant, full-field (i.e., spatially uniform) light stimuli. Thus, we set <italic>s</italic><sub><italic>k</italic></sub>(<italic>t</italic>) &#x0003D; 0 for <italic>k</italic> &#x0003D; 1, 2, 3, so that the cells received only Gaussian correlated noise &#x003B7;<sup>exc</sup><sub><italic>k</italic></sub> and &#x003B7;<sup>inh</sup><sub><italic>k</italic></sub> and constant excitatory and inhibitory conductances. Time-dependent conductances were generated and used as inputs to a simulation of three model RGCs. Simulation length was sufficient to ensure significance of all reported deviations from PME fits (see section 4). We found that the spiking distributions were strikingly well-modeled by a PME fit, as shown in the righthand panel of Figure <xref ref-type="fig" rid="F2">2A</xref>; <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M34"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) is 2.90 &#x000D7; 10<sup>&#x02212;5</sup> bits. This result is consistent with the very good fits found experimentally in Shlens et al. (<xref ref-type="bibr" rid="B49">2006</xref>) under constant light stimulation.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><bold>Results for RGC simulations with constant light and full-field flicker</bold>. <bold>(A&#x02013;C)</bold> (Left) A histogram and time series of stimulus, (center) a histogram of excitatory conductances and (right) the resulting distribution of spiking patterns. Stimuli are shown as deviations from a baseline intensity, expressed as a fraction of the baseline. Right panels show the probability distribution on spiking patterns <italic>P</italic> obtained from simulation (&#x0201C;Observed&#x0201D;; dark blue), and the corresponding pairwise approximation <inline-formula><mml:math id="M35"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula> (&#x0201C;PME&#x0201D;; light pink). Each row gives these results for a different stimulus condition. <bold>(A)</bold> No stimulus (Gaussian noise only). <bold>(B)</bold> Gaussian input, standard deviation 1/6, refresh rate 8 ms. <bold>(C)</bold> Binary input, standard deviation 1/3, refresh rate 8 ms. <bold>(D)</bold> Binary input, standard deviation 1/3, refresh rate 100 ms. For panel <bold>(D)</bold>, the data in the left panel differs. (Left, top panel) The excitatory filter <italic>L</italic><sup>exc</sup>(<italic>t</italic>) (Equation 7) is shown instead of a stimulus histogram; (Left, bottom panel) the normalized excitatory conductance, as a function of time (red dashed line), is superimposed on the stimulus (blue solid). (Center) The histogram of excitatory conductances and (right) the resulting distribution of spiking patterns. Both the form of the filter and the conductance trace illustrate that the LN model that processes light input acts as a (time-shifted) high pass filter.</p></caption>
<graphic xlink:href="fncom-08-00010-g0002.tif"/>
</fig>
<p>Next, we introduce temporal modulation into the full-field light stimuli such that each cell received the same stimulus, <italic>s</italic><sub><italic>k</italic></sub>(<italic>t</italic>) &#x0003D; <italic>s</italic>(<italic>t</italic>), where <italic>s</italic>(<italic>t</italic>) refreshed every few milliseconds with an independently chosen value from one of several marginal distributions. For our initial set of experiments, the marginal distribution was either Gaussian (as in Ganmor et al., <xref ref-type="bibr" rid="B18">2011</xref>) or binary (as used in Shlens et al., <xref ref-type="bibr" rid="B49">2006</xref>). For both choices, we explored inputs with a range of standard deviations (1/16, 1/12, 1/8, 1/6, 1/4, 1/3, or 1/2 of a baseline light intensity) and refresh rates (8, 40, or 100 ms). The shared stimulus produced strong pairwise correlation between conductances of neighboring cells. However, values of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M36"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) remained small, under 10<sup>&#x02212;2</sup> bits in all conditions tested.</p>
</sec>
<sec>
<title>2.2.3. Impact of stimulus spatial scale</title>
<p>We next asked whether PME models capture RGC responses to stimuli with varying spatial scales. We fixed stimulus dynamics to match the two cases that yielded the highest <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M37"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) under the full-field protocol: for both Gaussian and binary stimuli, we used 8 ms refresh rate and &#x003C3; &#x0003D; 1/2. The stimulus was generated as a random checkerboard with squares of variable size; each square in the checkerboard, or <italic>stixel</italic>, was drawn independently from the appropriate marginal distribution and updated at the corresponding refresh rate. The conductance input to each RGC was then given by convolving the light stimulus with its receptive field, where the stimulus was positioned with a fixed rotation and translation relative to the receptive fields. This position was drawn randomly at the beginning of each simulation and held constant throughout (see insets of Figures <xref ref-type="fig" rid="F3">3B,C</xref> for examples, and section 4 for further details).</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p><bold>Results for RGC simulations with light stimuli of varying spatial scale (&#x0201C;stixels&#x0201D;)</bold>. <bold>(A)</bold> Average <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M38"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) as a function of stixel size. Values were averaged over five stimulus positions, each with a different (random) stimulus rotation and translation; 512 &#x003BC;m corresponds to full-field stimuli. For the rest of the panels, data from the binary light distributions is shown; results from the Gaussian case are similar. <bold>(B,C)</bold> Probability of singlet and doublet spiking events, under stimulation by movies of 256 &#x003BC;m <bold>(B)</bold> and 60 &#x003BC;m <bold>(C)</bold> stixels. Event probabilities are plotted in 3-space, with the <italic>x</italic>, <italic>y</italic>, and <italic>z</italic> axes identifying the singlet (doublet) events 001 (011), 010 (101), and 100 (110), respectively. The black dashed line indicates perfect cell-to-cell homogeneity (e.g., <italic>P</italic>[(1, 0, 0)] &#x0003D; <italic>P</italic>[(0, 1, 0)] &#x0003D; <italic>P</italic>[(0, 0, 1)]). Both individual runs (dots) and averages over 20 runs (large circles) are shown, with averages outlined in black (singlet) and gray (doublet). Different colors indicate different stimulus positions. Insets: contour lines of the three receptive fields (at the 1 and 2 SD contour lines for the receptive field center; and at the zero contour line) superimposed on the stimulus checkerboard (for illustration, pictured in an alternating black/white pattern).</p></caption>
<graphic xlink:href="fncom-08-00010-g0003.tif"/>
</fig>
<p>The RGC spike patterns remained very well described by PME models for the full range of spatial scales. Figure <xref ref-type="fig" rid="F3">3A</xref> shows this by plotting <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M39"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) vs. stixel size. Values of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M40"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) increased with spatial scale, sharply rising beyond 128 &#x003BC;m, where a stixel had approximately the same size as a receptive field center, illustrating that introducing spatial scale via stixels produces even closer fits by PME models (the points at 512 &#x003BC;m correspond to the full-field simulations).</p>
<p>Values reported in Figure <xref ref-type="fig" rid="F3">3A</xref> are <italic>averages</italic> of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M41"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) produced by five random stimulus positions. At stixel sizes of 128 &#x003BC;m and 256 &#x003BC;m, the resulting spiking distributions differed significantly from position to position; in Figure <xref ref-type="fig" rid="F3">3B</xref>, we show the probabilities of the distinct singlet [e.g., <italic>P</italic>(1, 0, 0)] and doublet [e.g., <italic>P</italic>(1, 1, 0)] spiking events produced at 256 &#x003BC;m. Each stimulus position created a &#x0201C;cloud&#x0201D; of dots (identified by color); large dots show the average over 20 sub-simulations. Each sub-simulation was identified by a small dot of the same color; because the simulations were very well-resolved, most of them were contained within the large dots (and hence not visible in the figure). Heterogeneity across stimulus positioning is indicated by the distinct positioning of differently colored dots. At smaller spatial scales, the process of averaging stimuli over the receptive fields resulted in spiking distributions that were largely unchanged with stimulus position, as shown in Figure <xref ref-type="fig" rid="F3">3C</xref>, where singlet and doublet spiking probabilities are plotted for 60 &#x003BC;m stixels. Thus, filtered light inputs were largely homogeneous from cell to cell, as each receptive field sampled a similar number of independent, statistically identical inputs; the inset of Figure <xref ref-type="fig" rid="F3">3C</xref> shows the projection of input stixels onto cell receptive fields from an example with 60 &#x003BC;m stixels. The resulting excitatory conductances and spiking patterns were very close to cell-symmetric (see Figures <xref ref-type="supplementary-material" rid="SM2">S2B,C</xref>).</p>
<p>By contrast, spiking patterns showed significant heterogeneity from cell to cell when the stixel size was large, as illustrated in Figure <xref ref-type="fig" rid="F3">3B</xref>. This arises because each cell in the population may be located differently with respect to stixel boundaries, and therefore receive a distinct pattern of input activity; this is illustrated by the inset of Figure <xref ref-type="fig" rid="F3">3B</xref>, which shows the projection of input stixels onto cell receptive fields from one such simulation. However, PME models gave excellent fits to data regardless of heterogeneity in RGC responses (see Figures <xref ref-type="supplementary-material" rid="SM2">S2E,F</xref>); as seen in Figure <xref ref-type="fig" rid="F3">3A</xref>, over all 20 sub-simulations, and over all individual stixel positions, we found a maximal <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M42"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) value of 0.00811.</p>
</sec>
<sec>
<title>2.2.4. Conductance profiles and impact of stimulus filtering</title>
<p>Intrigued by the consistent finding of low values of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M43"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) from the RGC model circuit despite stimulation by a wide variety of highly correlated stimulus classes, we sought to further characterize the processing of light stimuli by this circuit. In particular, we examined the effects of different marginal statistics of light stimuli, standard deviation of full-field flicker, and refresh rate on the marginal distributions of excitatory conductances. We focused on excitatory conductances because they exhibit stronger correlations than inhibitory conductances in ON parasol RGCs (Trong and Rieke, <xref ref-type="bibr" rid="B55">2008</xref>).</p>
<p>With constant light stimulation (no temporal modulation) the excitatory conductances were unimodal and broadly Gaussian (Figure <xref ref-type="fig" rid="F2">2A</xref>, middle panel). For a short refresh rate (8 ms) or small flicker size (standard deviation 1/6 or 1/4 of baseline light intensity), temporal averaging via the filter <italic>L</italic><sup>exc</sup> and the approximately linear form of <italic>N</italic><sup>exc</sup> over these light intensities produced a unimodal, modestly skewed distribution of excitatory conductances, regardless of whether the flicker was drawn from a Gaussian or binary distribution (see Figures <xref ref-type="fig" rid="F2">2B,C</xref>, center panels). For a slower refresh rate (100 ms) and large flicker size (s.d. 1/3 or 1/2 of baseline light intensity), excitatory conductances had multi-modal and skewed features, again regardless of whether the flicker was drawn from a Gaussian or binary distribution (Figure <xref ref-type="fig" rid="F2">2D</xref>). Other parameters being equal, binary light input produced more skewed conductances. While some conductance distributions had multiple local maxima, these were never well separated, with the envelope of the distribution still resembling a skewed distribution.</p>
<p>The mechanism that leads to unimodal distributions of conductances, even when light stimuli are binary, is high-pass filtering&#x02014;a consequence of the differentiating linear filter in Equation (7) and illustrated in Figure <xref ref-type="fig" rid="F2">2D</xref>. To demonstrate this, we constructed an alternative filter with a more monophasic shape [Equation (9), illustrated in Figure <xref ref-type="supplementary-material" rid="SM1">S1</xref>] and compared the excitatory conductance distributions side-by-side. We saw a striking difference in the response to long time scale, binary stimuli: the distributions produced by the monophasic filter reflected the bimodal shape of the input. Interestingly, the resulting simulation produced eight-times greater <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M44"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) (Figure <xref ref-type="fig" rid="F4">4</xref>). This suggests that greater <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M45"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) may occur when ganglion cell inputs are primarily characterized via monophasic filters, e.g., at low mean light levels for which the retinal circuit acts to primarily integrate, rather than differentiate over time.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p><bold>Comparison of RGC simulations computed with the original ON parasol filter, vs. simulations using a more monophasic filter</bold>. <bold>(A)</bold> <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M46"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) for original vs. monophasic filter. Data is organized by stimulus refresh rate (8, 40, and 100 ms) and marginal statistics (Gaussian vs. binary). <bold>(B)</bold> Histograms of excitatory conductances for an illustrative stimulus class, under original (top) and monophasic (bottom) filters. The marginal statistics and refresh rate are illustrated by icons inside black circles; here, binary stimuli with refresh rate 100 ms. The input standard deviation (expressed as a fraction of baseline light intensity) was 1/2. <bold>(C)</bold> Time course of stimulus and resulting excitatory conductances, from simulation shown in <bold>(B)</bold>: original (top) vs. monophasic (bottom) filters.</p></caption>
<graphic xlink:href="fncom-08-00010-g0004.tif"/>
</fig>
<p>In Figure <xref ref-type="fig" rid="F4">4A</xref>, we examine this effect over all full-field stimulus conditions by plotting <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M47"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) from simulations with the monophasic filter, against <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M48"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) from simulations in which the original filter was used with the same stimulus type. An increase in <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M49"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) was observed across stimulus conditions, with a markedly larger effect for longer refresh rates. This consistent change could not be attributed to changes in lower order statistics; there was no consistent relationship between the change in pairwise model performance and either firing rate or pairwise correlations (data not shown). Instead, large effects in <italic>D</italic><sub>KL</sub> were accompanied by a striking increase in the bi- or multi-modality of excitatory conductances (see Figure <xref ref-type="fig" rid="F4">4B</xref>). In Figure <xref ref-type="fig" rid="F4">4C</xref>, we show an example stimulus and excitatory current trace taken from the simulation shown in Figure <xref ref-type="fig" rid="F4">4B</xref>: the monophasic filter allows the excitatory synaptic currents to track a long-timescale, bimodal stimulus with higher fidelity, transferring the bimodality of the stimulus into the synaptic currents. This finding was robust to specifics of the filtering process; we were able to reproduce the same results by designing integrating filters in different ways (data not shown).</p>
</sec>
<sec>
<title>2.2.5. Recurrent connectivity in the RGC circuit</title>
<p>We next considered the role of recurrence in shaping higher-order interactions by incorporating gap junction coupling into our simulations. We did this separately for each full-field stimulus condition described earlier. In each case, we added gap junction coupling with strengths from 1 to 16 times an experimentally measured value (Trong and Rieke, <xref ref-type="bibr" rid="B55">2008</xref>), and compared the resulting <italic>D</italic><sub>KL</sub> with that obtained without recurrent coupling (Figure <xref ref-type="fig" rid="F5">5</xref>).</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p><bold>The impact of recurrent coupling on RGC networks with full-field visual stimuli</bold>. The strength of gap junction connections was varied from a baseline level (relative magnitude <italic>g</italic> &#x0003D; 1, or absolute magnitude <italic>g</italic><sup>gap</sup> &#x0003D; 1.1 nS) to an order of magnitude larger (<italic>g</italic> &#x0003D; 16, or <italic>g</italic><sup>gap</sup> &#x0003D; 17.6 nS). In each panel, <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M50"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) obtained with coupling is plotted vs. the value obtained for the same stimulus ensemble without coupling, for each of 42 different stimulus ensembles. <bold>(A)</bold> <italic>g</italic><sup>gap</sup> &#x0003D; 1.1 nS (experimentally observed value); <bold>(B)</bold> <italic>g</italic><sup>gap</sup> &#x0003D; 4.4 nS; <bold>(C)</bold> <italic>g</italic><sup>gap</sup> &#x0003D; 8.8 nS; <bold>(D)</bold> <italic>g</italic><sup>gap</sup> &#x0003D; 17.6 nS.</p></caption>
<graphic xlink:href="fncom-08-00010-g0005.tif"/>
</fig>
<p>At the experimentally measured coupling strength (<italic>g</italic><sup>gap</sup> &#x0003D; 1.1 nS) itself, the fit of the pairwise model barely changed (Figure <xref ref-type="fig" rid="F5">5A</xref>) from the model without coupling. At twice the measured coupling strength (<italic>g</italic><sup>gap</sup> &#x0003D; 2.2 nS), recurrent coupling had increased higher-order interactions, as measured by larger values of <italic>D</italic><sub>KL</sub> for all tested stimulus conditions. Higher order interactions could be further increased, particularly for long refresh rates (100 ms), by increasing the coupling strength to four or eight times its baseline level (<italic>g</italic><sup>gap</sup> &#x0003D; 4.4 nS or <italic>g</italic><sup>gap</sup> &#x0003D; 8.8 nS; see Figures <xref ref-type="fig" rid="F5">5B,C</xref>). Consistent with the intuition that very strong coupling leads to &#x0201C;all-or-none&#x0201D; spiking patterns, <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M51"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) decreased as <italic>g</italic><sup>gap</sup> increased further, often to a level below what was seen in the absence of coupling (Figure <xref ref-type="fig" rid="F5">5D</xref>). In summary, the impact of coupling on <italic>D</italic><sub>KL</sub> is maximized at intermediate values of the coupling strength. However, the impact of recurrent coupling on the maximal values of <italic>D</italic><sub>KL</sub> evoked by visual stimuli is small overall, and almost negligible for experimentally measured coupling strengths.</p>
</sec>
<sec>
<title>2.2.6. Modeling heavy-tailed light stimuli in the RGC circuit</title>
<p>Finally, we repeated the full-field, recurrent, and alternate filter simulations previously described with light stimuli drawn from either Cauchy or heavy-tailed distributions: such distributions have been found to model the frequency of occurrence of luminance values in photographs of natural scenes (Ruderman and Bialek, <xref ref-type="bibr" rid="B42">1994</xref>). In contrast to previous results with Gaussian and bimodal inputs, here we found very low <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M52"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) over all stimulus conditions: the largest values found were more than an order of magnitude smaller than those obtained earlier. Specifically, for all conditions, we found <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M53"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) &#x0003C; 4.5 &#x000D7; 10<sup>&#x02212;4</sup>, over all 42 network realizations; for many simulations, this number did not meet a threshold for statistical significance (see section 4.1.7), indicating that <italic>P</italic> and <inline-formula><mml:math id="M54"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula> were not statistically distinguishable. Using a more monophasic filter resulted in no apparent consistent change to <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M55"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>). When gap junction coupling was added, <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M56"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) was maximized at an intermediate value; when <italic>g</italic><sup>gap</sup> &#x0003D; 8.8, all simulations produced a statistically significant <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M57"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) &#x02248; 3 &#x02212; 4 &#x000D7; 10<sup>&#x02212;3</sup>. However, overall levels remained relatively low, roughly 1/2 the value achieved with Gaussian or binary stimuli.</p>
<p>To explain these findings, we examined the excitatory input currents: we found that over a broad range of refresh rates and stimulus variances, the marginal distributions of excitatory input conductances produced were remarkably unimodal in shape, and showed little skewness (Figure <xref ref-type="fig" rid="F6">6A</xref>). By examining the time evolution of the filtered stimuli (see Figure <xref ref-type="fig" rid="F6">6B</xref>), we see that heavy-tailed distributions allow rare, large events, but at the expense of medium-size events which explore the full range of the linear-nonlinear model used for stimulus processing (compare the blue with the red/green traces). When combined with the Gaussian background noise, this produces near-Gaussian excitatory conductances and, as may be expected from our original full-field simulations, very low <italic>D</italic><sub>KL</sub>.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p><bold>Results for RGC simulations with heavy-tailed inputs</bold>. <bold>(A)</bold> Histograms of excitatory conductances, for the original (left) vs. monophasic (right) filter. The marginal statistics are heavy-tailed skew (top) and Cauchy (bottom) inputs, and refresh rate is 40 ms for both panels. The input standard deviation (expressed as a fraction of baseline light intensity) was 1/2 for both simulations. <bold>(B)</bold> Sample 100 ms stimuli, filtered by the original linear filter <italic>L</italic><sub>exc</sub> (top) and altered, monophasic filter <italic>L</italic><sub>exc,M</sub>(bottom). Cauchy (blue solid), Gaussian (red dashed), and bimodal (green dash-dotted) stimuli are shown.</p></caption>
<graphic xlink:href="fncom-08-00010-g0006.tif"/>
</fig>
<p>We hypothesize that the methodology of averaging over the entire stimulus ensemble may not capture the significance of rare events that may individually be detected with high fidelity: <italic>D</italic><sub>KL</sub> was low even for full-field, high variance stimuli, which presumably caused (infrequent) global spiking events. Additionally, an important avenue for future work would be to test the ability of our RGC model, which was trained on Gaussian stimuli, to accurately model the response of a ganglion cell to stimuli whose variance is dominated by large events. Recent work examining the adaptation of retinal filtering properties to higher-order input statistics found little evidence of adaptation; however, the stimuli used in this work incorporated significant kurtosis but not heavy tails (Tkacik et al., <xref ref-type="bibr" rid="B53">2012</xref>).</p>
</sec>
<sec>
<title>2.2.7. Summary of findings for RGC circuit</title>
<p>In summary, we probed the spiking response of a small array of RGC models to changes in light stimuli, gap junction coupling, and stimulus filtering properties, and identified two circumstances in which higher-order interactions were robustly generated in the spiking response. First, higher-order interactions were generated when excitatory currents had bimodal structure; we observed such structure when bimodal light stimuli was processed by a relatively monophasic filter. Secondly, higher-order interactions were maximized at an intermediate value of gap junction coupling; this value was, however, much larger (eight times) than the experimentally observed coupling strength.</p>
</sec>
</sec>
<sec>
<title>2.3. A simplified circuit that explains trends in RGC cell model</title>
<sec>
<title>2.3.1. Setup and motivation</title>
<p>In the previous section, we developed results for a computational model tuned to a very specific cell type; we now ask whether these findings will hold for a more general class of neural circuits, or whether they are the consequence of system-specific features. To answer this question, we considered a simplified model of neural spiking: a feedforward circuit in which three spiking cells sum their inputs and spike according to whether or not they cross a threshold. Such highly idealized models of spiking have a long history in neuroscience (McCulloch and Pitts, <xref ref-type="bibr" rid="B31">1943</xref>) and have been recently shown to predict the pairwise and higher-order activity of neural groups in both neural recordings and more complex dynamical spiking models (Nowotny and Huerta, <xref ref-type="bibr" rid="B35">2003</xref>; Tchumatchenko et al., <xref ref-type="bibr" rid="B52">2010</xref>; Yu et al., <xref ref-type="bibr" rid="B59">2011</xref>; Leen and Shea-Brown, <xref ref-type="bibr" rid="B25">2013</xref>).</p>
<p>In more detail, each cell <italic>j</italic> received an independent input <italic>I</italic><sub><italic>j</italic></sub> and a &#x0201C;triplet&#x0201D;&#x02014;(global) input <italic>I</italic><sub><italic>c</italic></sub> that is shared among all three cells. Comparison of the total input <italic>S</italic><sub><italic>j</italic></sub> &#x0003D; <italic>I</italic><sub><italic>c</italic></sub> &#x0002B; <italic>I</italic><sub><italic>j</italic></sub> with a threshold &#x00398; determined whether or not the cell spiked in that random draw. An additional parameter, <italic>c</italic>, identified the fraction of the total input variance &#x003C3;<sup>2</sup> originating from the global input; that is, <italic>c</italic> &#x02261; Var[<italic>I</italic><sub><italic>c</italic></sub>]/Var[<italic>I</italic><sub><italic>c</italic></sub> &#x0002B; <italic>I</italic><sub><italic>j</italic></sub>]. The global input was chosen from one of several marginal distributions, which included those used in the RGC model: Gaussian, bimodal, and heavy-tailed. The independent inputs <italic>I</italic><sub><italic>j</italic></sub> were, in all cases, chosen from a Gaussian distribution, consistent with our RGC model. When the common inputs are Gaussian, our model is equivalent to the Dichotomized Gaussian model previously studied by several groups (Amari et al., <xref ref-type="bibr" rid="B2">2003</xref>; Macke et al., <xref ref-type="bibr" rid="B26">2009</xref>, <xref ref-type="bibr" rid="B27">2011</xref>; Yu et al., <xref ref-type="bibr" rid="B59">2011</xref>), cf. (Tchumatchenko et al., <xref ref-type="bibr" rid="B52">2010</xref>). For further details, see section 4.2.</p>
<p>In the RGC model large effects in <italic>D</italic><sub>KL</sub> were accompanied by a striking increase in the bi- or multi-modality of excitatory conductances. Why are bimodal inputs, shared across cells, able to produce spiking responses that deviate from the pairwise model? We use our simple thresholding model to provide some intuition for how bimodal common inputs to thresholding cells lead to spiking probabilities that violate the constraints (Equation 3) which must hold for the pairwise model. For example, suppose that the common input <italic>I</italic><sub><italic>c</italic></sub> can take on values that cluster around two separated values, &#x003BC;<sub><italic>A</italic></sub> &#x0003C; &#x003BC;<sub><italic>B</italic></sub>, but rarely in the interval between; that is, the distribution of <italic>I</italic><sub><italic>c</italic></sub> is <italic>bimodal</italic>. If &#x003BC;<sub><italic>B</italic></sub> is large enough to push the cells over threshold but &#x003BC;<sub><italic>A</italic></sub> is not, then we see that any contribution to the right-hand side of Equation (3), <italic>p</italic><sub>2</sub>/<italic>p</italic><sub>1</sub>, depends only on the distribution of the independent inputs <italic>I</italic><sub><italic>j</italic></sub>; if either one or two cells spike, then the common input must have been drawn from the cluster of values around &#x003BC;<sub><italic>A</italic></sub>, because otherwise all three cells would have spiked.</p>
<p>To be concrete, let <italic>P</italic>[<bold>x</bold>] refer to the probability of spiking event <bold>x</bold> &#x0003D; (<italic>x</italic><sub>1</sub>, <italic>x</italic><sub>2</sub>, <italic>x</italic><sub>3</sub>), and <italic>P</italic>[<bold>x</bold> | <italic>I</italic><sub><italic>c</italic></sub> &#x02248; &#x003BC;<sub><italic>A</italic></sub>] refer to the probability that <bold>x</bold> occurs, conditioned on the event <italic>I</italic><sub><italic>c</italic></sub> &#x02248; &#x003BC;<sub><italic>A</italic></sub>. Then</p>
<disp-formula id="E8"><mml:math id="M58"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>+</mml:mo><mml:mtext>&#x02009;</mml:mtext><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02223;</mml:mo></mml:mrow> </mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>B</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>B</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>because <italic>P</italic>[(1, 0, 0) | I<sub><italic>c</italic></sub> &#x02248; &#x003BC;<sub><italic>B</italic></sub>] &#x0003D; 0: for the same reason,</p>
<disp-formula id="E9"><mml:math id="M59"><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>therefore</p>
<disp-formula id="E10"><mml:math id="M60"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mfrac><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>On the other hand,</p>
<disp-formula id="E11"><mml:math id="M61"><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>B</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02248;</mml:mo><mml:msub><mml:mi>&#x003BC;</mml:mi><mml:mi>A</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
<p>By changing the relative likelihood of drawing the common input from one cluster or the other, without changing the values of &#x003BC;<sub><italic>A</italic></sub> and &#x003BC;<sub><italic>B</italic></sub> themselves (that is, change <italic>P</italic>[<italic>I</italic><sub><italic>c</italic></sub> &#x02248; &#x003BC;<sub><italic>B</italic></sub>] and <italic>P</italic>[<italic>I</italic><sub><italic>c</italic></sub> &#x02248; &#x003BC;<sub><italic>A</italic></sub>] but leave the conditional probabilities (e.g., <italic>P</italic>[(1, 0, 0) | <italic>I</italic><sub><italic>c</italic></sub> &#x02248; &#x003BC;<sub><italic>A</italic></sub>]) fixed) one may change the ratio <italic>p</italic><sub>3</sub>/<italic>p</italic><sub>0</sub> <italic>without</italic> changing the ratio <italic>p</italic><sub>2</sub>/<italic>p</italic><sub>1</sub>. Hence the constraint specifying those network responses exactly describable by PME models can be violated when the common input is bimodal.</p>
<p>In contrast, we may instead consider a <italic>unimodal</italic> common input, of which a Gaussian is a natural example. Here, the distribution of the common input <italic>I</italic><sub><italic>c</italic></sub> is completely described by its mean and variance; both parameters can impact the ratio <italic>p</italic><sub>3</sub>/<italic>p</italic><sub>0</sub> (by altering the likelihood that the common input alone can trigger spikes) and the ratio <italic>p</italic><sub>2</sub>/<italic>p</italic><sub>1</sub>. Each value of <italic>I</italic><sub><italic>c</italic></sub> is consistent with both events <italic>p</italic><sub>1</sub> and <italic>p</italic><sub>2</sub>, with the relative likelihood of each event depending on the specific value of <italic>I</italic><sub><italic>c</italic></sub>; it is no longer clear how to separate the two events. In the following sections, we will confirm this intuition by direct evaluation of the resulting departure from pairwise statistics.</p>
</sec>
<sec>
<title>2.3.2. Model input distributions</title>
<p>Motivated by our observations of excitatory currents that arose in the RGC model, we chose several input distributions that allow us to explore other salient features, such as symmetry and the probability of large events. A distribution is called <italic>sub-Gaussian</italic> if the probability of large events decays rapidly with event size, so that it can be bounded above by a scaled Gaussian distribution (see section 4). We considered two sub-Gaussian distributions; the Gaussian itself, and a skewed distribution with a sub-Gaussian tail (hereafter referred to as &#x0201C;skewed&#x0201D;). We also considered the two &#x0201C;heavy-tailed&#x0201D; distributions used as stimuli to the RGC model&#x02014;the Cauchy distribution, and a skewed distribution with a Cauchy-like tail (hereafter referred to as &#x0201C;heavy-tailed skewed&#x0201D;). In these distributions, the probability of large events decays polynomially rather than exponentially.</p>
<p>For each choice of common input marginal, we varied the input parameters so as to explore a full range of firing rates and pairwise correlations: specifically, we varied the input correlation coefficient <italic>c</italic> in the range [0, 1], the <italic>total</italic> input standard deviation &#x003C3; in the range [0, 4], and the threshold &#x00398; in [0, 3]. In all cases the independent inputs <italic>I</italic><sub><italic>j</italic></sub> were chosen from a Gaussian distribution [of variance (1 &#x02212; <italic>c</italic>) &#x003C3;<sup>2</sup>]. For each choice of input parameters, we determine the resulting distribution on spiking states (as described in section 4) and compute the PME approximation.</p>
</sec>
<sec>
<title>2.3.3. Unimodal common inputs fail to produce significant higher-order interactions in three-cell feedforward circuits</title>
<p>We first considered common inputs chosen from a unimodal (e.g., Gaussian) distribution. If <italic>I</italic><sub><italic>c</italic></sub> is Gaussian, then the joint distribution of <bold>S</bold> &#x0003D; (<italic>S</italic><sub>1</sub>, <italic>S</italic><sub>2</sub>, <italic>S</italic><sub>3</sub>) is multivariate normal, and therefore characterized entirely by its means and covariances. Because the PME fit to a continuous distribution is precisely the multivariate normal that is consistent with the first and second moments, every such input distribution on <bold>S</bold> <italic>exactly</italic> coincides with its PME fit. However, even with Gaussian inputs, outputs (which are now in the binary state space {0, 1}<sup>3</sup>) will deviate from the PME fit (Amari et al., <xref ref-type="bibr" rid="B2">2003</xref>; Macke et al., <xref ref-type="bibr" rid="B26">2009</xref>). As shown below, non-Gaussian unimodal inputs can produce outputs with larger deviations. Nonetheless, these deviations are small for all cases in which inputs were chosen from a sub-Gaussian distribution, and PME models are quite accurate descriptions of circuits with a broad range of unimodal inputs.</p>
<p>We first considered circuits with either Gaussian or skewed common inputs. Over the full range of input parameters, distributions remained well fit by the pairwise model, with a maximum value of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M62"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) (of 0.0038 and 0.0035 for Gaussian and skewed inputs, respectively) achieved for high correlation values and &#x003C3; comparable to threshold. In Figure <xref ref-type="fig" rid="F7">7A</xref> we illustrate these trends with a contour plot of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M63"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) for a fixed value of threshold (here, &#x00398; &#x0003D; 1.5) and Gaussian common inputs (the analogous plot for skewed inputs is qualitatively very similar, Figure <xref ref-type="supplementary-material" rid="SM3">S3A</xref>).</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p><bold>Strength of higher-order interactions produced by the threshold model as input parameters vary, and the relationship of these higher-order interactions with other output firing statistics</bold>. <bold>(A)</bold> For Gaussian common inputs: <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M64"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) as a function of input correlation <italic>c</italic> and input standard deviation &#x003C3;, for a fixed threshold &#x00398; &#x0003D; 1.5. Color indicates <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M65"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>); see color bar for range. <bold>(B)</bold> For Gaussian common inputs: <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M66"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) vs. firing rate (Left) and the fraction of multi-information (&#x00394;) captured by the PME model vs. firing rate (Right). Each dot represents the value obtained from a single choice of the input parameters <italic>c</italic>, &#x003C3;, and &#x00398;; input parameters were varied over a broad range as described in section 2. Firing rate is defined as the probability of a spike occurring per cell per random draw of the sum-and-threshold model, as defined in Equation (16). Color indicates output correlation coefficient &#x003C1; ranging from black for &#x003C1; &#x02208; (0, 0.1), to white for &#x003C1; &#x02208; (0.9, 1), as illustrated in the color bars. <bold>(C,D)</bold>: as in <bold>(A,B)</bold>, but for Cauchy common inputs. <bold>(E,F)</bold>: as in <bold>(A,B)</bold>, but for bimodal common inputs.</p></caption>
<graphic xlink:href="fncom-08-00010-g0007.tif"/>
</fig>
<p>Clear patterns also emerged when we viewed <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M67"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) as a function of <italic>output</italic> spiking statistics rather than <italic>input</italic> statistics (as in Macke et al., <xref ref-type="bibr" rid="B27">2011</xref>). Non-linear spike generation can produce substantial differences between input and output correlations; this relationship can vary widely based on the specific non-linearity (Moreno et al., <xref ref-type="bibr" rid="B33">2002</xref>; de la Rocha et al., <xref ref-type="bibr" rid="B16">2007</xref>; Marella and Ermentrout, <xref ref-type="bibr" rid="B29">2008</xref>; Shea-Brown et al., <xref ref-type="bibr" rid="B47">2008</xref>; Vilela and Lindner, <xref ref-type="bibr" rid="B57">2009</xref>; Barreiro et al., <xref ref-type="bibr" rid="B5">2010</xref>, <xref ref-type="bibr" rid="B6">2012</xref>; Tchumatchenko et al., <xref ref-type="bibr" rid="B52">2010</xref>; Hong et al., <xref ref-type="bibr" rid="B19">2012</xref>). Figure <xref ref-type="fig" rid="F7">7B</xref> shows <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M68"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) and &#x00394; for all threshold values (including the data shown in Figure <xref ref-type="fig" rid="F7">7A</xref>), but now plotted with respect to the output firing rate. The data were segregated according to the Pearson&#x00027;s correlation coefficient &#x003C1; between the responses of cell pairs (<inline-formula><mml:math id="M69"><mml:mrow><mml:mi>&#x003C1;</mml:mi><mml:mo>&#x02261;</mml:mo><mml:mfrac><mml:mrow><mml:mtext>Cov</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:mtext>Var</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mtext>Var</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mover accent='true'><mml:mi>&#x003C1;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mi>&#x003BC;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mi>&#x003BC;</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003BC;</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mfrac></mml:mrow></mml:math></inline-formula>). For a fixed correlation, there was generally a one-to-one relationship between firing rate and <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M70"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>). For these distributions (Figure <xref ref-type="fig" rid="F7">7B</xref>, for Gaussian inputs; skewed inputs shown in Figure <xref ref-type="supplementary-material" rid="SM3">S3B</xref>), <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M71"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) was maximized at an intermediate firing rate. Additionally, <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M72"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) had a non-monotonic relationship with spike correlation: it increased from zero for low values of correlation, obtained a maximum for an intermediate value, and then decreased. These limiting behaviors agree with intuition: a spike pattern that is completely uncorrelated can be described by an independent distribution (a special case of PME model), and one that is perfectly correlated can be completely described via (perfect) pairwise interactions alone.</p>
<p>We next considered circuits in which inputs were drawn from one of two heavy-tailed distributions, the Cauchy distribution and a heavy-tailed skewed distribution, defined earlier. Here, distinctly different patterns emerge: for a fixed &#x00398;, <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M73"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) is maximized in regions of high input correlation and high input variance &#x003C3;, but relatively high values of <italic>D</italic><sub>KL</sub> are achievable across a wide range of input values (see Figure <xref ref-type="fig" rid="F7">7C</xref> for Cauchy inputs; heavy-tailed skewed in Figure <xref ref-type="supplementary-material" rid="SM3">S3C</xref>). However, the maximum achievable values of <italic>D</italic><sub>KL</sub> were achieved at intermediate <italic>output</italic> correlations &#x003C1; &#x02248; 0.4 (see Figure <xref ref-type="fig" rid="F7">7D</xref> for Cauchy inputs; heavy-tailed skewed shown in Figure <xref ref-type="supplementary-material" rid="SM3">S3D</xref>); this suggests that high input correlations do not result in high output correlations.</p>
<p>This somewhat unintuitive finding may be explained by the structure of the PDF of a heavy-tailed common input, which favors (infrequent) large events at the expense of medium-size events. For instance, the probability that a Cauchy input is above a given threshold (<italic>P</italic>[<italic>I</italic><sub><italic>c</italic></sub> &#x0003E; &#x00398; &#x0003E; <bold>E</bold>[<italic>I</italic><sub><italic>c</italic></sub>]]) is often much smaller than for a Gaussian distribution of the same variance. However, an input can trigger at best one single spiking event regardless of size: therefore a Cauchy common input generates fewer correlated spiking events with larger inputs, while a Gaussian common input triggers correlated spiking events with smaller, but more frequent, input values. As a result, heavy-tailed inputs are unable to explore the full range of output firing statistics: Figure <xref ref-type="fig" rid="F7">7D</xref> shows that high output correlations only occur at very low firing rates. Overall, <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M74"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) reaches higher numerical values than for sub-Gaussian inputs, possibly reflecting the higher-order statistics in the input. However, the maximal <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M75"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) attained still falls far short of exploring the full range of possible values (compare with Figure <xref ref-type="fig" rid="F1">1B</xref>).</p>
<p>Finally, we examine the behavior of the <italic>strain</italic>, which quantifies both the magnitude and sign of deviation from the pairwise model (see Ohiorhenuan and Victor, <xref ref-type="bibr" rid="B37">2010</xref>). It has been previously observed that the strain is negative for the DG model (Macke et al., <xref ref-type="bibr" rid="B27">2011</xref>), a condition that has been related to sparsity of the neural code and with which our results agree (data not shown). However, we found that any other choice of input marginal statistics, both positive and negative values are seen; for heavy-tailed common inputs, positive values predominated except at very low firing rates.</p>
</sec>
<sec>
<title>2.3.4. Bimodal triplet inputs can generate higher-order interactions in three-cell feedforward circuits</title>
<p>Having shown that a wide range of unimodal common inputs produced spike patterns that are well-approximated by PME fits, we next examined bimodal common inputs. Such inputs substantially increased departures from PME fits in the ganglion cell models described above. As in the previous section, we varied <italic>c</italic>, &#x003C3;, and &#x00398; so as to explore a full range of firing rates and pairwise correlations.</p>
<p>As a function of input parameter values, <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M76"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) is maximized for large input correlation and moderate input variance &#x003C3;<sup>2</sup> [see Figure <xref ref-type="fig" rid="F7">7E</xref>, which illustrates <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M77"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) for a fixed threshold &#x00398; &#x0003D; 1.5]. Figure <xref ref-type="fig" rid="F7">7F</xref> shows <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M78"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) values as a function of the firing rate and pairwise correlation elicited by the full range of possible bimodal inputs. We see that <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M79"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) is maximized at an intermediate (but relatively high: &#x003BD; &#x02248; 0.4) firing rate, and for intermediate-to-large correlation values (&#x003C1; &#x02248; 0.6 &#x02212; 0.8).</p>
<p>We find distinctly different results when we view &#x00394; (Equation 1), for these same simulations, as a function of output spiking statistics (right panels of Figures <xref ref-type="fig" rid="F7">7B,D,F</xref>). For unimodal, sub-Gaussian distributions (Figure <xref ref-type="fig" rid="F7">7B</xref>), &#x00394; is very close to 1, with the few exceptions at extreme firing rates. For heavy-tailed and bimodal inputs (Figures <xref ref-type="fig" rid="F7">7D,F</xref>), &#x00394; may be appreciably far from 1 (as small as 0.5) with the smallest numbers (suggesting a poor fit of the pairwise model) occurring for low correlation &#x003C1;. This highlights one interesting example where these two metrics for judging the quality of the pairwise model, <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M80"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) and &#x00394;, yield contrasting results.</p>
<p>Finally, we emphasize that while bimodal inputs can produce greater higher-order interactions than unimodal inputs, the values of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M81"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) accessible by feedforward circuits with global inputs remain far below their upper bounds at any given firing rate. The maximal values of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M82"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) reached by Cauchy and heavy-tailed skewed inputs were 0.0078 and 0.0153; bimodal common inputs reached a maximal value of 0.091. This is an order of magnitude smaller than possible departures among symmetric spike patterns (compare Figure <xref ref-type="fig" rid="F1">1B</xref>). The difference is illustrated in Figure <xref ref-type="supplementary-material" rid="SM4">S4</xref>, which compares the <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M83"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) values obtained in the thresholding model and those obtained by direct exhaustive search at each firing rate by superposing the datapoints on a single axis.</p>
</sec>
<sec>
<title>2.3.5. Mathematical analysis of unimodal vs. bimodal effects</title>
<p>The central finding above is that circuits with bimodal inputs can generate significantly greater higher-order interactions than circuits with unimodal inputs. To probe this further, we investigated the behavior of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M84"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) for the feedforward threshold model with a perturbation expansion in the limit of small common input. We found that as the strength of common input signals increased, circuits with bimodal inputs diverged from the PME fit more rapidly than circuits with unimodal inputs; the full calculation is given in the Appendix. In brief, we determined the leading order behavior of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M85"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) in the strength <italic>c</italic> of (weak) common input. <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M86"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) depended on <italic>c</italic><sup>3</sup> for unimodal distributions, i.e., the low order terms in <italic>c</italic> dropped out; for symmetric unimodal distributions, such as a Gaussian, <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M87"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) grew as <italic>c</italic><sup>4</sup>. For bimodal distributions, <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M88"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) grew as <italic>c</italic><sup>2</sup>. Because of the <italic>c</italic><sup>2</sup> dependence, rather than <italic>c</italic><sup>3</sup> or <italic>c</italic><sup>4</sup>, as the strength of common input signals <italic>c</italic> increases, circuits with bimodal inputs are predicted to produce greater deviations from their PME fits.</p>
</sec>
<sec>
<title>2.3.6. Impact of recurrent coupling</title>
<p>We next modified our thresholding model to incorporate the effects of recurrent coupling among the spiking cells. To mimic gap junction coupling in the RGC circuit, we considered all-to-all, excitatory coupling, and assumed that this coupling occurs on a faster timescale compared with the timescale over which inputs arrive at the cells.</p>
<p>Our previous model was extended as follows: if the inputs arriving at each cell elicited any spikes, there was a second stage at which the input to each neuron receiving a connection from a spiking cell was increased by an amount <italic>g</italic>. This represented a rapid depolarizing current, assumed for simplicity to add linearly to the input currents. If the second stage resulted in additional spikes, the process was repeated: recipient cells received an additional current <italic>g</italic>, and their summed inputs were again thresholded. The sequence terminated when no new spikes occurred on a given stage; e.g., for <italic>N</italic> &#x0003D; 3, there were a maximum of three stages. The spike pattern recorded on a given trial was the total number of spikes generated across all stages.</p>
<p>We then explored the impact of varying <italic>g</italic> for a single representative value of &#x003C3; and &#x00398;, and several values of the correlation coefficient <italic>c</italic>. We found that as <italic>g</italic> increased <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M89"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) varied smoothly, reflecting the underlying changes in the spike count distribution. For small <italic>c</italic> (<italic>c</italic> &#x0003D; 0.02 shown in Figure <xref ref-type="fig" rid="F8">8A</xref>), where the variance of common input is very small, the results varied little by input type: for all input types <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M90"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) reached an interior maximum near <italic>g</italic> &#x02248; 1.7. As <italic>c</italic> increases, the distinctions between inputs types become apparent (Figures <xref ref-type="fig" rid="F8">8B,C</xref> show <italic>c</italic> &#x0003D; 0.2, 0.5, respectively): for most input types and values of <italic>c</italic>, the value of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M91"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) reaches an interior maximum that exceeds its value without coupling (i.e., <italic>g</italic> &#x0003D; 0). However, overall values of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M92"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) remained modest, never exceeding 0.01 across the values explored here.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p><bold>The impact of recurrent coupling on the three-cell sum-and-threshold model</bold>. Each plot shows <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M93"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) as a function of <italic>g</italic>, for a specific value of the correlation coefficient. In all panels, input standard deviation &#x003C3; &#x0003D; 1, threshold &#x00398; &#x0003D; 1.5, <italic>N</italic> &#x0003D; 3 and symbols are as described in the legend for <bold>(C)</bold>. Abbreviations in the legend denote the marginal distribution of the common input: G, Gaussian; SK, skewed; C, Cauchy; HT, heavy-tailed skewed; B, bimodal. <bold>(A)</bold> For input correlation <italic>c</italic> &#x0003D; 0.02, <bold>(B)</bold> <italic>c</italic> &#x0003D; 0.2, and <bold>(C)</bold> <italic>c</italic> &#x0003D; 0.5.</p></caption>
<graphic xlink:href="fncom-08-00010-g0008.tif"/>
</fig>
</sec>
<sec>
<title>2.3.7. Summary of findings for simplified circuit model</title>
<p>We examined a highly idealized model of neural spiking, so as to explore the generality of our earlier findings in a small array of RGC models. We found that our main results from the RGC model&#x02014;that higher-order interactions were most significant when inputs had bimodal structure, and that when fast excitatory recurrence was added to the circuit, higher-order interactions were maximized at an intermediate value of the recurrence strength&#x02014;persisted in this simplified model. Moreover, we were able to show that the first of these findings is general, in that it holds over a complete exploration of parameter space.</p>
</sec>
</sec>
<sec>
<title>2.4. Scaling of higher-order interactions with population size</title>
<p>The results above suggest that unimodal, rather than bimodal, input statistics contribute to the success of PME models. Next, we examined whether this conclusion continues to hold when we increase network size. The permutation-symmetric architectures we have considered so far can be scaled up to more than three cells in several natural ways; for example, we can study <italic>N</italic> cells with a global common input.</p>
<p>We considered a sequence of models in which a set of <italic>N</italic> threshold spiking units received global input <italic>I</italic><sub><italic>c</italic></sub> [with mean 0 and variance &#x003C3;<sup>2</sup><italic>c</italic>] and an independent input <italic>I</italic><sub><italic>j</italic></sub> [with mean 0 and variance &#x003C3;<sup>2</sup> (1 &#x02212; <italic>c</italic>)]. As for the three-cell network models considered previously, the output of each cell was determined by summing and thresholding these inputs. Upon computing the probability distribution of network outputs (section 4), we fit a PME distribution. Again, we explored a range of &#x003C3;, <italic>c</italic>, and &#x00398; and recorded the maximum value of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M94"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) between the observed distribution <italic>P</italic> and its PME fit <inline-formula><mml:math id="M95"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>. Figure <xref ref-type="fig" rid="F9">9</xref> shows this <italic>D</italic><sub>KL</sub>/<italic>N</italic> [i.e., entropy per cell (Macke et al., <xref ref-type="bibr" rid="B26">2009</xref>)] for each class of marginal distributions.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p><bold>The significance of higher-order interactions increases with network size</bold>. <bold>(A)</bold> Normalized maximal deviation, <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M96"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>)/<italic>N</italic>, from the PME fit for the thresholding circuit model as network size <italic>N</italic> increases. For each <italic>N</italic> and common input distribution type, possible input parameters were in the following ranges: input correlation <italic>c</italic> &#x02208; [0, 1], input standard deciation &#x003C3; &#x02208; [0, 4], and threshold &#x00398; &#x02208; [0, 3]. <bold>(B)</bold> Example sample distributions for different types of common input: from top, bimodal, Gaussian, heavy-tailed skew, and Cauchy common inputs. For each input type, the distribution that maximized <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M97"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) for <italic>N</italic> &#x0003D; 16 is shown. Each distribution is illustrated with a bar plot contrasting the probabilities of spiking events in the true (dark blue) vs. pairwise maximum entropy (light pink) distributions.</p></caption>
<graphic xlink:href="fncom-08-00010-g0009.tif"/>
</fig>
<p>We found that the maximum <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M98"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>)/<italic>N</italic> increased roughly linearly with <italic>N</italic> for Gaussian, skewed and Cauchy inputs; for heavy-tailed skew and bimodal inputs, <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M99"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>)/<italic>N</italic> appeared to saturate after an initial increase (Figure <xref ref-type="fig" rid="F9">9</xref>). The relative ordering for unimodal inputs shifted as <italic>N</italic> increased; as <italic>N</italic> &#x02192; 16, the maximal achievable <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M100"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) for sub-Gaussian inputs overtook the values for heavy-tailed inputs. At all values of <italic>N</italic>, the values for Gaussian and skewed inputs tracked one another closely. Regardless, the values for all unimodal inputs remained substantially below the maximal value achievable for bimodal inputs. Figure <xref ref-type="fig" rid="F9">9B</xref> shows that the probability distributions produced by these inputs qualitatively agree with this trend: departures from PME were more visually pronounced for global bimodal inputs than for global unimodal inputs. In addition, the distributions for heavy-tailed and sub-Gaussian inputs differed qualitatively, offering a potential mechanism for different scaling behavior. Using the relationship between <italic>D</italic><sub>KL</sub> and likelihood ratios (described in section 2.1), at <italic>N</italic> &#x0003D; 16, the value <italic>D</italic><sub>KL</sub>/<italic>N</italic> &#x02248; 0.1 for bimodal global inputs corresponds to a likelihood ratio of 0.33 that a single draw from <italic>P</italic> (single network output) in fact came from the PME fit <inline-formula><mml:math id="M101"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula> rather than from <italic>P</italic>; a likelihood &#x0003C;0.01 is reached for four draws.</p>
<p>We next extended our model with recurrent coupling to <italic>N</italic> &#x0003E; 3 cells. In addition to the parameters for the uncoupled network, we varied the coupling strength, <italic>g</italic>, for each type of input. As in the <italic>N</italic> &#x0003D; 3 network, coupling was all-to-all. As for the small networks explored in an earlier section, <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M102"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) generally peaked at an intermediate value of the coupling strength <italic>g</italic>; however, the value of <italic>g</italic> decreased as population size <italic>N</italic> increased (illustrated in Figure <xref ref-type="fig" rid="F10">10A</xref>, for <italic>c</italic> &#x0003D; 0.2). This may be attributed to the increased potential impact of recurrence at larger population sizes; as <italic>N</italic> increases, the number of potential <italic>additional</italic> spikes that may be triggered increases; consequently the average recurrent excitation received by each cell increases, and therefore the probability that one or two spikes will trigger a cascade to <italic>N</italic> spikes. In Figure <xref ref-type="fig" rid="F10">10B</xref> we demonstrate that the impact of this effect may be captured by plotting <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M103"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) as a function of an <italic>effective</italic> coupling parameter, <italic>g</italic><sup>&#x0002A;</sup><italic>N</italic>/3. Here, we plot the curves for six population sizes (<italic>N</italic> &#x0003D; 3, 4, 6, 8, 10, and 12) and five common input types; each curve was scaled by normalizing <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M104"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) by its maximum value. For many sets of parameter values, the resulting curves line up remarkably well, suggesting a universal scaling with the effective coupling parameter.</p>
<fig id="F10" position="float">
<label>Figure 10</label>
<caption><p><bold>The impact of recurrent coupling on the sum-and-threshold model, for increasing population size</bold>. <bold>(A)</bold> <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M105"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) as a function of the coupling coefficient, <italic>g</italic>, for a specific value of population size <italic>N</italic>. In all plots, input standard deviation &#x003C3; &#x0003D; 1, threshold &#x00398; &#x0003D; 1.5 and input correlation <italic>c</italic> &#x0003D; 0.2. From top: <italic>N</italic> &#x0003D; 4; <italic>N</italic> &#x0003D; 8; <italic>N</italic> &#x0003D; 12. <bold>(B)</bold> <italic>D</italic><sup>norm</sup><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M106"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) as a function of the coupling coefficient, <italic>g</italic>, for populations sizes <italic>N</italic> &#x0003D; 3-12. For each curve, <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M107"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) was scaled by its maximal value and plotted as a function of the scaled coupling coefficient, <italic>g</italic><sup>&#x0002A;</sup><italic>N</italic>/3, to illustrate a universal scaling with effective coupling strength. The line style of each curve indicates the population size <italic>N</italic>, as listed in the legend. The marker and line color indicate the common input marginal, as listed in the legend for <bold>(A)</bold>. <bold>(C)</bold> (Top) Maximal value of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M108"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>)/<italic>N</italic>, achieved over a survey of parameter values <italic>c</italic>, &#x003C3;, &#x00398;, and <italic>g</italic>, as a function of the population size <italic>N</italic> (solid lines). For each input marginal type, a second curve shows the maximal value obtained over only feed-forward simulations (<italic>g</italic> &#x0003D; 0; dashed lines). The marker and line color indicate the common input marginal, as listed in the legend for <bold>(A)</bold>. (Bottom) Maximal value of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M109"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>)/<italic>k</italic>, achieved over a survey of parameter values <italic>c</italic>, &#x003C3;, &#x00398;, and <italic>g</italic>, as a function of the <italic>subsample</italic> population size <italic>k</italic>. Data was subsampled from the <italic>N</italic> &#x0003D; 8 data shown in the top panel, by restricting analysis to <italic>k</italic> out of <italic>N</italic> cells.</p></caption>
<graphic xlink:href="fncom-08-00010-g0010.tif"/>
</fig>
<p>We also explored the overall possible impact of recurrence on higher-order interactions, by surveying a range of circuit parameters <italic>c</italic>, &#x003C3;, &#x00398; and <italic>g</italic>. The top panel of Figure <xref ref-type="fig" rid="F10">10C</xref> shows the maximal <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M110"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) per neuron, for each type of input, up to population size <italic>N</italic> &#x0003D; 8. For unimodal inputs, recurrent coupling increased the available range of higher-order interactions modestly, compared with the range achieved with purely feedforward connections; however, these values remained significantly lower than those achieved for bimodal inputs.</p>
<p>Finally, we considered how higher-order interactions scale with population sampling size. The spike pattern distributions used to generate the last column of data points (<italic>N</italic> &#x0003D; 8) in the top panel of Figure <xref ref-type="fig" rid="F10">10C</xref> were reanalyzed by sub-sampling the spike pattern distributions on <italic>k</italic> &#x0003C; 8 cells. In each case, we chose our sub-population to be <italic>k</italic> nearest neighbors (for our setup, any subset of <italic>k</italic> cells is statistically identical). In the bottom panel of Figure <xref ref-type="fig" rid="F10">10C</xref>, we show the maximal value of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M111"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) per sub-sampled cell achieved over all input parameters (the curves for Gaussian, skewed and Cauchy inputs are so close together so as to be visually indistinguishable). This number increases or remains steady as <italic>k</italic> increases, indicating that sub-sampling a coupled network will depress the apparent higher-order interactions in the output spiking pattern.</p>
<p>To summarize, the greater impact of bimodal vs. unimodal input statistics on maximal values of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M112"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) persists in circuits with <italic>N</italic> &#x0003D; 3 cells up to <italic>N</italic> &#x0003D; 16 cells. Overall, for the circuit parameters producing maximal deviations from PME fits, it becomes easier to statistically distinguish between spiking distributions and their PME fits as the number of cells increases in feedforward networks.</p>
</sec>
</sec>
<sec sec-type="discussion" id="s3">
<title>3. Discussion</title>
<p>We used mechanistic models to identify input patterns and circuit mechanisms which produce spike patterns with significant higher-order interactions&#x02014;that is, with substantial deviations from predictions under a PME model. We focused on a tractable setting of small, symmetric circuits with common inputs. This revealed several general principles. First, we found that these circuits produced outputs that were much closer to PME predictions than required for a general spiking pattern. Second, bimodal input distributions produced stronger higher-order interactions than unimodal distributions. Third, recurrent excitatory or gap junction coupling could produce a further, moderate increase of higher-order correlations; the effect was greatest for coupling of intermediate strength.</p>
<p>These general results held for both an abstract threshold-and-spike model and for networks of non-linear integrate-and-fire units based on measured properties of one class of RGCs. Together with the facts that ON parasol cell filtering suppresses bimodality in light input, and that coupling among ON parasol cells is relatively weak, our findings provide an explanation for why their population activity is well captured by PME models.</p>
<sec>
<title>3.1. Comparison with empirical studies</title>
<p>How do our maximum entropy fits compare with empirical studies? In terms of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M113"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>)&#x02014;equivalently, the logarithm of the average relative likelihood that a sequence of data drawn from <italic>P</italic> was instead drawn from the model <inline-formula><mml:math id="M114"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>&#x02014;numbers obtained from our RGC models are very similar to those obtained by <italic>in vitro</italic> experiments on primate RGCs (Shlens et al., <xref ref-type="bibr" rid="B49">2006</xref>, <xref ref-type="bibr" rid="B48">2009</xref>). For example, in a survey of 20 numerical experiments under constant light conditions (each of length 100 ms, with spikes binned in 10 ms intervals), we find that <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M115"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) ranges between 0 and 0.00029: similarly excellent fits were found by Shlens et al. (<xref ref-type="bibr" rid="B49">2006</xref>) (in which cell triplets were stimulated by constant light for 60 s with spikes binned at 10 ms), with one example given of 0.0008 (inferred from a reported likelihood ratio of 0.99944). These values can increase by an order of magnitude under full-field stimulation, as well as spatio-temporally varying stixel simulations (bounded above by 0.007). We can view the 60 &#x003BC;m stixel simulations as a model of the checkerboard experiments of Shlens et al. (<xref ref-type="bibr" rid="B49">2006</xref>), for which close fits by the PME distribution were also observed (likelihood numbers were not reported). Similarly, the values of &#x00394; produced by our RGC model are close to those found by Schneidman et al. (<xref ref-type="bibr" rid="B44">2006</xref>); Shlens et al. (<xref ref-type="bibr" rid="B49">2006</xref>) under comparable stimulus conditions. We obtain &#x00394; &#x0003D; 99.5% (for cell group size <italic>N</italic> &#x0003D; 3) under constant illumination, which is near the range reported by Shlens et al. (<xref ref-type="bibr" rid="B49">2006</xref>) for the same bin size and stimulus conditions (98.6 &#x000B1; 0.5, <italic>N</italic> &#x0003D; 3 &#x02212; 7). For full-field stimuli we find a range of numbers from 95.7% to 99.3% (<italic>N</italic> &#x0003D; 3).</p>
<p>With regard to the circuit mechanisms behind these excellent fits by pairwise models, the findings that most directly address the experimental settings of Shlens et al. (<xref ref-type="bibr" rid="B49">2006</xref>, <xref ref-type="bibr" rid="B48">2009</xref>), are (1) the finding that in the threshold model, unimodal inputs generate minimal higher-order interactions, compared to bimodal inputs, and (2) the particular stimulus filtering properties of parasol cells can suppress bimodality that may be present in an input stimulus, resulting in a unimodal distribution of input currents. First, we believe that unimodal inputs are consistent with the white-noise checkboard stimuli used in Shlens et al. (<xref ref-type="bibr" rid="B49">2006</xref>, <xref ref-type="bibr" rid="B48">2009</xref>), where binary pixels were chosen to be small relative to the receptive field size; averaged over the spatial receptive field, they would be expected to yield a single Gaussian input by the central limit theorem. Second, temporal filtering may contribute to receipt of unimodal conductance inputs by cells for the full-field binary flicker stimuli that are delivered in Schneidman et al. (<xref ref-type="bibr" rid="B44">2006</xref>). With the 16.7 ms refresh rate used there, under the assumption that the filter time-scale of the cells studied in that paper is roughly similar to that of the ON parasol cell we consider, the filter would average a binary (and hence bimodal) stimulus into a unimodal shape (see Figure <xref ref-type="fig" rid="F2">2C</xref>, for example).</p>
<p>The simple threshold models that we have considered, meanwhile, give us a roadmap for how circuits could be driven in such a way as to lower &#x00394;. The right columns of Figures <xref ref-type="fig" rid="F7">7B,D,F</xref> show &#x00394; plotted as a function of firing rate for circuits of <italic>N</italic> &#x0003D; 3 cells receiving global common inputs; we observe that &#x00394; &#x02248; 1 for Gaussian inputs over a broad range of firing rates and pairwise correlation coefficients, but that values of &#x00394; can be depressed by 25&#x02013;50% in the presence of a bimodal common input. Indeed, Shlens et al. (<xref ref-type="bibr" rid="B49">2006</xref>) showed that adding global bimodal inputs to a purely pairwise model can lead to a comparable departure in &#x00394;. Our results are consistent with this finding, and explicitly demonstrate that the bimodality of the inputs&#x02014;as well as their global projection&#x02014;are characteristics that lead to this departure.</p>
</sec>
<sec>
<title>3.2. Consequences for specific neural circuits</title>
<p>Our results make predictions about when neural circuits are likely to generate higher-order interactions. A comprehensive study of our simple thresholding model shows that bimodal inputs generate greater beyond-pairwise interactions than unimodal inputs. This result can be extended to other circuits where a clear input&#x02013;output relationship exists, and be used to predict higher-order correlations by analyzing the impact of stimulus filtering on a statistically defined class of inputs. For example, the effect holds in our model of primate ON parasol cells, where a biphasic filter suppresses bimodality in a stimulus with a timescale matched to that of the filter. We can use these results to extrapolate to other classes of RGCs or other stimulus conditions in which filters are less biphasic (Victor, <xref ref-type="bibr" rid="B56">1999</xref>). Indeed, when we process long time-scale bimodal inputs through a preliminary model of the midget cell circuit, stimulus bimodality is no longer suppressed and is associated with higher-order interactions (see Figure <xref ref-type="fig" rid="F4">4</xref>). We predict that greater higher-order interactions will be found for stimuli or RGC circuits that elicit bimodal activity that is thresholded when generating spikes&#x02014;in comparison to the parasol circuits and stimuli studied in Shlens et al. (<xref ref-type="bibr" rid="B49">2006</xref>, <xref ref-type="bibr" rid="B48">2009</xref>). We believe that this principle will be further applicable in other sensory systems.</p>
<p>We found that recurrent excitatory connections further increase higher-order interactions, which are maximized at an intermediate recurrence strength; in particular, when the strength of an excitatory recurrent input was comparable to the distance between rest and threshold (Figure <xref ref-type="fig" rid="F8">8</xref>). For the primate ON parasol cells we considered, the experimentally measured strength of gap junction coupling would lead to an estimated membrane voltage jump of &#x02248; 1 mV in response to the firing of a neighboring RGC, while the voltage distance between the resting voltage and an approximate threshold is about 5&#x02013;10 mV (Trong and Rieke, <xref ref-type="bibr" rid="B55">2008</xref>). Consistent with this estimate, we found that in our ON parasol cell model, higher-order interactions were maximized when the strength of excitatory recurrence was eight times its experimentally measured value. The experimentally measured values of recurrence had little or no effect on higher-order interactions. We anticipate that this result may be used to predict whether recurrent coupling plays a role in generating higher-order interactions in other circuits where the average voltage jump produced by an electrical or synaptic connection can be measured.</p>
<p>To apply our findings to real circuits, we must also consider population size. A measurement from a neural circuit, in most cases, will be a subsample of a much larger, complete circuit. We addressed this question where it was computationally more tractable, for the thresholding model. Here, we found that the impact of higher-order interactions, as measured by entropy per cell unaccounted for by the pairwise model (<italic>D</italic><sub>KL</sub>/<italic>k</italic>), increases moderately as subsample size <italic>k</italic> increases. Since recurrent connectivity in our model is truly global, this is consistent with the suggestion of Roudi et al. (<xref ref-type="bibr" rid="B40">2009a</xref>) and others that the entropy can be expected to scale extensively with population size <italic>N</italic>, once <italic>N</italic> significantly exceeds the true spatial connectivity footprint: we may see different results with limited, local connectivity.</p>
</sec>
<sec>
<title>3.3. Scope and open questions</title>
<p>There are many aspects of circuits left unexplored by our study. Prominent among these is heterogeneity. Only a few of our simulations produce heterogeneous inputs to model RGCs, and all of our studies apply to cells with identical response properties. This is in contrast to studies such as Schneidman et al. (<xref ref-type="bibr" rid="B44">2006</xref>), which examine correlation structures among multiple cell types. For larger networks, feedforward connections with variable spatial profiles also occur, between the extremes of independent and global input connections examined here. It is also possible that more complex input statistics could lead to greater higher-order interactions (Bethge and Berens, <xref ref-type="bibr" rid="B8">2008</xref>). Finally, Figure <xref ref-type="fig" rid="F9">9</xref> indicates that some trends in <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M116"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) vs. N appear to become non-linear for <italic>N</italic> &#x02A86; 10; for larger networks, our qualitative findings could change.</p>
<p>Our study also leaves largely open the role of different retinal filters in generating higher-order interactions. We have found that the specific filtering properties of ON parasol cells suppress bimodality in light inputs, suggesting that other classes of RGCs, such as midget cells, may produce more robust higher-order interactions (compare panels in Figure <xref ref-type="fig" rid="F4">4B</xref>). This predicts a specific mechanism for the development of higher-order interactions in preparations that include multiple classes of ganglion cells (Schneidman et al., <xref ref-type="bibr" rid="B44">2006</xref>). For a complete picture, future studies will also need to account for the possible adaptation of stimulus filters in response to higher-order stimulus characteristics (Tkacik et al., <xref ref-type="bibr" rid="B53">2012</xref>); we did not consider the latter effect here, where our filter was fit to the response of a cell to Gaussian stimuli with specific mean and variance. An allied possibility is that multiple filters will be required, as was found when fitting the responses of salamander retinal cells to LN models (Fairhall et al., <xref ref-type="bibr" rid="B17">2006</xref>). Distinguishing the roles of linear filters vs. static non-linearities in determining which stimulus classes will give the greatest higher-order correlations is another important step. Finally, we considered circuits with a single step of inputs and simple excitatory or gap junction coupling; a plethora of other network features could also lead to higher-order interactions, including multi-layer feedforward structures, together with lateral and feedback coupling. We speculate that, in particular, such mechanisms could contribute to the higher-order interactions found in cortex (Tang et al., <xref ref-type="bibr" rid="B50">2008</xref>; Montani et al., <xref ref-type="bibr" rid="B32">2009</xref>; Ohiorhenuan et al., <xref ref-type="bibr" rid="B36">2010</xref>; Oizumi et al., <xref ref-type="bibr" rid="B38">2010</xref>; Koster et al., <xref ref-type="bibr" rid="B22">2013</xref>).</p>
<p>A final outstanding area of research is to link tractable network mechanisms for higher-order interactions with their impact (or lack of impact) on information encoded in neural populations (Kuhn et al., <xref ref-type="bibr" rid="B24">2003</xref>; Montani et al., <xref ref-type="bibr" rid="B32">2009</xref>; Oizumi et al., <xref ref-type="bibr" rid="B38">2010</xref>; Ganmor et al., <xref ref-type="bibr" rid="B18">2011</xref>; Cain and Shea-Brown, <xref ref-type="bibr" rid="B10">2013</xref>). A simple starting point is to consider rate-based population codes in which each stimulus produces a different &#x0201C;tuned&#x0201D; average spike count (see for e.g., chapter 3 of Dayan and Abbot, <xref ref-type="bibr" rid="B15">2001</xref>). One can then ask whether spike responses can be more easily decoded to estimate stimuli for the full population response (i.e., <italic>P</italic>) to each stimulus or for its pairwise approximation (<inline-formula><mml:math id="M117"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>). In our preliminary tests where higher-order correlations were created by inputs with bimodal distributions, we found examples where decoding of <italic>P</italic> vs. <inline-formula><mml:math id="M118"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula> differed substantially. However, a more complete study would be required before general conclusions about trends and magnitudes of the effect could be made; such a study would include complementary approach in which the full spike responses <italic>P</italic> are themselves decoded via a &#x0201C;mismatched&#x0201D; decoder based on the pairwise model <inline-formula><mml:math id="M119"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula> (Oizumi et al., <xref ref-type="bibr" rid="B38">2010</xref>). Overall, we hope that the present paper, as one of the first that connects circuit mechanisms to higher-order statistics of spike patterns, will contribute to future research that takes these next steps.</p>
</sec>
</sec>
<sec sec-type="materials and methods" id="s4">
<title>4. Materials and methods</title>
<sec>
<title>4.1. Experimentally-based model of a RGC circuit</title>
<p>We model the response of a individual RGC using data collected from a representative primate ON parasol cell, following methods in Murphy and Rieke (<xref ref-type="bibr" rid="B34">2006</xref>); Trong and Rieke (<xref ref-type="bibr" rid="B55">2008</xref>). Similar response properties were observed in recordings from 16 other cells. To measure the relationship between light stimuli and synaptic conductances, the retina was exposed to a full-field, white noise stimulus. The cell was voltage clamped at the excitatory (or inhibitory) reversal potential <italic>V</italic><sub><italic>E</italic></sub> &#x0003D; 0 mV (<italic>V</italic><sub><italic>I</italic></sub> &#x0003D; &#x02212;60 mV), and the inhibitory (or excitatory) currents were measured in response to the stimulus. These currents were then turned into equivalent conductances by dividing by the driving force of &#x000B1; 60 mV; in other words</p>
<disp-formula id="E12"><mml:math id="M120"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msup><mml:mi>g</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mi>I</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msup><mml:mo>/</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mi>V</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>E</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>;</mml:mo><mml:mtext>&#x02003;&#x02003;</mml:mtext><mml:mi>V</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>E</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mn>60</mml:mn><mml:mtext>&#x02009;mV</mml:mtext></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi>g</mml:mi><mml:mrow><mml:mtext>inh</mml:mtext></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mi>I</mml:mi><mml:mrow><mml:mtext>inh</mml:mtext></mml:mrow></mml:msup><mml:mo>/</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mi>V</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>I</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>;</mml:mo><mml:mtext>&#x02003;&#x02003;</mml:mtext><mml:mi>V</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>I</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>60</mml:mn><mml:mtext>&#x02009;mV</mml:mtext></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The time-dependent conductances <italic>g</italic><sup>exc</sup> and <italic>g</italic><sup>inh</sup> were now injected into a different cell using a dynamic clamp procedure (i.e., the input current was varied rapidly to maintain the correct relationship between the conductance and the membrane voltage) and the voltage was measured at a resolution of 0.1 ms.</p>
<sec>
<title>4.1.1. Stimulus filtering</title>
<p>To model the relationship between the light stimulus and synaptic conductances, the current measurements <italic>I</italic><sup>exc</sup> and <italic>I</italic><sup>inh</sup> were fit to a linear-nonlinear model:</p>
<disp-formula id="E13"><mml:math id="M121"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msup><mml:mi>g</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msup><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mi>N</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msup><mml:mi>L</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msup><mml:mo>&#x02217;</mml:mo><mml:mi>s</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:msup><mml:mi>&#x003B7;</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi>g</mml:mi><mml:mrow><mml:mtext>inh</mml:mtext></mml:mrow></mml:msup><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mi>N</mml:mi><mml:mrow><mml:mtext>inh</mml:mtext></mml:mrow></mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msup><mml:mi>L</mml:mi><mml:mrow><mml:mtext>inh</mml:mtext></mml:mrow></mml:msup><mml:mo>&#x02217;</mml:mo><mml:mi>s</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:msup><mml:mi>&#x003B7;</mml:mi><mml:mrow><mml:mtext>inh</mml:mtext></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>s</italic> is the stimulus, <italic>L</italic><sup>exc</sup> (<italic>L</italic><sup>inh</sup>) is a linear filter, <italic>N</italic><sup>exc</sup> (<italic>N</italic><sup>inh</sup>) is a non-linear function, and &#x003B7;<sup>exc</sup> (&#x003B7;<sup>inh</sup>) is a noise term. The linear filter was fit by the function</p>
<disp-formula id="E14"><label>(7)</label><mml:math id="M122"><mml:mrow><mml:msup><mml:mi>L</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msup><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>/</mml:mo><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:msup><mml:mi>exp</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mi>t</mml:mi><mml:mo>/</mml:mo><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>sin</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi><mml:mi>t</mml:mi><mml:mo>/</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>and the non-linear filter by the polynomial</p>
<disp-formula id="E15"><label>(8)</label><mml:math id="M123"><mml:mrow><mml:msup><mml:mi>N</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msup><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msub><mml:msup><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:msub><mml:mi>B</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msub><mml:mi>x</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mtext>exc</mml:mtext></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
<p>Fits minimized the mean-square distance between model and data. <italic>L</italic><sup>inh</sup> and <italic>N</italic><sup>inh</sup> were fit using the same parametrization.</p>
<p>The noise terms &#x003B7;<sup>exc</sup><sub><italic>k</italic></sub>, &#x003B7;<sup>inh</sup><sub><italic>k</italic></sub> were fit to reproduce the statistical characteristics of the residuals from this fitting. We simulated the noise terms &#x003B7;<sup>exc</sup> and &#x003B7;<sup>inh</sup> using Ornstein&#x02013;Uhlenbeck processes with the appropriate parameters; these were entirely characterized by the mean, standard deviation, and time constant of autocorrelation &#x003C4;<sub>&#x003B7;,exc</sub> (&#x003C4;<sub>&#x003B7;, inh</sub>), as well as pairwise correlation coefficients for noise terms entering neighboring cells. The noise correlation coefficients were estimated from the dual recordings of Trong and Rieke (<xref ref-type="bibr" rid="B55">2008</xref>).</p>
<p>Linear filter parameters computed (also listed in Table <xref ref-type="table" rid="T1">1</xref>) were <italic>P</italic><sub>exc</sub> &#x0003D; &#x02212;8 &#x000D7; 10<sup>4</sup> s<sup>&#x02212;1</sup>, <italic>n</italic><sub>exc</sub> &#x0003D; 3.6, &#x003C4;<sub>exc</sub> &#x0003D; 12 ms, <italic>T</italic><sub>exc</sub> &#x0003D; 105 ms, and <italic>P</italic><sub>inh</sub> &#x0003D; &#x02212;1.8 &#x000D7; 10<sup>5</sup> s<sup>&#x02212;1</sup>, <italic>n</italic><sub>inh</sub> &#x0003D; 3.0, &#x003C4;<sub>inh</sub> &#x0003D; 16 ms, <italic>T</italic><sub>inh</sub> &#x0003D; 120 ms. Non-linearity parameters were <italic>A</italic><sub>exc</sub> &#x0003D; &#x02212;8.3 &#x000D7; 10<sup>&#x02212;7</sup> nS, <italic>B</italic><sub>exc</sub> &#x0003D; 7 &#x000D7; 10<sup>&#x02212;3</sup> nS, <italic>C</italic><sub>exc</sub> &#x0003D; &#x02212;0.95 nS, and <italic>A</italic><sub>inh</sub> &#x0003D; 1.67 &#x000D7; 10<sup>&#x02212;6</sup> nS, <italic>B</italic><sub>inh</sub> &#x0003D; 6.2 &#x000D7; 10<sup>&#x02212;3</sup> nS, <italic>C</italic><sub>inh</sub> &#x0003D; 4.17 nS. Noise parameters were measured to be mean(&#x003B7;<sup>exc</sup><sub><italic>k</italic></sub>) &#x0003D; 30, std(&#x003B7;<sup>exc</sup><sub><italic>k</italic></sub>) &#x0003D; 500, &#x003C4;<sub>&#x003B7;,exc</sub> &#x0003D; 22 ms, and mean(&#x003B7;<sup>inh</sup><sub><italic>k</italic></sub>) &#x0003D; &#x02212;1200, std (&#x003B7;<sup>inh</sup><sub><italic>k</italic></sub>) &#x0003D; 780, &#x003C4;<sub>&#x003B7;,inh</sub> &#x0003D; 33 ms. In addition, excitatory (inhibitory) noise to different cells &#x003B7;<sup>exc</sup><sub><italic>k</italic></sub>, &#x003B7;<sup>exc</sup><sub><italic>j</italic></sub> (&#x003B7;<sup>inh</sup><sub><italic>k</italic></sub>, &#x003B7;<sup>inh</sup><sub><italic>j</italic></sub>) had a correlation coefficient of 0.3 (0.15).</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p><bold>Parameters used to model the transformation of stimuli into synaptic conductances for the RGC model, as described in Equations (7&#x02013;9)</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left"><bold>Model (MOD)</bold></th>
<th align="left"><bold><italic>P</italic><sub>MOD</sub> (s<sup>&#x02212;1</sup>)</bold></th>
<th align="left"><bold>&#x003C4;<sub>MOD</sub> (ms)</bold></th>
<th align="left"><bold><italic>n</italic><sub>MOD</sub></bold></th>
<th align="left"><bold><italic>T</italic><sub>MOD</sub> (ms)</bold></th>
<th align="left"><bold><italic>A</italic><sub>MOD</sub> (nS)</bold></th>
<th align="left"><bold><italic>B</italic><sub>MOD</sub> (nS)</bold></th>
<th align="left"><bold><italic>C</italic><sub>MOD</sub> (nS)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">exc</td>
<td align="center">&#x02212;8 &#x000D7; 10<sup>4</sup></td>
<td align="center">12</td>
<td align="center">3.6</td>
<td align="center">105</td>
<td align="center">&#x02212;8.3 &#x000D7; 10<sup>&#x02212;7</sup></td>
<td align="center">7 &#x000D7; 10<sup>&#x02212;3</sup></td>
<td align="center">&#x02212;0.95</td>
</tr>
<tr>
<td align="left">inh</td>
<td align="center">&#x02212;1.8 &#x000D7; 10<sup>5</sup></td>
<td align="center">16</td>
<td align="center">3.0</td>
<td align="center">120</td>
<td align="center">1.67 &#x000D7; 10<sup>&#x02212;6</sup></td>
<td align="center">6.2 &#x000D7; 10<sup>&#x02212;3</sup></td>
<td align="center">4.17</td>
</tr>
<tr>
<td align="left">exc,M</td>
<td align="center">&#x02212;3.2 &#x000D7; 10<sup>5</sup></td>
<td align="center">12</td>
<td align="center">2</td>
<td align="center">120<sup>&#x0002A;</sup></td>
<td align="center">&#x02212;8.3 &#x000D7; 10<sup>&#x02212;7</sup></td>
<td align="center">7 &#x000D7; 10<sup>&#x02212;3</sup></td>
<td align="center">&#x02212;0.95</td>
</tr>
<tr>
<td align="left">inh,M</td>
<td align="center">&#x02212;3.5 &#x000D7; 10<sup>5</sup></td>
<td align="center">13.2</td>
<td align="center">2</td>
<td align="center">132<sup>&#x0002A;</sup></td>
<td align="center">1.67 &#x000D7; 10<sup>&#x02212;6</sup></td>
<td align="center">6.2 &#x000D7; 10<sup>&#x02212;3</sup></td>
<td align="center">4.17</td>
</tr>
<tr>
<td align="center" colspan="8"><bold>Additional parameters for monophasic filters</bold></td>
</tr>
<tr>
<td align="left" colspan="4"><bold>Model (MOD)</bold></td>
<td align="center"><bold><italic>T</italic><sub>MOD, S</sub> (ms)</bold></td>
<td align="center"><bold><italic>T</italic><sub>MOD, C</sub> (ms)</bold></td>
<td align="center" colspan="2"><bold><italic>R</italic><sub>MOD</sub></bold></td>
</tr>
<tr>
<td align="left" colspan="4">exc,M</td>
<td align="center">120</td>
<td align="center">100</td>
<td align="center" colspan="2">0.8</td>
</tr>
<tr>
<td align="left" colspan="4">inh,M</td>
<td align="center">132</td>
<td align="center">110</td>
<td align="center" colspan="2">0.8</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Asterisks (<sup>&#x0002A;</sup>) indicate parameters that are superceded by later rows; note that the monophasic filter equations contain two filtering timescales&#x02014;for example <italic>T</italic><sub>exc,M,S</sub> and <italic>T</italic><sub>exc,M,C</sub>, for the excitatory monophasic filter&#x02014;and a relative weighting (e.g., <italic>R</italic><sub>exc,M</sub>)</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>For the filter demonstrated in Figure <xref ref-type="fig" rid="F4">4</xref>, we added a cosine component to the previous filter, i.e.,</p>
<disp-formula id="E16"><label>(9)</label><mml:math id="M124"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msup><mml:mi>L</mml:mi><mml:mrow><mml:mtext>exc,M</mml:mtext></mml:mrow></mml:msup><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mtext>exc,M</mml:mtext></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>/</mml:mo><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mrow><mml:mtext>exc,M</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mtext>exc,M</mml:mtext></mml:mrow></mml:msub></mml:mrow></mml:msup><mml:mi>exp</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mi>t</mml:mi><mml:mo>/</mml:mo><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mrow><mml:mtext>exc,M</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>sin</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi><mml:mi>t</mml:mi><mml:mo>/</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mtext>exc,M,S</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mtext>exc,M</mml:mtext></mml:mrow></mml:msub><mml:mi>cos</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x003C0;</mml:mi><mml:mi>t</mml:mi><mml:mo>/</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mtext>exc,M,C</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Here <italic>P</italic><sub>exc,M</sub> &#x0003D; &#x02212;3.2 &#x000D7; 10<sup>5</sup> s<sup>&#x02212;1</sup>, <italic>n</italic><sub>exc,M</sub> &#x0003D; 2, &#x003C4;<sub>exc,M</sub> &#x0003D; 12 ms, <italic>T</italic><sub>exc,M,S</sub> &#x0003D; 120 ms and <italic>T</italic><sub>exc,M,C</sub> &#x0003D; 100 ms, and <italic>P</italic><sub>inh,M</sub> &#x0003D; &#x02212;3.5 &#x000D7; 10<sup>5</sup> s<sup>&#x02212;1</sup>, <italic>n</italic><sub>inh,M</sub> &#x0003D; 2, &#x003C4;<sub>inh,M</sub> &#x0003D; 13.2 ms, <italic>T</italic><sub>inh,M,S</sub> &#x0003D; 132 ms and <italic>T</italic><sub>inh,M,C</sub> &#x0003D; 110 ms, while <italic>R</italic><sub>exc,M</sub> &#x0003D; <italic>R</italic><sub>inh,M</sub> &#x0003D; 0.8.</p>
</sec>
<sec>
<title>4.1.2. Voltage evolution</title>
<p>We create a model of the cell as a non-linear integrate-and-fire model using the method of Badel et al. (<xref ref-type="bibr" rid="B4">2007</xref>), in which the membrane voltage is assumed to respond as</p>
<disp-formula id="E17"><label>(10)</label><mml:math id="M125"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>d</mml:mi><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>V</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mtext>last</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>input</mml:mtext></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mi>C</mml:mi></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>where <italic>C</italic> is the cell capacitance, <italic>t</italic><sub>last</sub> is the time of the last spike before time <italic>t</italic>, and <italic>I</italic><sub>input</sub>(<italic>t</italic>) is a time-dependent input current. We use the current-clamp data, which yields cell voltage in response to the input current <italic>I</italic><sub>input</sub>(<italic>t</italic>) &#x0003D; &#x02212;<italic>g</italic><sup>exc</sup>(<italic>t</italic>)(<italic>V</italic> &#x02212; <italic>V</italic><sub><italic>E</italic></sub>) &#x02212; <italic>g</italic><sup>inh</sup>(<italic>V</italic> &#x02212; <italic>V</italic><sub><italic>I</italic></sub>), to fit a function <italic>F</italic>(<italic>V, t</italic>). When voltage data is segregated according to the time since the last spike <italic>t</italic> &#x02212; <italic>t</italic><sub>last</sub>, the <italic>I</italic> &#x02212; <italic>V</italic> curve is well fit by a function of the form</p>
<disp-formula id="E18"><label>(11)</label><mml:math id="M126"><mml:mrow><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>V</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mtext>last</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow></mml:mfrac><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>E</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mi>V</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x00394;</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>V</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>T</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:msub><mml:mi>&#x00394;</mml:mi><mml:mi>T</mml:mi></mml:msub></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where parameters are the membrane time constant &#x003C4;<sub><italic>m</italic></sub>, resting potential (<italic>E</italic><sub><italic>L</italic></sub>), spike width &#x00394;<sub><italic>T</italic></sub> and knee of the exponential curve <italic>V</italic><sub><italic>T</italic></sub>.</p>
<p>The values of these constants differed in each bin of voltage data; to estimate these constants, we first extracted their values from each mean <italic>I</italic> &#x02212; <italic>V</italic> curve. We found that these constants, as a function of <italic>t</italic> &#x02212; <italic>t</italic><sub>last</sub>, were well fit by either a single exponential or a difference of two exponentials, with relaxation to a baseline rate (as in Badel et al., <xref ref-type="bibr" rid="B4">2007</xref>, Figure <xref ref-type="fig" rid="F3">3</xref>). Specifically, we chose:</p>
<disp-formula id="E19"><label>(12)</label><mml:math id="M127"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mi>m</mml:mi></mml:msub></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mtext>last</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:msub><mml:mi>&#x003C4;</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>E</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:msub><mml:mi>E</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:msub><mml:mi>E</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mtext>last</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:msub><mml:mi>E</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msup><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mtext>last</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:msub><mml:mi>E</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>&#x00394;</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:msub><mml:mi>&#x00394;</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:msub><mml:mi>&#x00394;</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mtext>last</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:msub><mml:mi>&#x00394;</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msup><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mtext>last</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:msub><mml:mi>&#x00394;</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>V</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mtext>last</mml:mtext></mml:mrow></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>We obtained the coefficients by least-squares fitting to the above functional forms: specifically, we found that (up to four digits): (<italic>c</italic><sub>&#x003C4;<sub><italic>m</italic></sub>,1</sub>, <italic>c</italic><sub>&#x003C4;<sub><italic>m</italic></sub>,2</sub>, <italic>c</italic><sub>&#x003C4;<sub><italic>m</italic></sub>,3</sub>) &#x0003D; (0.3719 ms<sup>&#x02212;1</sup>, 0.5412 ms<sup>&#x02212;1</sup>, 13.2726 ms), (<italic>c</italic><sub><italic>E</italic><sub><italic>L</italic></sub>,1</sub>, <italic>c</italic><sub><italic>E</italic><sub><italic>L</italic></sub>,2</sub>, <italic>c</italic><sub><italic>E</italic><sub><italic>L</italic></sub>,3</sub>, <italic>c</italic><sub><italic>E</italic><sub><italic>L</italic></sub>,4</sub>) &#x0003D; (&#x02212;59.4858 mV, 5.8966 mV, 8.3076 ms, 233.1114 ms), (<italic>c</italic><sub>&#x00394;<sub><italic>T</italic></sub>,1</sub>, <italic>c</italic><sub>&#x00394;<sub><italic>T</italic></sub>,2</sub>, <italic>c</italic><sub>&#x00394;<sub><italic>T</italic></sub>,3</sub>, <italic>c</italic><sub>&#x00394;<sub><italic>T</italic></sub>,4</sub>) &#x0003D; (20.0487 ms, 19.0560 ms, 3.6280 ms, 2.4304 s), and (<italic>c</italic><sub><italic>V</italic><sub><italic>T</italic></sub>,1</sub>, <italic>c</italic><sub><italic>V</italic><sub><italic>T</italic></sub>,2</sub>, <italic>c</italic><sub><italic>V</italic><sub><italic>T</italic></sub>,3</sub>) &#x0003D; (&#x02212;44.3323 mV, 25.1812 mV, 4.7653 ms). Coefficients are also listed in Table <xref ref-type="table" rid="T2">2</xref>.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p><bold>Coefficients used to define refractory EIF model as specified in Equations (11, 12)</bold>.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left"><bold>Parameter (PAR)</bold></th>
<th align="center"><bold><italic>c</italic><sub>PAR,1</sub></bold></th>
<th align="center"><bold><italic>c</italic><sub>PAR,2</sub></bold></th>
<th align="center"><bold><italic>c</italic><sub>PAR,3</sub> (ms)</bold></th>
<th align="center"><bold><italic>c</italic><sub>PAR,4</sub> (ms)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">&#x003C4;<sub><italic>m</italic></sub>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;(actual fit: 1/&#x003C4;<sub><italic>m</italic></sub>)</td>
<td align="center">0.3719 ms<sup>&#x02212;1</sup></td>
<td align="center">0.5412 ms<sup>&#x02212;1</sup></td>
<td align="center">13.2726</td>
<td/>
</tr>
<tr>
<td align="left"><italic>V</italic><sub><italic>T</italic></sub></td>
<td align="center">&#x02212;44.3323 mV</td>
<td align="center">25.1812 mV</td>
<td align="center">4.7653</td>
<td/>
</tr>
<tr>
<td align="left"><italic>E</italic><sub><italic>L</italic></sub></td>
<td align="center">&#x02212;59.4858 mV</td>
<td align="center">5.8966 mV</td>
<td align="center">8.3076</td>
<td align="center">233.1114</td>
</tr>
<tr>
<td align="left">&#x00394;<sub><italic>T</italic></sub></td>
<td align="center">20.0487 ms</td>
<td align="center">19.0560 ms</td>
<td align="center">3.6280</td>
<td align="center">2430.4</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The parameters <italic>1</italic>/&#x003C4;<sub><italic>m</italic></sub> and <italic>V</italic><sub><italic>T</italic></sub> were fit to single exponentials as functions of time, with three free parameters. The parameters <italic>E</italic><sub><italic>L</italic></sub> and &#x00394;<sub><italic>T</italic></sub> were fit to differences of exponentials and therefore have four parameters. Units in the first and second columns are as stated; coefficients in the third and fourth column are in units of milliseconds (ms).</italic></p>
</table-wrap-foot>
</table-wrap>
<p>The capacitance was inferred from the voltage trace data by finding, at a voltage value where the voltage/membrane current relationship is approximately Ohmic, the value of <italic>C</italic> that minimizes error in the relation Equation (10) (Badel et al., <xref ref-type="bibr" rid="B4">2007</xref>). The estimated value was <italic>C</italic> &#x0003D; 28 pF.</p>
</sec>
<sec>
<title>4.1.3. Spiking dynamics: feedforward network</title>
<p>For simulations without electronic coupling, our model neuron comprises Equations (10, 11) for <italic>V</italic> &#x0003C; <italic>V</italic><sub>threshold</sub>; a spike was detected when <italic>V</italic> reached <italic>V</italic><sub>threshold</sub> &#x0003D; &#x02212;30 mV; voltage was then reset to <italic>V</italic><sub>reset</sub> &#x0003D; &#x02212;55 mV. The cell was then unable to spike for an absolute refractory period of &#x003C4;<sub>abs</sub> &#x0003D; 3 ms.</p>
<p>All simulations presented here were done in a three-cell network.</p>
</sec>
<sec>
<title>4.1.4. Spiking dynamics: recurrent network</title>
<p>Gap junction coupling was introduced as an additional current on the right-hand side of Equation (10):</p>
<disp-formula id="E20"><label>(13)</label><mml:math id="M128"><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mtext>gap</mml:mtext><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mi>C</mml:mi></mml:mfrac><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mi>g</mml:mi><mml:mrow><mml:mtext>gap</mml:mtext></mml:mrow></mml:msup></mml:mrow><mml:mi>C</mml:mi></mml:mfrac><mml:mstyle displaystyle='true'><mml:munder><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:menclose notation='updiagonalstrike'><mml:mo>=</mml:mo></mml:menclose><mml:mi>j</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mstyle></mml:mrow></mml:math></disp-formula>
<p>The coupling strength <italic>g</italic><sup>gap</sup> was held constant during a simulation. When coupling was present (i.e., when <italic>g</italic><sup>gap</sup> &#x02260; 0), <italic>g</italic><sup>gap</sup> was varied from the measured level (1.1 nS) (Trong and Rieke, <xref ref-type="bibr" rid="B55">2008</xref>) to 16 times this value (17.6 nS) between simulations. When present, coupling was all-to-all.</p>
<p>As in the feedforward model, Equations (10, 11) were integrated for <italic>V</italic> &#x0003C; <italic>V</italic><sub>threshold</sub>, and a spike was detected when <italic>V</italic> reached <italic>V</italic><sub>threshold</sub> &#x0003D; &#x02212;30 mV. To model the voltage trajectory immediately following a spike, an averaged spike waveform was extracted from voltage traces of the same ON parasol cell used to fit Equations (10, 11). This spike waveform was then used to replace 1 ms of the membrane voltage trajectory during and after a spike; at the end of the 1 ms, the voltage was released at approximately &#x02212;58 mV. The cell was unable to spike for an absolute refractory period of &#x003C4;<sub>abs</sub> &#x0003D; 3 ms. A relative refractory period was induced by introducing a declining threshold for the period of 3&#x02013;6 ms following a spike, after which <italic>V</italic><sub>threshold</sub> returns to &#x02212;30 mV.</p>
</sec>
<sec>
<title>4.1.5. Cell receptive field and stimulation</title>
<p>We defined each cell&#x00027;s stimulus as the linear convolution of an image with its receptive field. The receptive fields include an ON center and an OFF surround, as in Chichilnisky and Kalmar (<xref ref-type="bibr" rid="B11">2002</xref>):</p>
<disp-formula id="E21"><label>(14)</label><mml:math id="M129"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>s</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mover accent='true'><mml:mi>x</mml:mi><mml:mo>&#x02192;</mml:mo></mml:mover><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>exp</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover accent='true'><mml:mi>x</mml:mi><mml:mo>&#x02192;</mml:mo></mml:mover><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mover accent='true'><mml:mi>x</mml:mi><mml:mo>&#x02192;</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mi>T</mml:mi></mml:msup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>Q</mml:mi></mml:mstyle><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover accent='true'><mml:mi>x</mml:mi><mml:mo>&#x02192;</mml:mo></mml:mover><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mover accent='true'><mml:mi>x</mml:mi><mml:mo>&#x02192;</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mi>k</mml:mi><mml:mi>exp</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover accent='true'><mml:mi>x</mml:mi><mml:mo>&#x02192;</mml:mo></mml:mover><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mover accent='true'><mml:mi>x</mml:mi><mml:mo>&#x02192;</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>Q</mml:mi></mml:mstyle><mml:mi>r</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mover accent='true'><mml:mi>x</mml:mi><mml:mo>&#x02192;</mml:mo></mml:mover><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mover accent='true'><mml:mi>x</mml:mi><mml:mo>&#x02192;</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where the parameters <italic>k</italic> and 1/<italic>r</italic> give the relative strength and size of the surround. <bold>Q</bold> specifies the shape of the center and was chosen to have a 1 standard deviation (SD) radius of 50 &#x003BC;m and to be perfectly circular. The receptive field locations <inline-formula><mml:math id="M130"><mml:mover accent='true'><mml:mi>x</mml:mi><mml:mo>&#x02192;</mml:mo></mml:mover></mml:math></inline-formula><sub>1</sub>, <inline-formula><mml:math id="M131"><mml:mover accent='true'><mml:mi>x</mml:mi><mml:mo>&#x02192;</mml:mo></mml:mover></mml:math></inline-formula><sub>2</sub>, and <inline-formula><mml:math id="M132"><mml:mover accent='true'><mml:mi>x</mml:mi><mml:mo>&#x02192;</mml:mo></mml:mover></mml:math></inline-formula><sub>3</sub> were chosen so that the 1 SD outlines of the receptive field centers will tile the plane (i.e., they just touch). Other parameters used were <italic>k</italic> &#x0003D; 0.3, <italic>r</italic> &#x0003D; 0.675.</p>
<p>Stimulation images were defined on a 512 &#x003BC;m &#x000D7; 512 &#x003BC;m grid that overlapped all three receptive fields. For full-field stimuli, light intensity was chosen be spatially constant and refreshed every 8, 40, or 100 ms by choosing independently from the specified stimulus distribution (Gaussian, binary, Cauchy, or heavy-tailed skew). For spatially variable stimuli, a checkerboard pattern was imposed on the stimulation image: the intensity value in each checkerboard square was chosen independently and refreshed at the appropriate interval. The checkerboard pattern was first given a random rotation and translation relative to the receptive fields: this was chosen at the outset of each batch of stixel simulations (for a total of five rotation/translation pairs per stixel size, refresh rate, and stimulus distribution). Two example placements are shown in Figures <xref ref-type="supplementary-material" rid="SM2">S2A,D</xref> for 256 &#x003BC;m and 60 &#x003BC;m pixels respectively.</p>
</sec>
<sec>
<title>4.1.6. Numerical methods</title>
<p>All simulations and data analysis were performed using MATLAB. Equations (10, 11) were integrated using the Euler method for &#x0003E;10<sup>5</sup> ms with a time step of 0.1 ms. The synaptic noise terms, &#x003B7;<sup>exc</sup><sub><italic>k</italic></sub> and &#x003B7;<sup>inh</sup><sub><italic>k</italic></sub>, as well as the light input, were generated independently for each simulation. In response to uniform light stimuli, firing rates were 11.51 &#x000B1; 0.38 Hz (standard deviations given across a total of 60 cells; 3 cells each from 20 10<sup>5</sup> ms simulations); 10 ms bins were used to discretize the spiking output. Firing rates were higher for full-field stimuli, ranging from 12 to 43 Hz (firing rates increased with stimulus variance); therefore shorter (5 ms) bins were used to discretize spike output for all other simulations. With this range of firing rates and bin size, multiple spikes were very rare (occurring in &#x0003C;1% of occupied bins). Empirical spiking distributions were computed from the binned spike data.</p>
<p>For each stimulus condition, 20 simulations (or sub-simulations) were run, for a total integration time of &#x0003E; 20 &#x000D7; 10<sup>5</sup> ms. These 20 sub-simulations were used to estimate standard errors in both the probability distribution over spiking events and <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M133"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>). Numbers reported in section 2 are, unless specified otherwise, produced by collating the data from the 20 simulations.</p>
<p>To fit a maximum entropy model <inline-formula><mml:math id="M134"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula> to an empirical probability distribution <italic>P</italic>, we used standard methods that have been explained elsewhere (Malouf, <xref ref-type="bibr" rid="B28">2002</xref>). Briefly, we minimized the negative log-likelihood function:</p>
<disp-formula id="E22"><label>(15)</label><mml:math id="M135"><mml:mrow><mml:mi>L</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x003BB;</mml:mi></mml:mstyle><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mstyle displaystyle='true'><mml:munder><mml:mo>&#x02211;</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>x</mml:mi></mml:mstyle></mml:munder><mml:mi>P</mml:mi></mml:mstyle><mml:mrow><mml:mo>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>x</mml:mi></mml:mstyle><mml:mo>)</mml:mo></mml:mrow><mml:mi>log</mml:mi><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>x</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x003BB;</mml:mi></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where</p>
<disp-formula id="E23"><mml:math id="M136"><mml:mrow><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>x</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x003BB;</mml:mi></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mi>Z</mml:mi><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x003BB;</mml:mi></mml:mstyle><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mi>exp</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mstyle displaystyle='true'><mml:munder><mml:mo>&#x02211;</mml:mo><mml:mi>k</mml:mi></mml:munder><mml:mrow><mml:msub><mml:mi>&#x003BB;</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:mstyle><mml:msub><mml:mi>f</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>x</mml:mi></mml:mstyle><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>;</mml:mo></mml:mrow></mml:math></disp-formula>
<p><italic>Z</italic><sub>&#x003BB;</sub> is the partition function, <italic>f</italic><sub><italic>k</italic></sub>, <italic>k</italic> &#x0003D; 1, &#x02026;, <italic>M</italic> is a set of functions or &#x0201C;features&#x0201D; of the spiking state, and <bold>&#x003BB;</bold> is a vector of parameters, each of which serves as a Lagrange multiplier enforcing the constraint <bold>E</bold><sub><inline-formula><mml:math id="M137"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula></sub>[<italic>f</italic><sub><italic>k</italic></sub>]. For the pairwise (PME) model on <italic>N</italic> cells, <bold>&#x003BB;</bold> corresponds to <italic>N</italic> firing rates and <italic>N</italic>(<italic>N</italic> &#x02212; 1)/2 covariances, and the sum is over all possible spiking states of the system. For <italic>N</italic> &#x0003D; 3 there are six such parameters, and</p>
<disp-formula id="E24"><mml:math id="M138"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mi>log</mml:mi><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x0007B;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>&#x0007D;</mml:mo><mml:mo>,</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x003BB;</mml:mi></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x003BB;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x003BB;</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x003BB;</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x003BB;</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x003BB;</mml:mi><mml:mrow><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x003BB;</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mi>log</mml:mi><mml:msub><mml:mi>Z</mml:mi><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x003BB;</mml:mi></mml:mstyle></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The function in Equation (15) is a convex function of the parameters <bold>&#x003BB;</bold> which will be minimized precisely (and uniquely) when <inline-formula><mml:math id="M139"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula> matches the desired moments from <italic>P</italic>: e.g., <bold>E</bold><sub><italic>P</italic></sub>[<italic>x</italic><sub>1</sub>] &#x0003D; <bold>E</bold><sub><inline-formula><mml:math id="M140"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula></sub>[<italic>x</italic><sub>1</sub>]. Since <inline-formula><mml:math id="M141"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula> is in log-linear form, the result will be the <italic>maximum entropy</italic> distribution that matches the desired moments (Malouf, <xref ref-type="bibr" rid="B28">2002</xref>). In principle any unconstrained gradient descent method may be used; we used an implementation of the non-linear conjugate gradient method. The Kullback Leibler divergence <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M142"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) was computed using the identity <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M143"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) &#x0003D; <italic>S</italic>(<inline-formula><mml:math id="M144"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) &#x02212; <italic>S</italic>(<italic>P</italic>), where <italic>S</italic>(<italic>P</italic>) is the entropy of <italic>P</italic>, i.e., <italic>S</italic>(<italic>P</italic>) &#x0003D; &#x02212;&#x02211;<sub><bold>x</bold></sub> <italic>P</italic>(<bold>x</bold>) log <italic>P</italic>(<bold>x</bold>).</p>
</sec>
<sec>
<title>4.1.7. Convergence testing</title>
<p>To test our finding that the observed distributions were well-modeled by the PME fit, we also performed the PME analysis on each of the 20 simulations for each stimulus condition. While in general <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M145"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) can be quite sensitive to perturbations in <italic>P</italic>, the numbers remained small under this analysis. To confirm that our results for <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M146"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) are sufficiently resolved to remove bias from sampling, we performed an analysis in which we collect the 20 simulations in subgroups of 1, 2, 4, 5, 10, and 20, and plot the mean <italic>D</italic><sub>KL</sub> with estimated standard errors. As expected (e.g., Paninski, <xref ref-type="bibr" rid="B39">2003</xref>), bias decreases as the length of subgroup increases and asymptotes at&#x02014;or before&#x02014;the full simulation length.</p>
<p>To provide a cross-validation test for the significance of our reported <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M147"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) values, we divided our data into halves (which we denote <italic>P</italic><sub>1</sub> and <italic>P</italic><sub>2</sub>, each including data from 10 sub-simulations) and performed the PME analysis on one half (say <italic>P</italic><sub>1</sub>) to yield a model <inline-formula><mml:math id="M148"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula><sub>1</sub>. We then computed <italic>D</italic><sub>KL</sub>(<italic>P</italic><sub>2</sub>, <inline-formula><mml:math id="M149"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula><sub>1</sub>) and <italic>D</italic><sub>KL</sub>(<italic>P</italic><sub>2</sub>, <italic>P</italic><sub>1</sub>) (as in Yu et al., <xref ref-type="bibr" rid="B59">2011</xref>), which we refer to the <italic>cross-validated</italic> and <italic>empirical</italic> likelihood, respectively. The former tests whether the PME fit is robust to over-fitting; the latter tests how well-resolved our &#x0201C;true&#x0201D; distribution is in the first place. Most cross-validated likelihoods fall on or near the identity line; most empirical likelihoods are close to zero [and importantly, significantly smaller than either <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M150"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) or <italic>D</italic><sub>KL</sub>(<italic>P</italic><sub>2</sub>, <inline-formula><mml:math id="M151"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula><sub>1</sub>), indicating that <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M152"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) is accurately resolved]. We conclude that the deviations that we observe when these conditions are met can not be accounted for by the differences in testing and training data.</p>
</sec>
</sec>
<sec>
<title>4.2. Computation of spiking patterns in the simplified model</title>
<p>As a simplified model of a neural circuit, we consider a variant of the <italic>Dichotomized Gaussian</italic> (Amari et al., <xref ref-type="bibr" rid="B2">2003</xref>; Macke et al., <xref ref-type="bibr" rid="B26">2009</xref>, <xref ref-type="bibr" rid="B27">2011</xref>), in which correlated inputs are thresholded to produce an output spike pattern. To be concrete, a set of <italic>N</italic> threshold spiking units is forced by a common input <italic>I</italic><sub><italic>c</italic></sub> [drawn from a probability distribution <italic>P</italic><sub><italic>C</italic></sub>(<italic>y</italic>)] and an independent input <italic>I</italic><sub><italic>j</italic></sub> [drawn from a probability distribution <italic>P</italic><sub><italic>I</italic></sub>(<italic>y</italic>)]. To relate these functions to the other free parameters in the model, <italic>P</italic><sub><italic>C</italic></sub>(<italic>y</italic>) and <italic>P</italic><sub><italic>I</italic></sub>(<italic>y</italic>) were always chosen so that <italic>I</italic><sub><italic>j</italic></sub> and <italic>I</italic><sub><italic>c</italic></sub> had mean 0 and variances (1 &#x02212; <italic>c</italic>) &#x003C3;<sup>2</sup> and <italic>c</italic> &#x003C3;<sup>2</sup>, respectively (so that <italic>c</italic> yields the Pearson&#x00027;s correlation coefficient of the input to two cells). The output of each cell <italic>x</italic><sub><italic>j</italic></sub> is determined by summing and thresholding these inputs:</p>
<disp-formula id="E25"><label>(16)</label><mml:math id="M153"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mi>&#x00398;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <italic>H</italic> is the Heaviside function [<italic>H</italic>(<italic>x</italic>) &#x0003D; 1 if <italic>x</italic> &#x02265; 0; <italic>H</italic>(<italic>x</italic>) &#x0003D; 0 otherwise]. Conditioned on <italic>I</italic><sub><italic>c</italic></sub>, the probability of each spike is given by:</p>
<disp-formula id="E26"><mml:math id="M154"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>a</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mi>&#x00398;</mml:mi><mml:mo>&#x0003E;</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x0003E;</mml:mo><mml:mi>&#x00398;</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle='true'><mml:mrow><mml:msubsup><mml:mo>&#x0222B;</mml:mo><mml:mrow><mml:mi>&#x00398;</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mi>&#x0221E;</mml:mi></mml:msubsup><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>I</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mstyle><mml:mo stretchy='false'>(</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mi>d</mml:mi><mml:mi>y</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Similarly, we have the conditioned probability that <italic>x</italic><sub><italic>j</italic></sub> &#x0003D; 0:</p>
<disp-formula id="E27"><mml:math id="M155"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>a</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mi>&#x00398;</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x0003C;</mml:mo><mml:mi>&#x00398;</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mi>a</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle='true'><mml:mrow><mml:msubsup><mml:mo>&#x0222B;</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mi>&#x0221E;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x00398;</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mi>a</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>I</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mstyle><mml:mo stretchy='false'>(</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mi>d</mml:mi><mml:mi>y</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Because these are conditionally independent, the probability of any spiking event (<italic>x</italic><sub>1</sub>, <italic>x</italic><sub>2</sub>, &#x02026;, <italic>x</italic><sub><italic>N</italic></sub>) &#x0003D; (<italic>A</italic><sub>1</sub>, <italic>A</italic><sub>2</sub>, &#x02026;, <italic>A</italic><sub><italic>N</italic></sub>) is given by the integral of the product of the conditioned probabilities against the density of the common input.</p>
<disp-formula id="E28"><label>(17)</label><mml:math id="M156"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>N</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mi>N</mml:mi></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle='true'><mml:mrow><mml:msubsup><mml:mo>&#x0222B;</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mi>&#x0221E;</mml:mi></mml:mrow><mml:mi>&#x0221E;</mml:mi></mml:msubsup><mml:mi>d</mml:mi></mml:mrow></mml:mstyle><mml:mi>y</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>C</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x0220F;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi></mml:mstyle></mml:mrow></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x02223;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The integral in Equation (17) is numerically evaluated via an adaptive quadrature routine, such at Matlab&#x00027;s <monospace>quad</monospace> or <monospace>integral</monospace>.</p>
<p>Four distinct unimodal inputs were used; two with heavy tails (Cauchy and heavy-tailed with skew), and two with sub-Gaussian tails (Gaussian and skewed). A random variable <italic>X</italic> is <italic>sub-Gaussian</italic> if the probability of large events can be bounded above by a scaled Gaussian; that is, if there exist constants <italic>C</italic>, <italic>c</italic> &#x0003E; 0 such that</p>
<disp-formula id="E29"><mml:math id="M157"><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x0007C;</mml:mo><mml:mi>X</mml:mi><mml:mo>&#x0007C;</mml:mo><mml:mo>&#x0003E;</mml:mo><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02264;</mml:mo><mml:mi>C</mml:mi><mml:mi>exp</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mi>c</mml:mi><mml:msup><mml:mi>&#x003BB;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>for all &#x003BB; (e.g., see Tao, <xref ref-type="bibr" rid="B51">2012</xref>, p. 15).</p>
<p>Unimodal inputs <italic>I</italic><sub><italic>j</italic></sub>, <italic>I</italic><sub><italic>c</italic></sub> were chosen from different marginals with mean 0 and variances (1 &#x02212; <italic>c</italic>) &#x003C3;<sup>2</sup>, <italic>c</italic> &#x003C3;<sup>2</sup>, respectively (for simplicity, we use &#x003C3;<sup>2</sup> to refer to the variance of a generic probability distribution in the following three paragraphs). For Gaussian inputs with variance &#x003C3;<sup>2</sup>, <italic>P</italic>(<italic>x</italic>) &#x0221D; <italic>e</italic><sup>&#x02212;<italic>x</italic><sup>2</sup>/2&#x003C3;<sup>2</sup></sup>; for skewed inputs, <italic>P</italic>(<italic>x</italic>) &#x0221D; (<italic>x</italic> &#x0002B; &#x003BC;)<italic>e</italic><sup>&#x02212;(<italic>x</italic> &#x0002B; &#x003BC;)<sup>2</sup>/2<italic>a</italic></sup>, for <italic>x</italic> &#x0003E; &#x02212;&#x003BC;, where the parameter <italic>a</italic> sets the variance <inline-formula><mml:math id="M158"><mml:mrow><mml:mn>2</mml:mn><mml:mi>a</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mi>&#x003C0;</mml:mi><mml:mn>4</mml:mn></mml:mfrac><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:math></inline-formula> and shifting by <inline-formula><mml:math id="M159"><mml:mrow><mml:mi>&#x003BC;</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mfrac><mml:mrow><mml:mi>a</mml:mi><mml:mi>&#x003C0;</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:mfrac></mml:mrow></mml:msqrt></mml:mrow></mml:math></inline-formula> ensures that the mean of <italic>P</italic>(<italic>x</italic>) is zero.</p>
<p>The heavy-tailed unimodal inputs were chosen so that the rate of tail decay would mimic the <italic>I</italic><sup>&#x02212;2</sup> luminance statistics found in natural scenes (Ruderman and Bialek, <xref ref-type="bibr" rid="B42">1994</xref>):</p>
<disp-formula id="E30"><mml:math id="M160"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x0221D;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x02003;&#x02003;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mi>X</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mi>X</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x0221D;</mml:mo><mml:mfrac><mml:mi>x</mml:mi><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>3</mml:mn><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x02003;&#x02003;</mml:mtext><mml:mn>0</mml:mn><mml:mo>&#x02264;</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mi>X</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Bimodal inputs with variance &#x003C3;<sup>2</sup> were chosen in the following way: in all cases, <italic>P</italic>(<italic>x</italic>) was chosen to be a discrete distribution with support on two values {0, X} i.e., <italic>P</italic>(<italic>X</italic>) &#x0003D; <italic>p</italic> and <italic>P</italic>(0) &#x0003D; 1 &#x02212; <italic>p</italic>. If possible (i.e., if &#x003C3;<sup>2</sup> &#x02264; 1/4), <italic>X</italic> was chosen to be 1; otherwise, <italic>X</italic> was chosen so as to minimize the distance between 0 and <italic>X</italic>. Finally, <italic>P</italic>(<italic>x</italic>) was shifted to have the desired mean value.</p>
</sec>
<sec>
<title>Conflict of interest statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</sec>
</body>
<back>
<ack>
<p>This research was supported by NSF grants DMS-0817649 and 1056125 by a Career Award at the Scientific Interface from the Burroughs-Wellcome Fund (Eric Shea-Brown), by the Howard Hughes Medical Institute and by NIH grant EY-11850 (Fred Rieke), by a Trinity College Research Studentship (Julijana Gjorgjieva), and by an Early Career Award from the Mathematical Biosciences Institute (Andrea K. Barreiro).</p>
</ack>
<sec sec-type="supplementary material" id="s5">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="http://www.frontiersin.org/journal/10.3389/fncom.2014.00010/abstract">http://www.frontiersin.org/journal/10.3389/fncom.2014.00010/abstract</ext-link></p>
<supplementary-material xlink:href="Presentation1.ZIP" id="SM1" mimetype="application/zip" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Figure S1</label>
<caption><p><bold>Biphasic vs. monophasic filters used in simulations illustrated in Figure <xref ref-type="fig" rid="F4">4</xref></bold>.</p></caption>
</supplementary-material>
<supplementary-material xlink:href="Presentation1.ZIP" id="SM2" mimetype="application/zip" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Figure S2</label>
<caption><p><bold>Illustration of RGC simulations with light stimuli of varying spatial scale (&#x0201C;stixels&#x0201D;)</bold>. <bold>(A&#x02013;C)</bold> For stixel size 60 &#x003BC;m, results for one randomly chosen stimulus position. <bold>(A)</bold> Contour lines of the three receptive fields (at 0.5, 1, 1.5, and 2 SD; and at the zero contour line) superimposed on the stimulus checkerboard (for illustration, pictured in an alternating black/white pattern). The red scale bar indicates 100 &#x003BC;m. <bold>(B)</bold> Histograms of the excitatory conductances, for each cell. <bold>(C)</bold> Spike pattern distribution, as obtained from computational simulations of the RGC model (&#x0201C;Observed&#x0201D;; dark blue), and the corresponding pairwise fit (&#x0201C;PME&#x0201D;; light pink). All eight spike patterns are shown, to allow for the possibility of non-symmetric responses; the three different probabilities labeled <italic>p</italic><sub>1</sub> correspond to <italic>P</italic>[(1, 0, 0)], <italic>P</italic>[(0, 1, 0)], and <italic>P</italic>[(0, 0, 1)]. <bold>(D&#x02013;F)</bold> As in <bold>(A&#x02013;C)</bold>, but for stixel size 256 &#x003BC;m. Panels <bold>(E,F)</bold> demonstrate that for this input, both excitatory inputs and spiking responses are heterogenous across the RGCs.</p></caption>
</supplementary-material>
<supplementary-material xlink:href="Presentation1.ZIP" id="SM3" mimetype="application/zip" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Figure S3</label>
<caption><p><bold>Strength of higher-order interactions produced by the threshold model as input parameters vary; relationship with other output firing statistics</bold>. <bold>(A)</bold> For skewed common inputs: <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M161"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) as a function of input correlation <italic>c</italic> and input standard deviation &#x003C3;, for a fixed threshold &#x00398; &#x0003D; 1.5. Color indicates <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M162"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>); see color bar for range. <bold>(B)</bold> For skewed common inputs: <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M163"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) vs. firing rate <bold>E</bold>[<italic>x</italic><sub>1</sub>] (Left) and the fraction of multi-information (&#x00394;) captured by the PME model vs. firing rate <bold>E</bold>[<italic>x</italic><sub>1</sub>] (Right). In <bold>(B)</bold>, possible input parameters were varied over a broad range as described in section 2. Firing rate is defined as the probability of a spike occurring per cell per random draw of the sum-and-threshold model, as defined in Equation (16). Color indicates output correlation coefficient &#x003C1; ranging from black for &#x003C1; &#x02208; (0, 0.1), to white for &#x003C1; &#x02208; (0.9, 1), as illustrated in the color bars. <bold>(C,D)</bold>: as in <bold>(A,B)</bold>, but for heavy-tailed, skewed common inputs.</p></caption>
</supplementary-material>
<supplementary-material xlink:href="Presentation1.ZIP" id="SM4" mimetype="application/zip" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Figure S4</label>
<caption><p><bold>The range of higher-order interactions produced by the threshold model varies across input type</bold>. Here, all values of <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M164"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) produced by the three-cell threshold model (previously displayed in Figures <xref ref-type="fig" rid="F7">7</xref>, <xref ref-type="supplementary-material" rid="SM3">S3</xref>) are superimposed to show the contrast between different input distributions. By comparing these data with data from direct sampling of all symmetric spiking distributions on three cells (from Figure <xref ref-type="fig" rid="F1">1</xref> and shown here in yellow), one can see that only a limited set of output patterns are accessed by the feedforward thresholding model. Firing rate is defined as the probability of a spike occurring per cell per random draw of the sum-and-threshold model, as defined in Equation (16).</p></caption>
</supplementary-material>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Amari</surname> <given-names>S.</given-names></name></person-group> (<year>2001</year>). <article-title>Information geometry on hierarchy of probability distributions</article-title>. <source>IEEE Trans. Inf. Theory</source> <volume>47</volume>, <fpage>1701</fpage>&#x02013;<lpage>1711</lpage>. <pub-id pub-id-type="doi">10.1109/18.930911</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Amari</surname> <given-names>S.</given-names></name> <name><surname>Nakahara</surname> <given-names>H.</given-names></name> <name><surname>Wu</surname> <given-names>S.</given-names></name> <name><surname>Sakai</surname> <given-names>Y.</given-names></name></person-group> (<year>2003</year>). <article-title>Synchronous firing and higher-order interactions in neuron pool</article-title>. <source>Neur. Comp</source>. <volume>15</volume>, <fpage>127</fpage>&#x02013;<lpage>142</lpage>. <pub-id pub-id-type="doi">10.1162/089976603321043720</pub-id><pub-id pub-id-type="pmid">12590822</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Badel</surname> <given-names>L.</given-names></name> <name><surname>Lefort</surname> <given-names>S.</given-names></name> <name><surname>Berger</surname> <given-names>T. K.</given-names></name> <name><surname>Petersen</surname> <given-names>C. C. H.</given-names></name> <name><surname>Gerstner</surname> <given-names>W.</given-names></name> <name><surname>Richardson</surname> <given-names>M. J. E.</given-names></name></person-group> (<year>2008</year>). <article-title>Extracting non-linear integrate-and-fire models from experimental data using dynamic I&#x02013;V curves</article-title>. <source>Biol. Cybern</source>. <volume>99</volume>, <fpage>361</fpage>&#x02013;<lpage>370</lpage>. <pub-id pub-id-type="doi">10.1007/s00422-008-0259-4</pub-id><pub-id pub-id-type="pmid">19011924</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Badel</surname> <given-names>L.</given-names></name> <name><surname>Lefort</surname> <given-names>S.</given-names></name> <name><surname>Brette</surname> <given-names>R.</given-names></name> <name><surname>Petersen</surname> <given-names>C. C. H.</given-names></name> <name><surname>Gerstner</surname> <given-names>W.</given-names></name> <name><surname>Richardson</surname> <given-names>M. J. E.</given-names></name></person-group> (<year>2007</year>). <article-title>Dynamic I-V curves are reliable predictors of naturalistic pyramidal-neuron voltage traces</article-title>. <source>J. Neurophys</source>. <volume>99</volume>, <fpage>656</fpage>&#x02013;<lpage>666</lpage>. <pub-id pub-id-type="doi">10.1152/jn.01107.2007</pub-id><pub-id pub-id-type="pmid">18057107</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barreiro</surname> <given-names>A. K.</given-names></name> <name><surname>Shea-Brown</surname> <given-names>E. T.</given-names></name> <name><surname>Thilo</surname> <given-names>E. L.</given-names></name></person-group> (<year>2010</year>). <article-title>Timescales of spike-train correlation for neural oscillators with common drive</article-title>. <source>Phys. Rev. E</source> <volume>81</volume>, <fpage>011916</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevE.81.011916</pub-id><pub-id pub-id-type="pmid">20365408</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barreiro</surname> <given-names>A. K.</given-names></name> <name><surname>Thilo</surname> <given-names>E. L.</given-names></name> <name><surname>Shea-Brown</surname> <given-names>E. T.</given-names></name></person-group> (<year>2012</year>). <article-title>A-current and type I / type II transition determine collective spiking from common input</article-title>. <source>J. Neurophysiol</source>. <volume>108</volume>, <fpage>1631</fpage>&#x02013;<lpage>1645</lpage>. <pub-id pub-id-type="doi">10.1152/jn.00928.2011</pub-id><pub-id pub-id-type="pmid">22673330</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="editor"><name><surname>Baudry</surname> <given-names>M.</given-names></name> <name><surname>Taketani</surname> <given-names>M.</given-names></name></person-group> (eds). (<year>2006</year>). <source>Advances in Network Electrophysiology Using Multi-Electrode Arrays</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer Press</publisher-name>. <pub-id pub-id-type="doi">10.1007/0-387-25858-2_15</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bethge</surname> <given-names>M.</given-names></name> <name><surname>Berens</surname> <given-names>P.</given-names></name></person-group> (<year>2008</year>). <article-title>Near-maximum entropy models for binary neural representations of natural images</article-title>. <source>Adv. Neur. Inf. Proc. Syst</source>. <volume>20</volume>, <fpage>97</fpage>&#x02013;<lpage>104</lpage>.</citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bohte</surname> <given-names>S. M.</given-names></name> <name><surname>Spekreijse</surname> <given-names>H.</given-names></name> <name><surname>Roelfsema</surname> <given-names>P. R.</given-names></name></person-group> (<year>2000</year>). <article-title>The effects of pair-wise and higher-order correlations on the firing rate of a postsynaptic neuron</article-title>. <source>Neur. Comp</source>. <volume>12</volume>, <fpage>153</fpage>&#x02013;<lpage>179</lpage>. <pub-id pub-id-type="doi">10.1162/089976600300015934</pub-id><pub-id pub-id-type="pmid">10636937</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cain</surname> <given-names>N.</given-names></name> <name><surname>Shea-Brown</surname> <given-names>E.</given-names></name></person-group> (<year>2013</year>). <article-title>Impact of correlated neural activity on decision making performance</article-title>. <source>Neur. Comp</source>. <volume>25</volume>, <fpage>289</fpage>&#x02013;<lpage>327</lpage>. <pub-id pub-id-type="doi">10.1162/NECO_a_00398</pub-id><pub-id pub-id-type="pmid">23148409</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chichilnisky</surname> <given-names>E. J.</given-names></name> <name><surname>Kalmar</surname> <given-names>R. S.</given-names></name></person-group> (<year>2002</year>). <article-title>Functional asymmetries in ON and OFF ganglion cells of primate retina</article-title>. <source>J. Neurosci</source>. <volume>22</volume>, <fpage>2737</fpage>&#x02013;<lpage>2747</lpage>. <pub-id pub-id-type="pmid">11923439</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cocco</surname> <given-names>S.</given-names></name> <name><surname>Leibler</surname> <given-names>S.</given-names></name> <name><surname>Monasson</surname> <given-names>R.</given-names></name></person-group> (<year>2009</year>). <article-title>Neuronal couplings between retinal ganglion cells inferred by efficient inverse statistical physics methods</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A</source>. <volume>106</volume>, <fpage>14058</fpage>&#x02013;<lpage>14062</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0906705106</pub-id><pub-id pub-id-type="pmid">19666487</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cover</surname> <given-names>T. M.</given-names></name> <name><surname>Thomas</surname> <given-names>J. A.</given-names></name></person-group> (<year>1991</year>). <source>Elements of Information Theory</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>Wiley</publisher-name>. <pub-id pub-id-type="doi">10.1002/0471200611</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dacey</surname> <given-names>D.</given-names></name> <name><surname>Brace</surname> <given-names>S.</given-names></name></person-group> (<year>1992</year>). <article-title>A coupled network for parasol but not midget ganglion cells in the primate retina</article-title>. <source>Vis. Neurosci</source>. <volume>9</volume>, <fpage>279</fpage>&#x02013;<lpage>290</lpage>. <pub-id pub-id-type="doi">10.1017/S0952523800010695</pub-id><pub-id pub-id-type="pmid">1390387</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dayan</surname> <given-names>P.</given-names></name> <name><surname>Abbot</surname> <given-names>L.</given-names></name></person-group> (<year>2001</year>). <source>Theoretical Neuroscience: Computational and Mathematical Modeling of Neural Systems</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>. <pub-id pub-id-type="doi">10.1016/S0306-4522(00)00552-2</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>de la Rocha</surname> <given-names>J.</given-names></name> <name><surname>Doiron</surname> <given-names>B.</given-names></name> <name><surname>Shea-Brown</surname> <given-names>E.</given-names></name> <name><surname>Josic</surname> <given-names>K.</given-names></name> <name><surname>Reyes</surname> <given-names>A.</given-names></name></person-group> (<year>2007</year>). <article-title>Correlation between neural spike trains increases with firing rate</article-title>. <source>Nature</source> <volume>448</volume>, <fpage>802</fpage>&#x02013;<lpage>806</lpage>. <pub-id pub-id-type="doi">10.1038/nature06028</pub-id><pub-id pub-id-type="pmid">17700699</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fairhall</surname> <given-names>A.</given-names></name> <name><surname>Burlingame</surname> <given-names>C.</given-names></name> <name><surname>Narasimhan</surname> <given-names>R.</given-names></name> <name><surname>Harris</surname> <given-names>R.</given-names></name> <name><surname>Puchalla</surname> <given-names>K.</given-names></name> <name><surname>Berry</surname> <given-names>M.</given-names></name></person-group> (<year>2006</year>). <article-title>Selectivity for multiple stimulus features in retinal ganglion cells</article-title>. <source>J. Neurophys</source>. <volume>96</volume>, <fpage>2724</fpage>&#x02013;<lpage>2738</lpage>. <pub-id pub-id-type="doi">10.1152/jn.00995.2005</pub-id><pub-id pub-id-type="pmid">16914609</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ganmor</surname> <given-names>E.</given-names></name> <name><surname>Segev</surname> <given-names>R.</given-names></name> <name><surname>Schneidman</surname> <given-names>E.</given-names></name></person-group> (<year>2011</year>). <article-title>Sparse low-order interaction network underlies a highly correlated and learnable population code</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A</source>. <volume>108</volume>, <fpage>9679</fpage>&#x02013;<lpage>9684</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1019641108</pub-id><pub-id pub-id-type="pmid">21602497</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hong</surname> <given-names>S.</given-names></name> <name><surname>Ratte</surname> <given-names>S.</given-names></name> <name><surname>Prescott</surname> <given-names>S.</given-names></name> <name><surname>De Schutter</surname> <given-names>E.</given-names></name></person-group> (<year>2012</year>). <article-title>Single neuron firing properties impact correlation-based population coding</article-title>. <source>J. Neurosci</source>. <volume>32</volume>, <fpage>1413</fpage>&#x02013;<lpage>1428</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.3735-11.2012</pub-id><pub-id pub-id-type="pmid">22279226</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jaynes</surname> <given-names>E. T.</given-names></name></person-group> (<year>1957a</year>). <article-title>Information theory and statistical mechanics</article-title>. <source>Physiol. Rev</source>. <volume>106</volume>, <fpage>620</fpage>&#x02013;<lpage>630</lpage>. <pub-id pub-id-type="doi">10.1103/PhysRev.106.620</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jaynes</surname> <given-names>E. T.</given-names></name></person-group> (<year>1957b</year>). <article-title>Information theory and statistical mechanics II</article-title>. <source>Physiol. Rev</source>. <volume>108</volume>, <fpage>171</fpage>&#x02013;<lpage>190</lpage>. <pub-id pub-id-type="doi">10.1103/PhysRev.108.171</pub-id><pub-id pub-id-type="pmid">21928963</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Koster</surname> <given-names>U.</given-names></name> <name><surname>Sohl-Dickstein</surname> <given-names>J.</given-names></name> <name><surname>Gray</surname> <given-names>C. M.</given-names></name> <name><surname>Olshausen</surname> <given-names>B. A.</given-names></name></person-group> (<year>2013</year>). <article-title>Higher order correlations within cortical layers dominate functional connectivity in microcolumns</article-title>. ArXiv q-Bio/1301.0050.</citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krumin</surname> <given-names>M.</given-names></name> <name><surname>Shoham</surname> <given-names>S.</given-names></name></person-group> (<year>2009</year>). <article-title>Generation of spike trains with controlled auto- and cross-correlation functions</article-title>. <source>Neur. Comp</source>. <volume>21</volume>, <fpage>1642</fpage>&#x02013;<lpage>1664</lpage>. <pub-id pub-id-type="doi">10.1162/neco.2009.08-08-847</pub-id><pub-id pub-id-type="pmid">19191596</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kuhn</surname> <given-names>A.</given-names></name> <name><surname>Aertsen</surname> <given-names>A.</given-names></name> <name><surname>Rotter</surname> <given-names>S.</given-names></name></person-group> (<year>2003</year>). <article-title>Higher-order statistics of input ensembles and the response of simple model neurons</article-title>. <source>Neur. Comp</source>. <volume>15</volume>, <fpage>67</fpage>&#x02013;<lpage>101</lpage>. <pub-id pub-id-type="doi">10.1162/089976603321043702</pub-id><pub-id pub-id-type="pmid">12590820</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Leen</surname> <given-names>D.</given-names></name> <name><surname>Shea-Brown</surname> <given-names>E.</given-names></name></person-group> (<year>2013</year>). <article-title>A simple mechanism for higher-order correlations in integrate-and-fire neurons</article-title>. ArXiv q-Bio.NC/1306.5275.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Macke</surname> <given-names>J. H.</given-names></name> <name><surname>Berens</surname> <given-names>P.</given-names></name> <name><surname>Ecker</surname> <given-names>A. S.</given-names></name> <name><surname>Tolias</surname> <given-names>A. S.</given-names></name> <name><surname>Bethge</surname> <given-names>M.</given-names></name></person-group> (<year>2009</year>). <article-title>Generating spike trains with specified correlation coefficients</article-title>. <source>Neur. Comp</source>. <volume>21</volume>, <fpage>397</fpage>&#x02013;<lpage>423</lpage>. <pub-id pub-id-type="doi">10.1162/neco.2008.02-08-713</pub-id><pub-id pub-id-type="pmid">19196233</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Macke</surname> <given-names>J. H.</given-names></name> <name><surname>Opper</surname> <given-names>M.</given-names></name> <name><surname>Bethge</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <article-title>Common input explains higher-order correlations and entropy in a simple model of neural population activity</article-title>. <source>Phys. Rev. Lett</source>. <volume>106</volume>, <fpage>208102</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevLett.106.208102</pub-id><pub-id pub-id-type="pmid">21668265</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Malouf</surname> <given-names>R.</given-names></name></person-group> (<year>2002</year>). <article-title>A comparison of algorithms for maximum entropy parameter estimation</article-title>, in <source>Proceedings of the Sixth Conference on Natural Language Learning</source> (<publisher-loc>Stroudsburg, PA</publisher-loc>), <fpage>49</fpage>&#x02013;<lpage>55</lpage>. <pub-id pub-id-type="doi">10.3115/1118853.1118871</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marella</surname> <given-names>S.</given-names></name> <name><surname>Ermentrout</surname> <given-names>G. B.</given-names></name></person-group> (<year>2008</year>). <article-title>Class-II neurons display a higher degree of stochastic synchronization than class-I neurons</article-title>. <source>Phys. Rev. E</source> <volume>77</volume>, <fpage>041908</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevE.77.041918</pub-id><pub-id pub-id-type="pmid">18517667</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Martignon</surname> <given-names>L.</given-names></name> <name><surname>Deco</surname> <given-names>G.</given-names></name> <name><surname>Laskey</surname> <given-names>K.</given-names></name> <name><surname>Diamond</surname> <given-names>M.</given-names></name> <name><surname>Freiwald</surname> <given-names>W.</given-names></name> <name><surname>Vaadia</surname> <given-names>E.</given-names></name></person-group> (<year>2000</year>). <article-title>Neural coding: higher-order temporal patterns in the neurostatistics of cell assemblies</article-title>. <source>Neur. Comp</source>. <volume>12</volume>, <fpage>2621</fpage>&#x02013;<lpage>2653</lpage>. <pub-id pub-id-type="doi">10.1162/089976600300014872</pub-id><pub-id pub-id-type="pmid">11110130</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McCulloch</surname> <given-names>W. S.</given-names></name> <name><surname>Pitts</surname> <given-names>W.</given-names></name></person-group> (<year>1943</year>). <article-title>A logical calculus of the ideas immanent in nervous activity</article-title>. <source>Bull. Math. Biophys</source>. <volume>5</volume>, <fpage>115</fpage>&#x02013;<lpage>137</lpage>. <pub-id pub-id-type="doi">10.1007/BF02478259</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Montani</surname> <given-names>F.</given-names></name> <name><surname>Ince</surname> <given-names>R. A. A.</given-names></name> <name><surname>Senatore</surname> <given-names>R.</given-names></name> <name><surname>Arabzadeh</surname> <given-names>E.</given-names></name> <name><surname>Diamond</surname> <given-names>M. E.</given-names></name> <name><surname>Panzeri</surname> <given-names>S.</given-names></name></person-group> (<year>2009</year>). <article-title>The impact of high-order interactions on the rate of synchronous discharge and information transmission in somatosensory cortex</article-title>. <source>Phil. Trans. R. Soc. A</source> <volume>367</volume>, <fpage>3297</fpage>&#x02013;<lpage>3310</lpage>. <pub-id pub-id-type="doi">10.1098/rsta.2009.0082</pub-id><pub-id pub-id-type="pmid">19620125</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moreno</surname> <given-names>R.</given-names></name> <name><surname>de la Rocha</surname> <given-names>J.</given-names></name> <name><surname>Renart</surname> <given-names>A.</given-names></name> <name><surname>Parga</surname> <given-names>N.</given-names></name></person-group> (<year>2002</year>). <article-title>Response of spiking neurons to correlated inputs</article-title>. <source>Phys. Rev. Lett</source>. <volume>89</volume>, <fpage>288101</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevLett.89.288101</pub-id><pub-id pub-id-type="pmid">12513181</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Murphy</surname> <given-names>G. J.</given-names></name> <name><surname>Rieke</surname> <given-names>F.</given-names></name></person-group> (<year>2006</year>). <article-title>Network variability limits stimulus-evoked spike timing precision in retinal ganglion cells</article-title>. <source>Neuron</source> <volume>52</volume>, <fpage>511</fpage>&#x02013;<lpage>524</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuron.2006.09.014</pub-id><pub-id pub-id-type="pmid">17088216</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nowotny</surname> <given-names>T.</given-names></name> <name><surname>Huerta</surname> <given-names>R.</given-names></name></person-group> (<year>2003</year>). <article-title>Explaining synchrony in feed-forward networks: are McCulloch-Pitts neurons good enough?</article-title> <source>Biol. Cybern</source>. <volume>89</volume>, <fpage>237</fpage>&#x02013;<lpage>241</lpage>. <pub-id pub-id-type="doi">10.1007/s00422-003-0431-9</pub-id><pub-id pub-id-type="pmid">14605888</pub-id></citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ohiorhenuan</surname> <given-names>I. E.</given-names></name> <name><surname>Mechler</surname> <given-names>F.</given-names></name> <name><surname>Purpura</surname> <given-names>K. P.</given-names></name> <name><surname>Schmid</surname> <given-names>A. M.</given-names></name> <name><surname>Hu</surname> <given-names>Q.</given-names></name> <name><surname>Victor</surname> <given-names>J. D.</given-names></name></person-group> (<year>2010</year>). <article-title>Sparse coding and high-order correlations in fine-scale cortical networks</article-title>. <source>Nature</source> <volume>466</volume>, <fpage>617</fpage>&#x02013;<lpage>621</lpage>. <pub-id pub-id-type="doi">10.1038/nature09178</pub-id><pub-id pub-id-type="pmid">20601940</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ohiorhenuan</surname> <given-names>I. E.</given-names></name> <name><surname>Victor</surname> <given-names>J. D.</given-names></name></person-group> (<year>2010</year>). <article-title>Information-geometric measure of 3-neuron firing patterns characterizes scale-dependence in cortical networks</article-title>. <source>J. Comp. Neurosci</source>. <volume>30</volume>, <fpage>125</fpage>&#x02013;<lpage>141</lpage>. <pub-id pub-id-type="doi">10.1007/s10827-010-0257-0</pub-id><pub-id pub-id-type="pmid">20635129</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Oizumi</surname> <given-names>M.</given-names></name> <name><surname>Ishii</surname> <given-names>T.</given-names></name> <name><surname>Ishibashi</surname> <given-names>K.</given-names></name> <name><surname>Okada</surname> <given-names>M.</given-names></name></person-group> (<year>2010</year>). <article-title>Mismatched decoding in the brain</article-title>. <source>J. Neurosci</source>. <volume>30</volume>, <fpage>4815</fpage>&#x02013;<lpage>4826</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.4360-09.2010</pub-id><pub-id pub-id-type="pmid">20357132</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Paninski</surname> <given-names>L.</given-names></name></person-group> (<year>2003</year>). <article-title>Estimation of entropy and mutual information</article-title>. <source>Neur. Comp</source>. <volume>15</volume>, <fpage>1191</fpage>&#x02013;<lpage>1253</lpage>. <pub-id pub-id-type="doi">10.1162/089976603321780272</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roudi</surname> <given-names>Y.</given-names></name> <name><surname>Nirenberg</surname> <given-names>S.</given-names></name> <name><surname>Latham</surname> <given-names>P. E.</given-names></name></person-group> (<year>2009a</year>). <article-title>Pairwise maximum entropy models for studying large biological systems: when they can work and when they can&#x00027;t</article-title>. <source>PLoS Comp. Biol</source>. <volume>5</volume>:<fpage>e1000380</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1000380</pub-id><pub-id pub-id-type="pmid">19424487</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roudi</surname> <given-names>Y.</given-names></name> <name><surname>Tyrcha</surname> <given-names>J.</given-names></name> <name><surname>Hertz</surname> <given-names>J.</given-names></name></person-group> (<year>2009b</year>). <article-title>Ising model for neural data: model quality and approximate methods for extracting functional connectivity</article-title>. <source>Phys. Rev. E</source> <volume>79</volume>, <fpage>051915</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevE.79.051915</pub-id><pub-id pub-id-type="pmid">19518488</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ruderman</surname> <given-names>D. L.</given-names></name> <name><surname>Bialek</surname> <given-names>W.</given-names></name></person-group> (<year>1994</year>). <article-title>Statistics of natural images: scaling in the woods</article-title>. <source>Phys. Rev. Lett</source>. <volume>73</volume>, <fpage>814</fpage>&#x02013;<lpage>818</lpage>. <pub-id pub-id-type="doi">10.1103/PhysRevLett.73.814</pub-id><pub-id pub-id-type="pmid">10057546</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Santos</surname> <given-names>G. S.</given-names></name> <name><surname>Gireesh</surname> <given-names>E. D.</given-names></name> <name><surname>Plenz</surname> <given-names>D.</given-names></name> <name><surname>Nakahara</surname> <given-names>H.</given-names></name></person-group> (<year>2010</year>). <article-title>Hierarchical interaction structure of neural activities in cortical slice cultures</article-title>. <source>J. Neurosci</source>. <volume>30</volume>, <fpage>8720</fpage>&#x02013;<lpage>8733</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.6141-09.2010</pub-id><pub-id pub-id-type="pmid">20592194</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schneidman</surname> <given-names>E.</given-names></name> <name><surname>Berry (II)</surname> <given-names>M. J.</given-names></name> <name><surname>Segev</surname> <given-names>R.</given-names></name> <name><surname>Bialek</surname> <given-names>W.</given-names></name></person-group> (<year>2006</year>). <article-title>Weak pairwise correlations imply strongly correlated network states in a neural population</article-title>. <source>Nature</source> <volume>440</volume>, <fpage>1007</fpage>&#x02013;<lpage>1012</lpage>. <pub-id pub-id-type="doi">10.1038/nature04701</pub-id><pub-id pub-id-type="pmid">16625187</pub-id></citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schneidman</surname> <given-names>E.</given-names></name> <name><surname>Still</surname> <given-names>S.</given-names></name> <name><surname>Berry (II)</surname> <given-names>M. J.</given-names></name> <name><surname>Bialek</surname> <given-names>W.</given-names></name></person-group> (<year>2003</year>). <article-title>Network information and connected correlations</article-title>. <source>Phys. Rev. Lett</source>. <volume>91</volume>, <fpage>238701</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevLett.91.238701</pub-id><pub-id pub-id-type="pmid">14683220</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sharpe</surname> <given-names>L. T.</given-names></name> <name><surname>Whittle</surname> <given-names>P.</given-names></name> <name><surname>Nordby</surname> <given-names>K.</given-names></name></person-group> (<year>1993</year>). <article-title>Spatial integration and sensitivity changes in the human rod visual system</article-title>. <source>J. Physiol</source>. <volume>461</volume>, <fpage>235</fpage>&#x02013;<lpage>246</lpage>. <pub-id pub-id-type="pmid">8350263</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shea-Brown</surname> <given-names>E.</given-names></name> <name><surname>Josi&#x00107;</surname> <given-names>K.</given-names></name> <name><surname>Doiron</surname> <given-names>B.</given-names></name> <name><surname>de la Rocha</surname> <given-names>J.</given-names></name></person-group> (<year>2008</year>). <article-title>Correlation and synchrony transfer in integrate-and-fire neurons: basic properties and consequences for coding</article-title>. <source>Phys. Rev. Lett</source>. <volume>100</volume>, <fpage>108102</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevLett.100.108102</pub-id><pub-id pub-id-type="pmid">18352234</pub-id></citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shlens</surname> <given-names>J.</given-names></name> <name><surname>Field</surname> <given-names>G. D.</given-names></name> <name><surname>Gauthier</surname> <given-names>J. L.</given-names></name> <name><surname>Greschner</surname> <given-names>M.</given-names></name> <name><surname>Sher</surname> <given-names>A.</given-names></name> <name><surname>Litke</surname> <given-names>A. M.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>The structure of large-scale synchronized firing in primate retina</article-title>. <source>J. Neurosci</source>. <volume>29</volume>, <fpage>5022</fpage>&#x02013;<lpage>5031</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.5187-08.2009</pub-id><pub-id pub-id-type="pmid">19369571</pub-id></citation>
</ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shlens</surname> <given-names>J.</given-names></name> <name><surname>Field</surname> <given-names>G. D.</given-names></name> <name><surname>Gauthier</surname> <given-names>J. L.</given-names></name> <name><surname>Grivich</surname> <given-names>M. I.</given-names></name> <name><surname>Petrusca</surname> <given-names>D.</given-names></name> <name><surname>Sher</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2006</year>). <article-title>The structure of multi-neuron firing patterns in primate retina</article-title>. <source>J. Neurosci</source>. <volume>26</volume>, <fpage>8254</fpage>&#x02013;<lpage>8266</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.1282-06.2006</pub-id><pub-id pub-id-type="pmid">16899720</pub-id></citation>
</ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tang</surname> <given-names>A.</given-names></name> <name><surname>Jackson</surname> <given-names>D.</given-names></name> <name><surname>Hobbs</surname> <given-names>J.</given-names></name> <name><surname>Smith</surname> <given-names>J. L.</given-names></name> <name><surname>Patel</surname> <given-names>H.</given-names></name> <name><surname>Prieto</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2008</year>). <article-title>A maximum entropy model applied to spatial and temporal correlations from cortical networks <italic>in vitro</italic></article-title>. <source>J. Neurosci</source>. <volume>28</volume>, <fpage>505</fpage>&#x02013;<lpage>518</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.3359-07.2008</pub-id><pub-id pub-id-type="pmid">18184793</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tao</surname> <given-names>T.</given-names></name></person-group> (<year>2012</year>). <source>Topics in Random Matrix Theory</source>. <publisher-loc>Providence, RI</publisher-loc>: <publisher-name>American Mathematical Society</publisher-name>. <pub-id pub-id-type="doi">10.1142/S2010326311500018</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tchumatchenko</surname> <given-names>T.</given-names></name> <name><surname>Malyshev</surname> <given-names>A.</given-names></name> <name><surname>Geisel</surname> <given-names>T.</given-names></name> <name><surname>Wolf</surname> <given-names>F.</given-names></name></person-group> (<year>2010</year>). <article-title>Correlations and synchrony in threshold neuron models</article-title>. <source>Phys. Rev. Lett</source>. <volume>104</volume>, <fpage>058102</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevLett.104.058102</pub-id><pub-id pub-id-type="pmid">20366796</pub-id></citation>
</ref>
<ref id="B53">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Tkacik</surname> <given-names>G.</given-names></name> <name><surname>Ghosh</surname> <given-names>A.</given-names></name> <name><surname>Schneidman</surname> <given-names>E.</given-names></name> <name><surname>Segev</surname> <given-names>R.</given-names></name></person-group> (<year>2012</year>). <article-title>Retinal adaptation and invariance in changes to higher-order stimulus statistics</article-title>. arXiv:1201.3552.</citation>
</ref>
<ref id="B54">
<citation citation-type="other"><person-group person-group-type="author"><name><surname>Tkacik</surname> <given-names>G.</given-names></name> <name><surname>Schneidman</surname> <given-names>E.</given-names></name> <name><surname>Berry II</surname> <given-names>M. J.</given-names></name> <name><surname>Bialek</surname> <given-names>W.</given-names></name></person-group> (<year>2009</year>). <article-title>Spin glass models for a network of real neurons</article-title>. arXiv:0912.5409.</citation>
</ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Trong</surname> <given-names>P. K.</given-names></name> <name><surname>Rieke</surname> <given-names>F.</given-names></name></person-group> (<year>2008</year>). <article-title>Origin of correlated activity between parasol retinal ganglion cells</article-title>. <source>Nat. Neurosci</source>. <volume>11</volume>, <fpage>1343</fpage>&#x02013;<lpage>1351</lpage>. <pub-id pub-id-type="doi">10.1038/nn.2199</pub-id><pub-id pub-id-type="pmid">18820692</pub-id></citation>
</ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Victor</surname> <given-names>J.</given-names></name></person-group> (<year>1999</year>). <article-title>Temporal aspects of neural coding in the retina and lateral geniculate</article-title>. <source>Netw. Comput. Neur. Syst</source>. <volume>10</volume>, <fpage>1</fpage>&#x02013;<lpage>66</lpage>. <pub-id pub-id-type="doi">10.1088/0954-898X/10/4/201</pub-id><pub-id pub-id-type="pmid">10695759</pub-id></citation>
</ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vilela</surname> <given-names>R. D.</given-names></name> <name><surname>Lindner</surname> <given-names>B.</given-names></name></person-group> (<year>2009</year>). <article-title>Comparative study of different integrate-and-fire neurons: spontaneous activity, dynamical response, and stimulus-induced correlation</article-title>. <source>Phys. Rev. E</source> <volume>80</volume>, <fpage>031909</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevE.80.031909</pub-id><pub-id pub-id-type="pmid">19905148</pub-id></citation>
</ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>S.</given-names></name> <name><surname>Huang</surname> <given-names>D.</given-names></name> <name><surname>Singer</surname> <given-names>W.</given-names></name> <name><surname>Nikoli&#x00107;</surname> <given-names>D.</given-names></name></person-group> (<year>2008</year>). <article-title>A small world of neuronal synchrony</article-title>. <source>Cereb. Cortex</source> <volume>18</volume>, <fpage>2891</fpage>&#x02013;<lpage>2901</lpage>. <pub-id pub-id-type="doi">10.1093/cercor/bhn047</pub-id><pub-id pub-id-type="pmid">18400792</pub-id></citation>
</ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>S.</given-names></name> <name><surname>Yang</surname> <given-names>H.</given-names></name> <name><surname>Nakahara</surname> <given-names>H.</given-names></name> <name><surname>Santos</surname> <given-names>G.</given-names></name> <name><surname>Nikoli&#x00107;</surname> <given-names>D.</given-names></name> <name><surname>Plenz</surname> <given-names>D.</given-names></name></person-group> (<year>2011</year>). <article-title>Higher-order interactions characterized in cortical activity</article-title>. <source>J. Neurosci</source>. <volume>31</volume>, <fpage>17514</fpage>&#x02013;<lpage>17526</lpage>. <pub-id pub-id-type="doi">10.1523/JNEUROSCI.3127-11.2011</pub-id><pub-id pub-id-type="pmid">22131413</pub-id></citation>
</ref>
</ref-list>
<app-group>
<app id="A1">
<title>Appendix</title>
<sec>
<title>A.1. A measure of higher-order interactions: <italic>D</italic><sub><italic>KL</italic></sub>(<italic>P</italic>, <inline-formula><mml:math id="M165"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>)</title>
<p>We begin by observing that when <inline-formula><mml:math id="M166"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula> is a maximum entropy distribution that approximates <italic>P</italic> (that is, it is log-linear, with coefficients chosen to enforce equality of a set of moments), then the KL-distance may be written as a difference of entropies (Cover and Thomas, <xref ref-type="bibr" rid="B13">1991</xref>; Malouf, <xref ref-type="bibr" rid="B28">2002</xref>):</p>
<disp-formula id="E31"><mml:math id="M167"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mtext>KL</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mi>S</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>Here, the entropy of a probability distribution <italic>P</italic> on {0, 1}<sup>3</sup> is given</p>
<disp-formula id="E32"><label>(18)</label><mml:math id="M168"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mi>S</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mi>log</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mn>3</mml:mn><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mi>log</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mn>3</mml:mn><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mi>log</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mi>log</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>if we use the fact that the distributions are permutation-symmetric [i.e., <italic>p</italic><sub>1</sub> &#x02261; <italic>P</italic>(1, 0, 0) &#x0003D; <italic>P</italic>(0, 1, 0) &#x0003D; <italic>P</italic>(0, 0, 1)]. We take the logarithms in the definitions of the entropy <italic>S</italic> and KL-divergence <italic>D</italic><sub>KL</sub> to be base 2, so that any numerical values of these quantities are in units of bits. Using the fact that <italic>P</italic> must normalize to 1, we rewrite</p>
<disp-formula id="E33"><mml:math id="M169"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mi>S</mml:mi><mml:mtext>&#x0200B;</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mn>3</mml:mn><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mn>3</mml:mn><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>log</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mn>3</mml:mn><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mn>3</mml:mn><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x02212;</mml:mo><mml:mn>3</mml:mn><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mi>log</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mn>3</mml:mn><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mi>log</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mi>log</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where the set of admissible distributions may now be described by the convex tetrahedron in &#x0211D;<sup>3</sup>, <graphic xlink:href="fncom-08-00010-i0001.tif"/> &#x0003D; {<italic>p</italic><sub>1</sub>, <italic>p</italic><sub>2</sub>, <italic>p</italic><sub>3</sub> &#x02265; 0; 3<italic>p</italic><sub>1</sub> &#x0002B; 3<italic>p</italic><sub>2</sub> &#x0002B; <italic>p</italic><sub>3</sub> &#x02264; 1}</p>
<p>We note that the set of distributions which satisfies a desired set of lower order moments is given by an affine subspace (in &#x0211D;<sup>3</sup>, a line) which intersects this tetrahedron:</p>
<disp-formula id="E34"><mml:math id="M170"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mi>&#x003BC;</mml:mi><mml:mo>&#x02261;</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mo stretchy='false'>[</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>]</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mn>2</mml:mn><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mover accent='true'><mml:mi>&#x003C1;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mo>&#x02261;</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mo stretchy='false'>[</mml:mo><mml:msubsup><mml:mi>X</mml:mi><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:msubsup><mml:mo stretchy='false'>]</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Denoting this set by <graphic xlink:href="fncom-08-00010-i0001.tif"/><sub>&#x003BC;, <inline-formula><mml:math id="M171"><mml:mover accent='true'><mml:mi>&#x003C1;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover></mml:math></inline-formula></sub>, we note that <graphic xlink:href="fncom-08-00010-i0001.tif"/><sub>&#x003BC;, <inline-formula><mml:math id="M172"><mml:mover accent='true'><mml:mi>&#x003C1;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover></mml:math></inline-formula></sub> is a convex set and that <italic>S</italic>(<inline-formula><mml:math id="M173"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) is constant on each <graphic xlink:href="fncom-08-00010-i0001.tif"/><sub>&#x003BC;, <inline-formula><mml:math id="M174"><mml:mover accent='true'><mml:mi>&#x003C1;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover></mml:math></inline-formula></sub>.</p>
<p>By straightforward differentiation we can check that the Hessian of &#x02212;<italic>S</italic>(<italic>P</italic>) is positive definite, as long as the probabilities <italic>p</italic><sub>0</sub>, <italic>p</italic><sub>1</sub>, etc. are strictly greater than zero:</p>
<disp-formula id="E35"><mml:math id="M175"><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mi>D</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mi>S</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mfrac><mml:mn>3</mml:mn><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:mfrac></mml:mrow></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mrow><mml:mfrac><mml:mn>3</mml:mn><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mfrac></mml:mrow></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow></mml:mfrac></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:mrow></mml:mfrac><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mn>9</mml:mn></mml:mtd><mml:mtd><mml:mn>9</mml:mn></mml:mtd><mml:mtd><mml:mn>3</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>9</mml:mn></mml:mtd><mml:mtd><mml:mn>9</mml:mn></mml:mtd><mml:mtd><mml:mn>3</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>3</mml:mn></mml:mtd><mml:mtd><mml:mn>3</mml:mn></mml:mtd><mml:mtd><mml:mn>1</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>Therefore &#x02212;<italic>S</italic>(<italic>P</italic>) is convex on <graphic xlink:href="fncom-08-00010-i0001.tif"/><sub>&#x003BC;, <inline-formula><mml:math id="M176"><mml:mover accent='true'><mml:mi>&#x003C1;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover></mml:math></inline-formula></sub>; since <italic>S</italic>(<inline-formula><mml:math id="M177"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) is constant, <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M178"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) is likewise convex on <graphic xlink:href="fncom-08-00010-i0001.tif"/><sub>&#x003BC;, <inline-formula><mml:math id="M179"><mml:mover accent='true'><mml:mi>&#x003C1;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover></mml:math></inline-formula></sub>. As a consequence, if <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M180"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) has a local minimum, then it is unique and a global minimum as well. Since <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M181"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) &#x02265; 0 with equality if and only if <italic>P</italic> &#x0003D; <inline-formula><mml:math id="M182"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>, this minimum must be achieved occurs when <italic>P</italic> &#x0003D; <inline-formula><mml:math id="M183"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>; the maximum is likewise achieved on the boundary of the admissible region <graphic xlink:href="fncom-08-00010-i0001.tif"/><sub>&#x003BC;, <inline-formula><mml:math id="M184"><mml:mover accent='true'><mml:mi>&#x003C1;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover></mml:math></inline-formula></sub>.</p>
</sec>
<sec>
<title>A.2. A measure of higher-order interactions: strain</title>
<p>We define the <italic>strain</italic>,</p>
<disp-formula id="E36"><label>(19)</label><mml:math id="M185"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mi>&#x003B3;</mml:mi><mml:mo>=</mml:mo><mml:mi>log</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:msubsup><mml:mi>p</mml:mi><mml:mn>1</mml:mn><mml:mn>3</mml:mn></mml:msubsup></mml:mrow><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:msubsup><mml:mi>p</mml:mi><mml:mn>2</mml:mn><mml:mn>3</mml:mn></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:mi>log</mml:mi><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mi>log</mml:mi><mml:msub><mml:mi>p</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mn>3</mml:mn><mml:mi>log</mml:mi><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mn>3</mml:mn><mml:mi>log</mml:mi><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>a potential measure of the importance of higher-order interactions (Ohiorhenuan and Victor, <xref ref-type="bibr" rid="B37">2010</xref>). By Equation (3), we can see that &#x003B3; &#x0003D; 0 precisely for a pairwise maximum entropy (PME) distribution. We will show that as the distribution (<italic>p</italic><sub>0</sub>, <italic>p</italic><sub>1</sub>, <italic>p</italic><sub>2</sub>, <italic>p</italic><sub>3</sub>) is moved away from the constraint surface while fixing lower-order moments, the strain increases monotonically.</p>
<p>From the definition of lower-order moments,</p>
<disp-formula id="E37"><mml:math id="M186"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mi>&#x003BC;</mml:mi><mml:mo>=</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mo stretchy='false'>[</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>]</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mn>2</mml:mn><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mover accent='true'><mml:mi>&#x003C1;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mo stretchy='false'>[</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>X</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy='false'>]</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>we can verify that in order to keep &#x003BC;, <inline-formula><mml:math id="M187"><mml:mover accent='true'><mml:mi>&#x003C1;</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover></mml:math></inline-formula> constant, if <italic>p</italic><sub>1</sub> increases by <italic>z</italic> (i.e., <italic>p</italic><sub>1</sub> &#x02192; <italic>p</italic><sub>1</sub> &#x0002B; <italic>z</italic>), then we must also have <italic>p</italic><sub>2</sub> &#x02192; <italic>p</italic><sub>2</sub> &#x02212; <italic>z</italic> and <italic>p</italic><sub>3</sub> &#x02192; <italic>p</italic><sub>3</sub> &#x0002B; <italic>z</italic>. Then if each probability is strictly positive, then the derivative</p>
<disp-formula id="E38"><mml:math id="M188"><mml:mrow><mml:mfrac><mml:mrow><mml:mo>&#x02202;</mml:mo><mml:mi>&#x003B3;</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x02202;</mml:mo><mml:mi>z</mml:mi></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mi>z</mml:mi></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mn>3</mml:mn><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mn>3</mml:mn><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mi>z</mml:mi></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mfrac><mml:mn>3</mml:mn><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mi>z</mml:mi></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mfrac><mml:mn>3</mml:mn><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mi>z</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>is strictly positive as well. In particular, it is strictly positive at <italic>z</italic> &#x0003D; 0 and will remain positive until <italic>z</italic> reaches a value such that one of the denominators reaches 0. Therefore &#x003B3; increases monotonically for <italic>z</italic> &#x0003E; 0 and decreases monotonically for <italic>z</italic> &#x0003C; 0.</p>
</sec>
<sec>
<title>A.3. An analytical explanation for unimodal vs. bimodal effects</title>
<p>We consider an analytical argument to support the numerical results that bimodal inputs generate larger deviations from PME model fits than unimodal inputs. As a metric, we consider <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M189"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>)&#x02014;where <italic>P</italic> and <inline-formula><mml:math id="M190"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula> are again the true and model distributions, respectively&#x02014;when we perturb an independent spiking distribution by adding a common, global input of variance <italic>c</italic>. To simplify notation, the small parameter in the calculation will be denoted <inline-formula><mml:math id="M191"><mml:mrow><mml:mi>&#x003F5;</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:mi>c</mml:mi></mml:msqrt></mml:mrow></mml:math></inline-formula>.</p>
<p>We now compute <italic>S</italic>(<italic>P</italic>) and <italic>S</italic>(<inline-formula><mml:math id="M192"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) (defined in an earlier Appendix) by deriving a series expansion for each set of event probabilities. We can compute the true distribution <italic>P</italic> using the expressions derived in Equation (18); to recap, let the common input <italic>I</italic><sub><italic>c</italic></sub> have probability density <italic>p</italic>(<italic>I</italic><sub><italic>c</italic></sub>), and the independent input to each cell, <italic>x</italic>, have density <italic>p</italic><sub><italic>s</italic></sub>(<italic>x</italic>). Let &#x00398; be the threshold for generating a spike (i.e., a &#x0201C;1&#x0201D; response). For each cell, a spike is generated if <italic>x</italic> &#x0002B; <italic>I</italic><sub><italic>c</italic></sub> &#x0003E; &#x00398;, i.e., with probability</p>
<disp-formula id="E39"><mml:math id="M193"><mml:mrow><mml:mi>d</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle='true'><mml:mrow><mml:msubsup><mml:mo>&#x0222B;</mml:mo><mml:mrow><mml:mi>&#x00398;</mml:mi><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mi>&#x0221E;</mml:mi></mml:msubsup><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mstyle><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mi>d</mml:mi><mml:mi>x</mml:mi><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
<p>Given <italic>I</italic><sub><italic>c</italic></sub>, this is conditionally independent for each cell. We can therefore write our probabilities by integrating over <italic>I</italic><sub><italic>c</italic></sub> as follows:</p>
<disp-formula id="E40"><label>(20)</label><mml:math id="M194"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle='true'><mml:mrow><mml:msubsup><mml:mo>&#x0222B;</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mi>&#x0221E;</mml:mi></mml:mrow><mml:mi>&#x0221E;</mml:mi></mml:msubsup><mml:mi>p</mml:mi></mml:mrow></mml:mstyle><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>d</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mn>3</mml:mn></mml:msup><mml:mi>d</mml:mi><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle='true'><mml:mrow><mml:msubsup><mml:mo>&#x0222B;</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mi>&#x0221E;</mml:mi></mml:mrow><mml:mi>&#x0221E;</mml:mi></mml:msubsup><mml:mi>p</mml:mi></mml:mrow></mml:mstyle><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>d</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mi>d</mml:mi><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle='true'><mml:mrow><mml:msubsup><mml:mo>&#x0222B;</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mi>&#x0221E;</mml:mi></mml:mrow><mml:mi>&#x0221E;</mml:mi></mml:msubsup><mml:mi>p</mml:mi></mml:mrow></mml:mstyle><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>d</mml:mi><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle='true'><mml:mrow><mml:msubsup><mml:mo>&#x0222B;</mml:mo><mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mi>&#x0221E;</mml:mi></mml:mrow><mml:mi>&#x0221E;</mml:mi></mml:msubsup><mml:mi>p</mml:mi></mml:mrow></mml:mstyle><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mn>3</mml:mn></mml:msup><mml:mi>d</mml:mi><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>We develop a perturbation argument in the limit of very weak common input. That is, <italic>p</italic>(<italic>I</italic><sub><italic>c</italic></sub>) is close to a delta function centered at <italic>I</italic><sub><italic>c</italic></sub> &#x0003D; 0. Take <italic>p</italic>(<italic>I</italic><sub><italic>c</italic></sub>) to be a scaled function</p>
<disp-formula id="E41"><label>(21)</label><mml:math id="M195"><mml:mrow><mml:mi>p</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>&#x003F5;</mml:mi></mml:mfrac><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>We place no constraints on <italic>f</italic>(<italic>x</italic>), other than that it must be normalized (<bold>E</bold>[1] &#x0003D; 1) and that its moments must be finite (so that <bold>E</bold>[<italic>I</italic><sub><italic>c</italic></sub>], <bold>E</bold>[<italic>I</italic><sup>2</sup><sub><italic>c</italic></sub>], and so forth will exist, where <bold>E</bold>[<italic>g</italic>(<italic>x</italic>)] &#x02261; &#x0222B;<sup>&#x0221E;</sup><sub>&#x02212;&#x0221E;</sub> <italic>g</italic>(<italic>x</italic>) <italic>f</italic>(<italic>x</italic>) <italic>dx</italic>).</p>
<p>For the moment, assume that the function <italic>f</italic>(<italic>x</italic>) has a single maximum at <italic>x</italic> &#x0003D; 0. To evaluate the integrals above, we Taylor-expand <italic>d</italic>(<italic>x</italic>) around <italic>x</italic> &#x0003D; 0. Anticipating a sixth-order term to survive, we keep all terms up to this order. This gives, for small <italic>x</italic>,</p>
<disp-formula id="E42"><mml:math id="M196"><mml:mrow><mml:mi>d</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02248;</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mn>6</mml:mn></mml:munderover><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:mstyle><mml:msup><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>+</mml:mo><mml:mi>O</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mn>7</mml:mn></mml:msup><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where <italic>a</italic><sub>1</sub> &#x0003D; <italic>p</italic><sub><italic>s</italic></sub>(&#x00398;) (the other coefficients <italic>a</italic><sub>2</sub> &#x02212; <italic>a</italic><sub>6</sub> can be given similarly in terms of the independent input distribution at &#x00398;). Substituting this into the expressions for <italic>p</italic><sub>0</sub>, etc., above, with <italic>p</italic>(<italic>I</italic><sub><italic>c</italic></sub>) given as in Equation (21), gives us each event as a series in &#x003F5;; for example,</p>
<disp-formula id="E43"><mml:math id="M197"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mi>d</mml:mi><mml:mn>0</mml:mn><mml:mn>3</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>3</mml:mn><mml:msub><mml:mi>a</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:msubsup><mml:mi>d</mml:mi><mml:mn>0</mml:mn><mml:mn>2</mml:mn></mml:msubsup><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mo stretchy='false'>[</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>&#x003F5;</mml:mi><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>3</mml:mn><mml:msubsup><mml:mi>a</mml:mi><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:msubsup><mml:msub><mml:mi>d</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mn>3</mml:mn><mml:msub><mml:mi>a</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:msubsup><mml:mi>d</mml:mi><mml:mn>0</mml:mn><mml:mn>2</mml:mn></mml:msubsup><mml:mo stretchy='false'>)</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mo stretchy='false'>[</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo stretchy='false'>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mi>&#x003F5;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where expectations are, again, with respect to the unscaled PDF <italic>f</italic>(<italic>x</italic>). The entropy <italic>S</italic>(<italic>P</italic>) is now given by using these series expansions in Equation (18).</p>
<p>We note that our derivation does not rely on the fact that the distribution of common input is peaked at <italic>I</italic><sub><italic>c</italic></sub> &#x0003D; 0 in particular. For example, we could have a common input centered around &#x003BC;. The common input distribution function would be of the form</p>
<disp-formula id="E44"><mml:math id="M198"><mml:mrow><mml:mi>p</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>&#x003F5;</mml:mi></mml:mfrac><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>Changing &#x003F5; regulates the variance, but doesn&#x00027;t change the mean or the peak (assuming, without loss of generality, that the peak of <italic>f</italic> occurs at zero). The peak of <italic>p</italic>(<italic>I</italic><sub><italic>c</italic></sub>) now occurs at &#x003BC;, and the appropriate Taylor expansion of <italic>d</italic>(<italic>x</italic>) is</p>
<disp-formula id="E45"><mml:math id="M199"><mml:mrow><mml:mi>d</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02248;</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>&#x003BC;</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mn>6</mml:mn></mml:munderover><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:mstyle><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mi>&#x003BC;</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mi>k</mml:mi></mml:msup><mml:mo>+</mml:mo><mml:mi>O</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mn>7</mml:mn></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
<p>where the coefficients <italic>b</italic><sub><italic>k</italic></sub> now depend on the local behavior of <italic>d</italic> around &#x003BC;. The expectations that appear in the expansion of <italic>p</italic><sub>3</sub>, and so forth, are now centered moments taken around &#x003BC;; the calculations are otherwise identical. In other words, the perturbation expansion requires the <italic>variance</italic> of the common input to be small, but not the mean.</p>
<p>For bimodal inputs, we consider a common input with a probability distribution of the following form:</p>
<disp-formula id="E46"><mml:math id="M200"><mml:mrow><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mi>&#x003F5;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mi>&#x003F5;</mml:mi></mml:mfrac><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msup><mml:mi>&#x003F5;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mfrac><mml:mn>1</mml:mn><mml:mi>&#x003F5;</mml:mi></mml:mfrac><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>so that most of the probability distribution is peaked at zero, but there is a second peak of higher order (here taken at <italic>I</italic><sub><italic>c</italic></sub> &#x0003D; 1, without loss of generality). Again, we approximate the integrals given in Equation (20), and therefore the entropy <italic>S</italic>(<italic>P</italic>), by Taylor expanding <italic>d</italic>(<italic>x</italic>);</p>
<disp-formula id="E47"><mml:math id="M201"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mi>d</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02248;</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mn>6</mml:mn></mml:munderover><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:mstyle><mml:msup><mml:mi>x</mml:mi><mml:mi>k</mml:mi></mml:msup><mml:mo>+</mml:mo><mml:mi>O</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mn>7</mml:mn></mml:msup><mml:mo stretchy='false'>)</mml:mo><mml:mo>;</mml:mo><mml:mtext>&#x02003;</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x02248;</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x02248;</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mn>6</mml:mn></mml:munderover><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:mstyle><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>k</mml:mi></mml:msup><mml:mo>+</mml:mo><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>7</mml:mn></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>;</mml:mo><mml:mtext>&#x02003;</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>&#x02248;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>around the two peaks 0 and 1, respectively. For each integral we have the same contributions from the unimodal case, multiplied by (1 &#x02212; &#x003F5;<sup>2</sup>), as well as the corresponding contributions from the second peak multiplied by &#x003F5;<sup>2</sup> (these weightings are chosen so that the common input has variance of order &#x003F5;<sup>2</sup>, as in the unimodal case). This makes clear at what order every term enters.</p>
<p>We now construct an expansion for the PME model <inline-formula><mml:math id="M202"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>:</p>
<disp-formula id="E48"><mml:math id="M203"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>Z</mml:mi></mml:mfrac><mml:mi>exp</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>&#x003BB;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mrow><mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x003BB;</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>We approach this problem by describing &#x003BB;<sub>1</sub> and &#x003BB;<sub>2</sub> as a series in &#x003F5;. We match coefficients by forcing the first and second moments of <inline-formula><mml:math id="M204"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula> to match those of <italic>P</italic>&#x02014;as they must. Specifically, take</p>
<disp-formula id="E49"><mml:math id="M205"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>&#x003BB;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mover accent='true'><mml:mi>&#x003BB;</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover><mml:mo>+</mml:mo><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mn>6</mml:mn></mml:munderover><mml:mrow><mml:msup><mml:mi>&#x003F5;</mml:mi><mml:mi>k</mml:mi></mml:msup></mml:mrow></mml:mstyle><mml:msub><mml:mi>u</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mi>&#x003F5;</mml:mi><mml:mn>7</mml:mn></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>&#x003BB;</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mn>6</mml:mn></mml:munderover><mml:mrow><mml:msup><mml:mi>&#x003F5;</mml:mi><mml:mi>k</mml:mi></mml:msup></mml:mrow></mml:mstyle><mml:msub><mml:mi>v</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mi>&#x003F5;</mml:mi><mml:mn>7</mml:mn></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M206"><mml:mrow><mml:msub><mml:mi>&#x003BB;</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mover accent='true'><mml:mi>&#x003BB;</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>, &#x003BB;<sub>2</sub> &#x0003D; 0 are the corresponding parameters from the independent case. The events <inline-formula><mml:math id="M207"><mml:mover accent='true'><mml:mi>p</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula><sub>0</sub>, <inline-formula><mml:math id="M208"><mml:mover accent='true'><mml:mi>p</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula><sub>1</sub>, <inline-formula><mml:math id="M209"><mml:mover accent='true'><mml:mi>p</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula><sub>2</sub>, and <inline-formula><mml:math id="M210"><mml:mover accent='true'><mml:mi>p</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula><sub>3</sub> can be written as a series in &#x003F5;. We then require that the mean and centered second moments of <inline-formula><mml:math id="M211"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula> match those of <italic>P</italic>; that is</p>
<disp-formula id="E50"><mml:math id="M212"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mn>2</mml:mn><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mover accent='true'><mml:mi>p</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover><mml:mn>1</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mn>2</mml:mn><mml:msub><mml:mover accent='true'><mml:mi>p</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mover accent='true'><mml:mi>p</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover><mml:mn>3</mml:mn></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mn>2</mml:mn><mml:msub><mml:mi>p</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mn>3</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mover accent='true'><mml:mi>p</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mover accent='true'><mml:mi>p</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover><mml:mn>3</mml:mn></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mover accent='true'><mml:mi>p</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover><mml:mn>1</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:mn>2</mml:mn><mml:msub><mml:mover accent='true'><mml:mi>p</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover><mml:mn>2</mml:mn></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mover accent='true'><mml:mi>p</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover><mml:mn>3</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>At each order <italic>k</italic>, this yields a system of two linear equations in <italic>u</italic><sub><italic>k</italic></sub> and <italic>v</italic><sub><italic>k</italic></sub>; we solve, inductively, up to the desired order; we now have <inline-formula><mml:math id="M213"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>, and therefore <italic>S</italic>(<inline-formula><mml:math id="M214"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>), as a series in &#x003F5;.</p>
<p>Finally, we combine the two series to find that in the <italic>unimodal</italic> case,</p>
<disp-formula id="E51"><label>(22)</label><mml:math id="M215"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mtext>KL</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>S</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:msup><mml:mi>&#x003F5;</mml:mi><mml:mn>6</mml:mn></mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msubsup><mml:mi>a</mml:mi><mml:mn>1</mml:mn><mml:mn>6</mml:mn></mml:msubsup><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>2</mml:mn><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:msup><mml:mrow><mml:mo stretchy='false'>[</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>]</mml:mo></mml:mrow><mml:mn>3</mml:mn></mml:msup><mml:mo>&#x02212;</mml:mo><mml:mn>3</mml:mn><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mo stretchy='false'>[</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>]</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mo stretchy='false'>[</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo stretchy='false'>]</mml:mo><mml:mo>+</mml:mo><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>E</mml:mi></mml:mstyle><mml:mo stretchy='false'>[</mml:mo><mml:msup><mml:mi>x</mml:mi><mml:mn>3</mml:mn></mml:msup><mml:mo stretchy='false'>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>3</mml:mn></mml:msup><mml:msubsup><mml:mi>d</mml:mi><mml:mn>0</mml:mn><mml:mn>3</mml:mn></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>+</mml:mo><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mi>&#x003F5;</mml:mi><mml:mn>7</mml:mn></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>If the first two odd moments of the distribution are zero (something we can expect for &#x0201C;symmetric&#x0201D; distributions, such as a Gaussian), then this sixth-order term is zero as well.</p>
<p>For the <italic>bimodal</italic> case</p>
<disp-formula id="E52"><mml:math id="M216"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>D</mml:mi><mml:mrow><mml:mtext>KL</mml:mtext></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>S</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x02212;</mml:mo><mml:mi>S</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>=</mml:mo><mml:msup><mml:mi>&#x003F5;</mml:mi><mml:mn>4</mml:mn></mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>6</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>3</mml:mn></mml:msup><mml:msubsup><mml:mi>d</mml:mi><mml:mn>0</mml:mn><mml:mn>3</mml:mn></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msup><mml:mi>&#x003F5;</mml:mi><mml:mn>5</mml:mn></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>This last term depends on the distance <italic>d</italic><sub>1</sub> &#x02212; <italic>d</italic><sub>0</sub>, in other words, how much more likely the independent input is to push the cell over threshold when common input is &#x0201C;ON&#x0201D;. We can also view this as depending on the ratio <inline-formula><mml:math id="M217"><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:mrow></mml:mfrac></mml:mrow></mml:math></inline-formula>, which gives the fraction of previously non-spiking cells that now spike as a result of the common input.</p>
<p><italic>The main point here, of course, is that <italic>D</italic><sub>KL</sub>(<italic>P</italic>, <inline-formula><mml:math id="M218"><mml:mover accent='true'><mml:mi>P</mml:mi><mml:mo>&#x002DC;</mml:mo></mml:mover></mml:math></inline-formula>) is of order &#x003F5;<sup>4</sup> rather than</italic> &#x003F5;<sup>6</sup>. So, as the strength of a common binary vs. unimodal input increases, spiking distributions depart from the PME more rapidly.</p>
</sec>
</app>
</app-group>
</back>
</article>
