<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Appl. Math. Stat.</journal-id>
<journal-title>Frontiers in Applied Mathematics and Statistics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Appl. Math. Stat.</abbrev-journal-title>
<issn pub-type="epub">2297-4687</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">769355</article-id>
<article-id pub-id-type="doi">10.3389/fams.2021.769355</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Applied Mathematics and Statistics</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Statistical Modeling for the Effects of Vegetative Growth on Power Distribution System Reliability</article-title>
<alt-title alt-title-type="left-running-head">Pan</alt-title>
<alt-title alt-title-type="right-running-head">Statistical Modeling for Vegetative Growth</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Pan</surname>
<given-names>Juming</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1450862/overview"/>
</contrib>
</contrib-group>
<aff>Department of Mathematics, Rowan University, <addr-line>Glassboro</addr-line>, <addr-line>NJ</addr-line>, <country>United&#x20;States</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1433997/overview">Min Wang</ext-link>, University of Texas at San Antonio, United&#x20;States</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1465129/overview">Liucang Wu</ext-link>, Kunming University of Science and Technology, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1371460/overview">Fangyuan Zhang</ext-link>, Texas Tech University, United&#x20;States</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Juming Pan, <email>pan@rowan.edu</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Statistics, a section of the journal Frontiers in Applied Mathematics and Statistics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>20</day>
<month>10</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>7</volume>
<elocation-id>769355</elocation-id>
<history>
<date date-type="received">
<day>01</day>
<month>09</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>05</day>
<month>10</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Pan.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Pan</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>The purpose of this paper is to examine the effects of vegetative growth on the reliability of electric power distribution system under normal (storm exclusion) operating conditions, and to determine an effective vegetation maintenance schedule. Generalized statistical linear regression models, including Poisson, Negative Binomial, Zero-Inflated, and their mixed model variants are developed and are applied into a 5-years outage data along with vegetation maintenance history from a power company in Midwestern United&#x20;States. From the methodological point of view, advanced statistical models such as zero-inflated models and mixed models are utilized the first time on outage data and provided good fit to the occurrence of outages. In practice, numerical results from this study suggest that an optimal cycle length of every 6&#xa0;years could be greatly helpful for power companies in devising a cost-effective schedule, improving system reliability, and maintaining customer satisfaction.</p>
</abstract>
<kwd-group>
<kwd>power distribution system</kwd>
<kwd>vegetation maintenance</kwd>
<kwd>vegetationrelated outage</kwd>
<kwd>optimal cycle length</kwd>
<kwd>statistical modeling</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Virtually all power distribution circuitry in the United&#x20;States operate in a multi-earth-grounded configuration. This means that an electrical fault will exist if an electrical connection is made to ground. If the normal growth of vegetation under or next to distribution lines is not held in check by pruning, growth will eventually bring it close enough to the wires to create a conductive path to earth; causing a fault which will lead to an outage.</p>
<p>To ensure high levels of reliability in the distribution system, vegetative maintenance (tree trimming and other vegetation control measures) is periodically conducted by electric utilities. Since the costs of such maintenance counts a large fraction of the total amount spent on distribution system maintenance [<xref ref-type="bibr" rid="B1">1</xref>] and can be millions of dollars even for a small utility [<xref ref-type="bibr" rid="B2">2</xref>], the cycle length is generally selected by economics. It might seem like a longer trimming cycle would be the most economical choice because it spreads the cost out over time. However, this thinking is flawed because the amount of work that has to be done increases very rapidly with time. The work load increases due to two factors. First, the amount of biomass that must be cut and handled increases rapidly with each year of growth and second, as the tree grows closer to the energized lines, workers must use caution to keep from getting hurt or killed. In fact, short cycles are expensive due to the amount of work and long cycles are also costly due to excessive biomass and loss of productivity. Therefore, an advanced understanding on the relationship between vegetative growth and system reliability can help the power companies determine an effective vegetation maintenance schedule.</p>
<p>Some studies have been made to quantitatively analyze the effects of vegetative growth on distribution system outages. The paper [<xref ref-type="bibr" rid="B3">3</xref>] presented a time series and a non-linear machine learning regression model for predicting the number of vegetation-related outages that occur in power distribution systems on a monthly basis. In the article [<xref ref-type="bibr" rid="B1">1</xref>], several direct failure-rate models were investigated to predict the time-varying, vegetation-related failure rates of overhead distribution power lines. The authors in [<xref ref-type="bibr" rid="B4">4</xref>] developed statistical models for estimating the impacts of tree trimming on electric power system outages under normal operating conditions. Another study [<xref ref-type="bibr" rid="B5">5</xref>] reported a maintenance-scheduling algorithm that determines the optimal location and time for performing vegetation maintenance on overhead distribution feeders using a vegetation failure rate model. However, the number of literature on this critical issue is still limited due to the difficulty of collecting outage data, therefore this problem has not been investigated to the fullest extent&#x20;yet.</p>
<p>In this paper, a novel approach to quantify the impact of vegetation growth on electric power system outages is proposed. Utilizing statistical and machine learning predictive models, the proposed method advances the understanding on the relationship between vegetative growth and system reliability and enables effective and timely decision-making actions by power industry. What distinguishes this approach from the other studies in the literature could be summarized in the following three aspects:</p>
<p>&#x2022; From a theoretical perspective, to the best of our knowledge, ours is the first study that includes some statistical models such as zero-inflated Poisson and zero-inflated Negative Binomials, as well as the mixed models into the analysis of actual outage data. Considering a majority of 0&#x2019;s in the data and the clustering nature, those models provide better fit to the occurrence of outages than the existing methods.</p>
<p>&#x2022; One main reason for the perceived lack of studies on this critical problem is the limited access to a sufficient amount of outage data. Our approach takes advantage of a considerable amount of vegetation-related outage data that is provided by a power company in Midwestern United&#x20;States, thus allows a sound statistical basis and strong conclusions to be&#x20;drawn.</p>
<p>&#x2022; In practice, our study explicitly suggests an optimal cycle length of every 6&#xa0;years, which could be greatly helpful for power companies in devising a cost-effective schedule, improving system reliability, and maintaining customer satisfaction.</p>
<p>The rest of this paper is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> describes the outage data. In <xref ref-type="sec" rid="s3">Section 3</xref>, generalized statistical linear regression models, including Poisson, Negative Binomial, Zero-Inflated, and their mixed model variants are introduced and developed. In <xref ref-type="sec" rid="s4">Section 4</xref>, we fit the statistical models on a real outage data, the results are presented and discussed. The paper ends with discussion and future research directions.</p>
</sec>
<sec id="s2">
<title>2 Data Description</title>
<p>The analysis in this paper is based on data from a power company in Midwestern U.S, containing 431&#x20;vegetation-caused outages on 144 circuits for 2012 through 2016. A vegetation-caused outage is defined as any outage caused when the vegetation gets close enough to sway into the line or to create a path for a tree dwelling animal to bridge the air-gap between energized lines and vegetation, under normal (storm exclusion) operating conditions. For this data, the date and time of each outage, the duration of the outage, the name of the circuit where the outage occurred, the number of years since the last routine vegetation maintenance was performed (year of vegetative growth), and the number of customers affected by the outage are recorded. <xref ref-type="fig" rid="F1">Figure&#x20;1</xref> shows the histogram of vegetation-caused outages on each circuit. It can be seen that most circuits only have 1 outage, and the distribution is right skewed.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Histogram of the vegetation-caused outages on each circuit.</p>
</caption>
<graphic xlink:href="fams-07-769355-g001.tif"/>
</fig>
</sec>
<sec id="s3">
<title>3 Statistical Models</title>
<p>In statistics, count data is a type of data in which the observations can take only the non-negative integer values, for example, the number of power outages on a circuit in this paper. When modeling count data, the classical ordinary least-squares (OLS) regression is often inappropriate because the homoscedasticity and normality assumptions are violated. The violation of the basic OLS assumptions can result in inaccurate estimates of standard errors, and misleading <italic>p</italic>-values and consequent confidence intervals [<xref ref-type="bibr" rid="B6">6</xref>]. A class of generalized linear regression models has been developed for modeling count data. These models have a number of advantages over an ordinary linear regression model, including a skew, discrete distribution, and the restriction of predicted values to non-negative numbers&#x20;[<xref ref-type="bibr" rid="B7">7</xref>].</p>
<p>The most widely used model for count data is Poisson model. The Poisson model is made up of a Poisson probability mass function (PMF) denoted as P (<italic>y</italic>
<sub>
<italic>i</italic>
</sub> &#x3d; <italic>k</italic>) used to calculate the probability of observing <italic>k</italic> events given a mean event rate of <italic>&#x3bb;</italic>, and a link function that is used to express the mean rate <italic>&#x3bb;</italic> as a function of the regression variables <bold>
<italic>X</italic>
</bold>. The Poisson PMF is<disp-formula id="e3_1">
<mml:math id="m1">
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msup>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>!</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mspace width="2em"/>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1,2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
</mml:math>
<label>(3.1)</label>
</disp-formula>where <italic>y</italic>
<sub>
<italic>i</italic>
</sub> is the observed count for the <italic>i</italic>th row in the dataset, <italic>&#x3bb;</italic>
<sub>
<italic>i</italic>
</sub> is the event rate corresponding to the <italic>i</italic>th sample. The link function of the Poisson regression model is expressed as<disp-formula id="e3_2">
<mml:math id="m2">
<mml:mi>ln</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b2;</mml:mi>
<mml:mo>,</mml:mo>
</mml:math>
<label>(3.2)</label>
</disp-formula>where &#x2009;ln (&#x22c5;) is the natural logarithm function, <bold>
<italic>x</italic>
</bold>
<sub>
<bold>
<italic>i</italic>
</bold>
</sub> &#x3d; [<italic>x</italic>
<sub>
<italic>i</italic>1</sub>, <italic>x</italic>
<sub>
<italic>i</italic>2</sub>, &#x2026; , <italic>x</italic>
<sub>
<italic>ip</italic>
</sub>] is the regression variables in the <italic>i</italic>th row, and <bold>
<italic>&#x3b2;</italic>
</bold> is the vector of regression coefficients.</p>
<p>The Poisson model assumes that the mean and variance of the errors are equal, but usually in practice the variance of the errors is larger than the mean. In these cases, we say that the data is &#x201c;overdispersed&#x201d;. When overdispersion arises, the Poisson model is not proper, and an alternative is a Negative Binomial (NB) model. The Negative Binomial distribution is a form of the Poisson distribution in which the distribution&#x2019;s parameter is itself considered a random variable. The variation of this parameter can account for a variance of the data that is higher than the mean. With a NB model, the count data are assumed to follow a NB PMF as what follows,<disp-formula id="e3_3">
<mml:math id="m3">
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="normal">&#x393;</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">&#x393;</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi mathvariant="normal">&#x393;</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
</mml:math>
<label>(3.3)</label>
</disp-formula>where &#x393;(&#x22c5;) is the gamma function, <italic>&#x3b1;</italic> is the overdispersion parameter, and the regression model is same as for the Poisson model in&#x20;(3.2).</p>
<p>When count data has both excess zeros and large counts, zero-inflated Poisson regression (ZIP [<xref ref-type="bibr" rid="B8">8</xref>]) is a practical way to deal with such situation. It assumes that with probability <italic>p</italic> the only possible observation is zero, and with probability 1 &#x2212; <italic>p</italic> a Poisson (<italic>&#x3bb;</italic>) random variable is observed. The intuition behind the ZIP model is that there is a second underlying process that is determining whether a count is zero or non-zero. A ZIP distribution can be written as<disp-formula id="e3_4">
<mml:math id="m4">
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="cases">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msup>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>if</mml:mtext>
<mml:mspace width="0.28em"/>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo>;</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msup>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>!</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>if</mml:mtext>
<mml:mspace width="0.28em"/>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1,2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(3.4)</label>
</disp-formula>The Poisson mean <italic>&#x3bb;</italic>
<sub>
<italic>i</italic>
</sub> and the probability <italic>p</italic>
<sub>
<italic>i</italic>
</sub> are linked to the explanatory variables through the log link in (3.2) and logit link as<disp-formula id="e3_5">
<mml:math id="m5">
<mml:mi mathvariant="normal">l</mml:mi>
<mml:mi mathvariant="normal">o</mml:mi>
<mml:mi mathvariant="normal">g</mml:mi>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">t</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b3;</mml:mi>
<mml:mo>,</mml:mo>
</mml:math>
<label>(3.5)</label>
</disp-formula>where <bold>
<italic>z</italic>
</bold>
<sub>
<bold>
<italic>i</italic>
</bold>
</sub> is the vector of covariates for the <italic>i</italic>th subject, and <bold>
<italic>&#x3b3;</italic>
</bold> is the vector of the corresponding regression coefficients.</p>
<p>The zero-inflated Negative Binomial (ZINB) regression is used for count data that exhibits overdispersion and excess zeros. The data distribution combines the NB distribution and the logit distribution. The model can be expressed as<disp-formula id="e3_6">
<mml:math id="m6">
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="cases">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>if</mml:mtext>
<mml:mspace width="0.28em"/>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo>;</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>if</mml:mtext>
<mml:mspace width="0.28em"/>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1,2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(3.6)</label>
</disp-formula>where <italic>g</italic> (<italic>y</italic>
<sub>
<italic>i</italic>
</sub>) &#x3d; <italic>P</italic> (<italic>y</italic>
<sub>
<italic>i</italic>
</sub> &#x3d; <italic>k</italic>) is defined in (3.3), and the link functions are the same with the ZIP model in (3.2) and&#x20;(3.5).</p>
</sec>
<sec id="s4">
<title>4 Statistical Analysis</title>
<p>In this section, we fit the models on the outage data set and draw conclusions on the basis of these fitted models. The Akaike Information Criterion (AIC [<xref ref-type="bibr" rid="B9">9</xref>]) and Bayesian Information Criterion (BIC [<xref ref-type="bibr" rid="B10">10</xref>]) are utilized to compare models. The AIC of a fitted model is defined as<disp-formula id="e4_1">
<mml:math id="m7">
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mi mathvariant="normal">I</mml:mi>
<mml:mi mathvariant="normal">C</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>&#x2217;</mml:mo>
<mml:mi>ln</mml:mi>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>&#x2217;</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
</mml:math>
<label>(4.1)</label>
</disp-formula>and the BIC is expressed as<disp-formula id="e4_2">
<mml:math id="m8">
<mml:mi mathvariant="normal">B</mml:mi>
<mml:mi mathvariant="normal">I</mml:mi>
<mml:mi mathvariant="normal">C</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>&#x2217;</mml:mo>
<mml:mi>ln</mml:mi>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>log</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2217;</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
</mml:math>
<label>(4.2)</label>
</disp-formula>where <inline-formula id="inf1">
<mml:math id="m9">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> is the maximum value of the likelihood function for the model, <italic>k</italic> is the number of estimated parameters in the model, <italic>n</italic> is the sample size. In comparing models, a smaller AIC and BIC is favorable.</p>
<sec id="s4-1">
<title>4.1 Total Outages v.s. Cumulative Growth Years</title>
<p>First, we focus on the relationship between the total vegetation-caused outages on each circuit and its cumulative vegetative growth year. Over the 5-years of reliability data on the current 7&#x20;years maintenance cycle, there are eight possible trimming patterns for each of the circuit. For example, for all circuits last maintained in 2012, the years of growth after the trimming for 2012&#x2013;2016 are 0,1,2,3,4, respectively, therefore are assigned pattern 1 and this pattern has a cumulative years of growth 0 &#x2b; 1&#x20;&#x2b; 2&#x20;&#x2b; 3&#x20;&#x2b; 4&#x20;&#x3d; 10. Likewise, the rest of the patterns look like in <xref ref-type="table" rid="T1">Table&#x20;1</xref>. Assuming all the circuits are independent, using the total outages on a circuit as the response and the cumulative years of vegetative growth as the predictor is to fulfill the independence assumption of classical regression models. In <xref ref-type="fig" rid="F2">Figure&#x20;2</xref> we can observe an overall increasing trend in the mean outages as cumulative years&#x20;grow.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Trimming patterns and cumulative growth&#x20;years.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Pattern</th>
<th align="center">2012</th>
<th align="center">2013</th>
<th align="center">2014</th>
<th align="center">2015</th>
<th align="center">2016</th>
<th align="center">Cumulative years</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">1</td>
<td align="center">0</td>
<td align="center">1</td>
<td align="center">2</td>
<td align="center">3</td>
<td align="center">4</td>
<td align="center">10</td>
</tr>
<tr>
<td align="left">2</td>
<td align="center">1</td>
<td align="center">2</td>
<td align="center">3</td>
<td align="center">4</td>
<td align="center">5</td>
<td align="center">15</td>
</tr>
<tr>
<td align="left">3</td>
<td align="center">2</td>
<td align="center">3</td>
<td align="center">4</td>
<td align="center">5</td>
<td align="center">6</td>
<td align="center">20</td>
</tr>
<tr>
<td align="left">4</td>
<td align="center">3</td>
<td align="center">4</td>
<td align="center">5</td>
<td align="center">6</td>
<td align="center">7</td>
<td align="center">25</td>
</tr>
<tr>
<td align="left">5</td>
<td align="center">4</td>
<td align="center">5</td>
<td align="center">6</td>
<td align="center">7</td>
<td align="center">0</td>
<td align="center">22</td>
</tr>
<tr>
<td align="left">6</td>
<td align="center">5</td>
<td align="center">6</td>
<td align="center">7</td>
<td align="center">0</td>
<td align="center">1</td>
<td align="center">19</td>
</tr>
<tr>
<td align="left">7</td>
<td align="center">6</td>
<td align="center">7</td>
<td align="center">0</td>
<td align="center">1</td>
<td align="center">2</td>
<td align="center">16</td>
</tr>
<tr>
<td align="left">8</td>
<td align="center">7</td>
<td align="center">0</td>
<td align="center">1</td>
<td align="center">2</td>
<td align="center">3</td>
<td align="center">13</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Mean outages over cumulative&#x20;years.</p>
</caption>
<graphic xlink:href="fams-07-769355-g002.tif"/>
</fig>
<p>For each circuit <italic>i</italic>, <italic>i</italic>&#x20;&#x3d; 1, &#x2026; , 144, we sum the total number of vegetation-caused incidents on circuit <italic>i</italic> in the 5-years period and define the vector the response variable <italic>y</italic>
<sub>
<italic>i</italic>
</sub>, and let the predictor <italic>x</italic>
<sub>
<italic>i</italic>
</sub> be the cumulative years of growth from 2012 to 2016. We then apply the four regression models discussed in <xref ref-type="sec" rid="s3">Section 3</xref> on this data. In order to fit for the ZIP and ZINB models, we subtract the number of outages by 1 so that to have a majority of 0&#x2019;s in the data. Statistical software <monospace>R</monospace> is used to facilitate this analysis. The &#x201c;glm&#x201d; and &#x201c;glm.nb&#x201d; functions are implemented to fit the Poisson and NB models, both of which are in the package &#x201c;MASS&#x201d;, and the &#x201c;zeroinfl&#x201d; function within the package &#x201c;pscl&#x201d; is applied to fit the ZIP and ZINB models. The regression parameter estimates and <italic>p</italic>-values, as well as the values of AIC and BIC from the four models are reported in <xref ref-type="table" rid="T2">Table&#x20;2</xref>.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>The regression parameter estimates, <italic>p</italic>-values, values of AIC and BIC from the Poisson, NB, ZIP, and ZINB models.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Model</th>
<th align="center">Estimate</th>
<th align="center">
<italic>p</italic>-value</th>
<th align="center">AIC</th>
<th align="center">BIC</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Poisson</td>
<td align="char" char=".">0.046</td>
<td align="char" char=".">0.001</td>
<td align="char" char=".">662.406</td>
<td align="char" char=".">668.346</td>
</tr>
<tr>
<td align="left">NB</td>
<td align="char" char=".">0.051</td>
<td align="char" char=".">0.042</td>
<td align="char" char=".">551.112</td>
<td align="char" char=".">560.021</td>
</tr>
<tr>
<td align="left">ZIP</td>
<td align="char" char=".">0.019</td>
<td align="char" char=".">0.202</td>
<td align="char" char=".">606.478</td>
<td align="char" char=".">618.357</td>
</tr>
<tr>
<td align="left">ZINB</td>
<td align="char" char=".">0.035</td>
<td align="char" char=".">0.226</td>
<td align="char" char=".">553.482</td>
<td align="char" char=".">568.331</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Overall, NB-type models fit the data substantially better than Poisson models do. It is not surprising as the response variable is overdispersed with a variance of 6.216 and a mean of 1.993. Based on the values of AIC and BIC, we find that the NB model fits the data the best. The results demonstrate that vegetative growth has a significant positive effect on vegetation-caused outages.</p>
<p>To decide the optimal cycle length, we introduce some indicator variables <italic>p</italic>
<sub>
<italic>j</italic>
</sub> as follows,<disp-formula id="equ1">
<mml:math id="m10">
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="cases">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mn>1</mml:mn>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mspace width="1em"/>
<mml:mtext>if&#x2009;a&#x2009;trimming&#x2009;pattern&#x2009;in&#x2009;Table</mml:mtext>
<mml:mn>1</mml:mn>
<mml:mtext>&#x2009;includes&#x2009;vegetative&#x2009;growth&#x2009;year</mml:mtext>
<mml:mi>j</mml:mi>
<mml:mo>;</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mn>0</mml:mn>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mspace width="1em"/>
<mml:mtext>otherwise</mml:mtext>
<mml:mo>.</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:math>
</disp-formula>Along the 7&#xa0;years trimming cycle, <italic>j</italic>&#x20;&#x3d; 0, &#x2026; , 7. For example, growth year of 7 is included in trimming patterns 4, 5, 6, 7, 8, thus<disp-formula id="equ2">
<mml:math id="m11">
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>7</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="cases">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mn>1</mml:mn>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mspace width="1em"/>
<mml:mtext>if&#x2009;a&#x2009;circuit&#x2009;has&#x2009;a&#x2009;pattern&#x2009;</mml:mtext>
<mml:mn>4</mml:mn>
<mml:mtext>,&#x2009;</mml:mtext>
<mml:mn>5</mml:mn>
<mml:mtext>,&#x2009;</mml:mtext>
<mml:mn>6</mml:mn>
<mml:mtext>,&#x2009;</mml:mtext>
<mml:mn>7</mml:mn>
<mml:mtext>,&#x2009;or&#x2009;</mml:mtext>
<mml:mn>8</mml:mn>
<mml:mtext>;</mml:mtext>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mn>0</mml:mn>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mspace width="1em"/>
<mml:mtext>if&#x2009;a&#x2009;circuit&#x2009;has&#x2009;a&#x2009;pattern&#x2009;</mml:mtext>
<mml:mn>1</mml:mn>
<mml:mtext>,&#x2009;</mml:mtext>
<mml:mn>2</mml:mn>
<mml:mtext>,&#x2009;or&#x2009;</mml:mtext>
<mml:mn>3</mml:mn>
<mml:mo>.</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:math>
</disp-formula>
</p>
<p>If <italic>p</italic>
<sub>7</sub> is a significant predictor on outages, it means the patterns involving year 7 have significant more outages than other patterns, in other words, growth year 7 tends to have more outages than other years. We then fit the NB model with the response <italic>y</italic>
<sub>
<italic>i</italic>
</sub> on each <italic>p</italic>
<sub>
<italic>j</italic>
</sub>, individually. The results are included in <xref ref-type="table" rid="T3">Table&#x20;3</xref>. <xref ref-type="table" rid="T3">Table&#x20;3</xref> shows that growth year 7 tends to have significantly more outages than other years, therefore the power company should consider shortening the current 7&#xa0;years trimming cycle to a 6&#xa0;years cycle, in order to decrease vegetation-caused outages.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>The regression parameter estimates and <italic>p</italic>-values for the indicator variables.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Indicator</th>
<th align="center">Estimate</th>
<th align="center">
<italic>p</italic>-value</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">
<italic>p</italic>
<sub>0</sub>
</td>
<td align="char" char=".">0.201</td>
<td align="char" char=".">0.400</td>
</tr>
<tr>
<td align="left">
<italic>p</italic>
<sub>1</sub>
</td>
<td align="char" char=".">&#x2212;0.075</td>
<td align="char" char=".">0.764</td>
</tr>
<tr>
<td align="left">
<italic>p</italic>
<sub>2</sub>
</td>
<td align="char" char=".">&#x2212;0.403</td>
<td align="char" char=".">0.118</td>
</tr>
<tr>
<td align="left">
<italic>p</italic>
<sub>3</sub>
</td>
<td align="char" char=".">&#x2212;0.481</td>
<td align="char" char=".">0.054</td>
</tr>
<tr>
<td align="left">
<italic>p</italic>
<sub>4</sub>
</td>
<td align="char" char=".">&#x2212;0.212</td>
<td align="char" char=".">0.417</td>
</tr>
<tr>
<td align="left">
<italic>p</italic>
<sub>5</sub>
</td>
<td align="char" char=".">0.221</td>
<td align="char" char=".">0.423</td>
</tr>
<tr>
<td align="left">
<italic>p</italic>
<sub>6</sub>
</td>
<td align="char" char=".">0.205</td>
<td align="char" char=".">0.336</td>
</tr>
<tr>
<td align="left">
<italic>p</italic>
<sub>7</sub>
</td>
<td align="char" char=".">0.431</td>
<td align="char" char=".">0.040</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4-2">
<title>4.2 Outages v.s. Vegetative Growth Years</title>
<p>In this section, we will quantify the relationship between the number vegetative-caused outages in each year on each circuit and the number of years of vegetative growth when the outage occurred. If we consider each circuit as a cluster, thus the data has a clustered structure and the outages on the same circuit across different years are correlated. Statistical models introduced in <xref ref-type="sec" rid="s3">Section 3</xref> can not be directly employed to clustered data since they all assume the observations are independent. The violation of the independence assumption might result in misleading conclusions. The mixed effects models treat clustered data adequately and assumes two sources of variation, within cluster and between clusters. Two types of coefficients are distinguished in the mixed model: fixed effects and random effects (or cluster specific effects). The fixed effects have the same meaning as in classical statistics, the random effects are random and are estimated as posterior means&#x20;[<xref ref-type="bibr" rid="B11">11</xref>].</p>
<p>For count data specially, a generalized linear mixed model, i.e.,&#x20;a Poisson generalized linear mixed (GLM) model with a random intercept, is conducted for this analysis. This model assumes that the conditional distribution of the count data is the Poisson distribution as in (3.2), but it adds an extra random term in the link function as shown in (4.3),<disp-formula id="e4_3">
<mml:math id="m12">
<mml:mi>ln</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi>&#x3b2;</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
</mml:math>
<label>(4.3)</label>
</disp-formula>where <italic>b</italic>
<sub>
<italic>i</italic>
</sub> is the random effect in cluster <italic>i</italic> and represents circuit-specific variability. By accommodating both fixed and random effects, this model provides an effective and flexible way of representing the mean as well as the covariance structure of the data. Similarly, we can have the GLM variant for NB, ZIP, ZINB models in Section&#x20;3.1.</p>
<p>We then define the response variable <italic>y</italic>
<sub>
<italic>ij</italic>
</sub> as the total number of vegetative-caused outages incidents on circuit <italic>i</italic> on growth year <italic>j</italic>, and let the fixed effect <italic>x</italic>
<sub>
<italic>ij</italic>
</sub> be the growth year for circuit <italic>i</italic> on year <italic>j</italic>, the circuit effect is consider random in the mixed model. For each circuit, we only focus on those years having vegetative-caused outages, thus <italic>y</italic>
<sub>
<italic>ij</italic>
</sub> &#x3e; 0 for all (<italic>i</italic>, <italic>j</italic>). This generates a data set with 306 observations from 144 circuits. <xref ref-type="fig" rid="F3">Figure&#x20;3</xref> shows an increasing trend of the mean vegetative-caused outages over growth&#x20;years.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Mean vegetative-caused outages on different growth&#x20;years.</p>
</caption>
<graphic xlink:href="fams-07-769355-g003.tif"/>
</fig>
<p>In order to fit for the ZIP and ZINB GLM models, we subtract the number of outages by 1, so that we have a majority of 0&#x2019;s in the data. The &#x201c;glmmPQL&#x201d; function in package &#x201c;MASS&#x201d;, &#x201c;glmer.nb&#x201d; function in package &#x201c;lme4&#x201d;, &#x201c;glmmTMB&#x201d; function in package &#x201c;glmmTMB&#x201d;, and &#x201c;gam&#x201d; function in package &#x201c;mgcv&#x201d;, are implemented to fit the GLM models for Poisson, NB, ZIP, and ZINB, respectively. The fixed effect estimates and <italic>p</italic>-values, as well as the values of AIC and BIC from the four models are reported in <xref ref-type="table" rid="T4">Table&#x20;4</xref>.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>The fixed effect estimates, <italic>p</italic>-values, values of AIC and BIC from the Poisson, NB, ZIP, and ZINB GLM models.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Mixed model</th>
<th align="center">Estimate</th>
<th align="center">
<italic>p</italic>-value</th>
<th align="center">AIC</th>
<th align="center">BIC</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Poisson</td>
<td align="char" char=".">0.097</td>
<td align="char" char=".">0.022</td>
<td align="char" char=".">524.800</td>
<td align="char" char=".">536.000</td>
</tr>
<tr>
<td align="left">NB</td>
<td align="char" char=".">0.097</td>
<td align="char" char=".">0.039</td>
<td align="char" char=".">518.097</td>
<td align="char" char=".">532.991</td>
</tr>
<tr>
<td align="left">ZIP</td>
<td align="char" char=".">0.096</td>
<td align="char" char=".">0.026</td>
<td align="char" char=".">524.831</td>
<td align="char" char=".">536.002</td>
</tr>
<tr>
<td align="left">ZINB</td>
<td align="char" char=".">0.114</td>
<td align="char" char=".">0.018</td>
<td align="char" char=".">506.765</td>
<td align="char" char=".">521.632</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The estimates from different models in <xref ref-type="table" rid="T4">Table&#x20;4</xref> are fairly close and are all have a small <italic>p</italic>-value, thus we can conclude that the vegetative growth has significant effect on the number of outages. Based on the ZINB model which has the smallest AIC and BIC, we estimate that as the growth year increases by 1, the expected number of outages increases by exp (0.114) &#x3d; 1.12 or about&#x20;12%.</p>
<p>To decide the optimal cycle length, we define several indicator variables <italic>g</italic>
<sub>
<italic>k</italic>
</sub> as follows,<disp-formula id="equ3">
<mml:math id="m13">
<mml:msub>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="cases">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mn>1</mml:mn>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mspace width="1em"/>
<mml:mtext>if&#x2009;the&#x2009;outage&#x2009;occurred&#x2009;at&#x2009;growth&#x2009;year</mml:mtext>
<mml:mi>k</mml:mi>
<mml:mo>;</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mn>0</mml:mn>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mspace width="1em"/>
<mml:mtext>otherwise</mml:mtext>
<mml:mo>,</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:math>
</disp-formula>where <italic>k</italic>&#x20;&#x3d; 0, &#x2026; , 7. For example,<disp-formula id="equ4">
<mml:math id="m14">
<mml:msub>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>7</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="cases">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mn>1</mml:mn>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mspace width="1em"/>
<mml:mtext>if&#x2009;the&#x2009;outage&#x2009;occurred&#x2009;at&#x2009;growth&#x2009;year&#x2009;</mml:mtext>
<mml:mn>7</mml:mn>
<mml:mtext>;</mml:mtext>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mn>0</mml:mn>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mspace width="1em"/>
<mml:mtext>if&#x2009;the&#x2009;outage&#x2009;occurred&#x2009;at&#x2009;growth&#x2009;year&#x2009;0,&#x2009;</mml:mtext>
<mml:mn>1</mml:mn>
<mml:mtext>,&#x2009;</mml:mtext>
<mml:mn>2</mml:mn>
<mml:mtext>,&#x2009;</mml:mtext>
<mml:mn>3</mml:mn>
<mml:mtext>,&#x2009;</mml:mtext>
<mml:mn>4</mml:mn>
<mml:mtext>,&#x2009;</mml:mtext>
<mml:mn>5</mml:mn>
<mml:mtext>,&#x2009;or&#x2009;</mml:mtext>
<mml:mn>6</mml:mn>
<mml:mo>.</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:math>
</disp-formula>If <italic>g</italic>
<sub>7</sub> is a relevant predictor, it would suggest that growth year 7 tends to have more outages than other years. We then fit the ZINB GLM model with the response <italic>y</italic>
<sub>
<italic>ij</italic>
</sub> on each <italic>g</italic>
<sub>
<italic>k</italic>
</sub>, individually. The results are shown in <xref ref-type="table" rid="T5">Table&#x20;5</xref>. We observe from <xref ref-type="table" rid="T5">Table&#x20;5</xref> that both growth years 6 and 7 have more outages than other years. Notice that the conclusions drawn from <xref ref-type="sec" rid="s4-1">Section 4.1</xref> and <xref ref-type="sec" rid="s4-2">Section 4.2</xref> are consistent, both of which show that there is a significant statistical relationship between the reliability of the electric power distribution system and the number of years of vegetative growth on distribution circuits, and advocate a 6&#xa0;years cycle as the reasonable vegetation maintenance schedule.</p>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>The GLM regression model parameter estimates and <italic>p</italic>-values for the indicator variables.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Predictor</th>
<th align="center">Coefficient</th>
<th align="center">
<italic>p</italic>-value</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">
<italic>g</italic>
<sub>0</sub>
</td>
<td align="char" char=".">&#x2212;0.042</td>
<td align="char" char=".">0.901</td>
</tr>
<tr>
<td align="left">
<italic>g</italic>
<sub>1</sub>
</td>
<td align="char" char=".">&#x2212;0.0092</td>
<td align="char" char=".">0.767</td>
</tr>
<tr>
<td align="left">
<italic>g</italic>
<sub>2</sub>
</td>
<td align="char" char=".">&#x2212;0.571</td>
<td align="char" char=".">0.137</td>
</tr>
<tr>
<td align="left">
<italic>g</italic>
<sub>3</sub>
</td>
<td align="char" char=".">&#x2212;0.649</td>
<td align="char" char=".">0.089</td>
</tr>
<tr>
<td align="left">
<italic>g</italic>
<sub>4</sub>
</td>
<td align="char" char=".">&#x2212;0.124</td>
<td align="char" char=".">0.682</td>
</tr>
<tr>
<td align="left">
<italic>g</italic>
<sub>5</sub>
</td>
<td align="char" char=".">&#x2212;0.238</td>
<td align="char" char=".">0.462</td>
</tr>
<tr>
<td align="left">
<italic>g</italic>
<sub>6</sub>
</td>
<td align="char" char=".">0.791</td>
<td align="char" char=".">0.004</td>
</tr>
<tr>
<td align="left">
<italic>g</italic>
<sub>7</sub>
</td>
<td align="char" char=".">0.429</td>
<td align="char" char=".">0.062</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s5">
<title>5 Concluding Remarks</title>
<p>This paper investigates the role of vegetative growth on distribution reliability. Some statistical models including Poisson, Negative Binomial, Zero-Inflated models and their variants are utilized. Based on this study, the following conclusions can be drawn.<list list-type="simple">
<list-item>
<p>1) There is a significant statistical relationship between the reliability of the electric power distribution system and the number of years of vegetative growth on distribution circuits.</p>
</list-item>
<list-item>
<p>2) An optimal vegetation maintenance should be scheduled every 6&#xa0;years.</p>
</list-item>
<list-item>
<p>3) The Negative Binomial model and its modifications are particularly effective at fitting vegetation-caused outages.</p>
</list-item>
</list>
</p>
<p>The study can advance the understanding on the relationship between vegetative growth and system reliability and should be useful to help the power companies determine an effective vegetation maintenance schedule. However, there are three main limitations of the models and future research could be conducted to obtain more accurate results.</p>
<p>First, we limited the attention on the effect of vegetative growth on power system reliability, other possible factors are not included due to lack of access to related data. As a result, the performance of the analysis may be improved by the inclusion of additional climate and geographical information. The second limitation lies on the fact that the models are fit with data from only normal (storm exclusion) operating conditions, which means that our models might underestimate the benefits of tree pruning. Additional data and study would be needed to learn the impacts of tree maintenance on distribution system reliability under storm conditions. The third limitation is that the analysis is based on data from only one company, consequently the model may not fit the data from another power company, especially if the region where the tree types, temperature, population density, wind regimes, and precipitation patterns is substantially different from Midwestern United&#x20;States. If more data are available, it will further test the effectiveness of the statistical models, and enhance insights on the importance of vegetative maintenance in improving power system reliability.</p>
</sec>
</body>
<back>
<sec id="s6">
<title>Data Availability Statement</title>
<p>The datasets presented in this article are not readily available because company confidential. Requests to access the datasets should be directed to pan@rowan.edu.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>JP developed the method, conducted the numerical experiments, and wrote the manuscript.</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of Interest</title>
<p>The author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<label>1.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Radmer</surname>
<given-names>DT</given-names>
</name>
<name>
<surname>Kuntz</surname>
<given-names>PA</given-names>
</name>
<name>
<surname>Christie</surname>
<given-names>RD</given-names>
</name>
<name>
<surname>Venkata</surname>
<given-names>SS</given-names>
</name>
<name>
<surname>Fletcher</surname>
<given-names>RH</given-names>
</name>
</person-group>. <article-title>Predicting vegetation-related failure rates for overhead distribution feeders</article-title>. <source>IEEE Trans Power Deliv</source> (<year>2002</year>) <volume>17</volume>:<fpage>4</fpage>. <pub-id pub-id-type="doi">10.1109/tpwrd.2002.804006</pub-id> </citation>
</ref>
<ref id="B2">
<label>2.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lovlace</surname>
<given-names>WR</given-names>
</name>
</person-group>. <article-title>Vegetation Management on Distribution Line Right-of-Way Are You Getting Top Value for Your Money</article-title>. <source>Proc Rural Electric Power Conf</source> (<year>1996</year>): <fpage>B5</fpage>. <pub-id pub-id-type="doi">10.1109/REPCON.1996.495241</pub-id> </citation>
</ref>
<ref id="B3">
<label>3.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Doostan</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Sohrabi</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Chowdhury</surname>
<given-names>B</given-names>
</name>
</person-group>. <article-title>A data-driven approach for predicting vegetation-related outages in power distribution systems</article-title>. <source>Int Trans Electr Energ Syst.</source> (<year>2020</year>) <volume>30</volume>:<fpage>1</fpage>. <pub-id pub-id-type="doi">10.1002/2050-7038.12154</pub-id> </citation>
</ref>
<ref id="B4">
<label>4.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guikema</surname>
<given-names>SD</given-names>
</name>
<name>
<surname>Davidson</surname>
<given-names>RA</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>H</given-names>
</name>
</person-group>. <article-title>Statistical models of the effects of tree trimming on power system outages</article-title>. <source>IEEE Trans Power Deliv</source> (<year>2006</year>) <volume>21</volume>:<fpage>3</fpage>. <pub-id pub-id-type="doi">10.1109/tpwrd.2005.860238</pub-id> </citation>
</ref>
<ref id="B5">
<label>5.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kuntz</surname>
<given-names>PA</given-names>
</name>
<name>
<surname>Christie</surname>
<given-names>RD</given-names>
</name>
<name>
<surname>Venkata</surname>
<given-names>SS</given-names>
</name>
</person-group>. <article-title>Optimal vegetation maintenance scheduling of overhead electric power distribution systems</article-title>. <source>IEEE Trans Power Deliv</source> (<year>2002</year>) <volume>17</volume>:<fpage>4</fpage>. <pub-id pub-id-type="doi">10.1109/tpwrd.2002.804007</pub-id> </citation>
</ref>
<ref id="B6">
<label>6.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Du</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Park</surname>
<given-names>YT</given-names>
</name>
<name>
<surname>Theera-Ampornpunt</surname>
<given-names>N</given-names>
</name>
<name>
<surname>McCullough</surname>
<given-names>JS</given-names>
</name>
<name>
<surname>Speedie</surname>
<given-names>SM</given-names>
</name>
</person-group>. <article-title>The use of count data models in biomedical informatics evaluation research</article-title>. <source>J&#x20;Am Med Inform Assoc</source> (<year>2012</year>) <volume>19</volume>:<fpage>39</fpage>&#x2013;<lpage>44</lpage>. <pub-id pub-id-type="doi">10.1136/amiajnl-2011-000256</pub-id> </citation>
</ref>
<ref id="B7">
<label>7.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Cameron</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Trivedi</surname>
<given-names>P</given-names>
</name>
</person-group>. <source>Regression Analysis of Count Data</source>. <edition>&#x201d; 2nd ed.</edition> <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name> (<year>1998</year>). </citation>
</ref>
<ref id="B8">
<label>8.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lambert</surname>
<given-names>D</given-names>
</name>
</person-group>. <article-title>Zero-inflated Poisson Regression, with an Application to Defects in Manufacturing</article-title>. <source>Technometrics</source> (<year>1992</year>) <volume>34</volume>:<fpage>1</fpage>&#x2013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.2307/1269547</pub-id> </citation>
</ref>
<ref id="B9">
<label>9.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Akaike</surname>
<given-names>H</given-names>
</name>
</person-group>. <article-title>A new look at the statistical model identification</article-title>. <source>IEEE Trans Automat Contr</source> (<year>1974</year>) <volume>19</volume>(<issue>6</issue>):<fpage>716</fpage>&#x2013;<lpage>23</lpage>. <pub-id pub-id-type="doi">10.1109/tac.1974.1100705</pub-id> </citation>
</ref>
<ref id="B10">
<label>10.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schwarz</surname>
<given-names>G</given-names>
</name>
</person-group>. <article-title>Estimating the dimension of a model</article-title>. <source>Ann Stat</source> (<year>1978</year>) <volume>6</volume>:<fpage>461</fpage>&#x2013;<lpage>4</lpage>. <pub-id pub-id-type="doi">10.1214/aos/1176344136</pub-id> </citation>
</ref>
<ref id="B11">
<label>11.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Laird</surname>
<given-names>NM</given-names>
</name>
<name>
<surname>Ware</surname>
<given-names>JH</given-names>
</name>
<name>
<surname>Random</surname>
<given-names>&#x201c;</given-names>
</name>
</person-group>. <article-title>Random-effects models for longitudinal data</article-title>. <source>Biometrics</source> (<year>1982</year>) <volume>38</volume>:<fpage>963</fpage>&#x2013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.2307/2529876</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>