<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Energy Res.</journal-id>
<journal-title>Frontiers in Energy Research</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Energy Res.</abbrev-journal-title>
<issn pub-type="epub">2296-598X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">730058</article-id>
<article-id pub-id-type="doi">10.3389/fenrg.2021.730058</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Energy Research</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>An Association Rules-Based Method for Outliers Cleaning of Measurement Data in the Distribution Network</article-title>
<alt-title alt-title-type="left-running-head">Kuang et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">AR-Based Method for Outliers Cleaning</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Kuang</surname>
<given-names>Hua</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Qin</surname>
<given-names>Risheng</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1303863/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>He</surname>
<given-names>Mi</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>He</surname>
<given-names>Xin</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Duan</surname>
<given-names>Ruimin</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Guo</surname>
<given-names>Cheng</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Meng</surname>
<given-names>Xian</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>Yunnan Power Grid Co., Ltd., <addr-line>Kunming</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>Electric Power Research Institute of Yunnan Power Grid Company Ltd., <addr-line>Kunming</addr-line>, <country>China</country>
</aff>
<aff id="aff3">
<label>
<sup>3</sup>
</label>Kunming Power Supply Company of Yunnan Power Grid Co., Ltd., <addr-line>Kunming</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1115169/overview">Qiuye Sun</ext-link>, Northeastern University, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1435665/overview">Zhijun Qin</ext-link>, Guangxi University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1435705/overview">Shunbo Lei</ext-link>, University of Michigan, United&#x20;States</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1121823/overview">Xuguang Hu</ext-link>, Northeastern University, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Risheng Qin, <email>qinrisheng2020@126.com</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Smart Grids, a section of the journal Frontiers in Energy Research</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>13</day>
<month>10</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>9</volume>
<elocation-id>730058</elocation-id>
<history>
<date date-type="received">
<day>24</day>
<month>06</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>27</day>
<month>09</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Kuang, Qin, He, He, Duan, Guo and Meng.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Kuang, Qin, He, He, Duan, Guo and Meng</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>For any power system, the reliability of measurement data is essential in operation, management and also in planning. However, it is inevitable that the measurement data are prone to outliers, which may impact the results of data-based applications. In order to improve the data quality, the outliers cleaning method for measurement data in the distribution network is studied in this paper. The method is based on a set of association rules (AR) that are automatically generated form historical measurement data. First, the association rules are mining in conjunction with the density-based spatial clustering of application with noise (DBSCAN), k-means and Apriori technique to detect outliers. Then, for the outliers repairing process after outliers detection, the proposed method uses a distance-based model to calculate the repairing cost of outliers, which describes the similarity between outlier and normal data. Besides, the Mahalanobis distance is employed in the repairing cost function to reduce the errors, which could implement precise outliers cleaning of measurement data in the distribution network. The test results for the simulated datasets with artificial errors verify that the superiority of the proposed outliers cleaning method for outliers detection and repairing.</p>
</abstract>
<kwd-group>
<kwd>association rules</kwd>
<kwd>outliers cleaning</kwd>
<kwd>outliers detection</kwd>
<kwd>outliers repairing</kwd>
<kwd>measurement data</kwd>
<kwd>distribution network</kwd>
</kwd-group>
<contract-num rid="cn001">YNKJXM20191369</contract-num>
<contract-sponsor id="cn001">China Southern Power Grid<named-content content-type="fundref-id">10.13039/501100005311</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>With the evolution of smart grids, the intelligent monitoring equipment and system are becoming an integral component of the distribution network, collecting a substantial volume of data in order to manage the status and provide timely updates in the network (<xref ref-type="bibr" rid="B1">Alimardani et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B28">Wang et&#x20;al., 2018</xref>). Among them, the supervisory control and data acquisition (SCADA) system provides a large number of operation data and analysis results, which brings great convenience for operators to evaluate the planning and operation of distribution system. For instance, the data structure is complex, many types of the data, and the sampling period/frequency of data are also different. For distribution network dispatching control system, poor quality data may lead to wrong decisions, which will have a great impact on the stable operation of power grid. Hence, it is essential that to clean the outliers of measurement data in the distribution network.</p>
<p>The distribution network is an important part of production, transmission, and consumption, which plays a critical role in the delivery of electric power. In the planning and operation of distribution network, the availability of accurate measurement data has a considerable impact on dispatching operations and control of the distribution network. For instance, the analysis of measurement data in the distribution network can assist in taking action against fault detection, dispatching, load forecasting, power quality, tariff settings, and so forth. (<xref ref-type="bibr" rid="B8">Hayes et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B2">Cai et&#x20;al., 2021</xref>; <xref ref-type="bibr" rid="B26">Wang et&#x20;al., 2019</xref>). Moreover, it solves the problems that distribution networks frequently face in terms of integrated energy planning, distributed energy storage, and demand-side management, respectively (<xref ref-type="bibr" rid="B23">Thams et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B12">Liu et&#x20;al., 2019</xref>). Generally, the majority of the researches in the distribution network, for the analysis and prediction of the measurement data, is focus on the feature selection or parametric optimization of the model (<xref ref-type="bibr" rid="B11">Liu et&#x20;al., 2020</xref>). However, due to the complex topology features and communication disturbances, the accuracy of distribution network measurement data is not always satisfactory, making it susceptible to data anomalies such as outliers or missing data (<xref ref-type="bibr" rid="B21">Shi et&#x20;al., 2019</xref>). To fill in missing data, denoise while detecting outliers, and repair inconsistencies, data cleaning is the first and most crucial step. Obviously, it has a decisive influence on the final result: if the dataset is incomplete in terms of data cleaning and preprocessing. This means that the established analysis and prediction model will not be accurate and efficient, which may no longer be suitable for the planning and operation of distribution system. For instance, due to external disturbances, data recorded in smart electric meters is abruptly modified because a transmission error for control commands, such as electric quantity or associated parametric information is reset to outliers, or even data missing (<xref ref-type="bibr" rid="B16">Nascimento et&#x20;al., 2012</xref>). And in DC microgrids, the large-scale converters with inhomogeneous initial values are widely appeared due to soft-starting operation, which make the input-output maps error will be large (<xref ref-type="bibr" rid="B27">Wang et&#x20;al., 2021</xref>). Under these circumstances, efficient preprocessing via data cleaning aids in improving the quality and accuracy of subsequent analysis and decision-making outcomes, which can successfully guide the planning and operation of the distribution network.</p>
<p>Researchers have extensively conducted many outliers cleaning studies to improve the data quality and decision-making results, including outliers detection and repairing. For outliers detection, with the rapid development of machine learning technology, many machine learning algorithms have been utilized to improve the accuracy in power systems. In literature (<xref ref-type="bibr" rid="B17">Nemati et&#x20;al., 2018</xref>), a constraint and association rule-based current transmission capability forecasting method was proposed for outliers detection in substation metering equipment. However, this model is complex and computationally intensive, which is not suitable for the detection of bad data in a large number of transformer districts. In literature (<xref ref-type="bibr" rid="B7">Esmalifalak et&#x20;al., 2014</xref>), support vector machine (SVM) has been investigated for detecting the outliers injected into the measurement data from power grid. Since SVM is a supervised learning method, it necessitates labeling the data in order to train the model. However, in practice, obtaining a considerable volume of tagged data is difficult. In literature (<xref ref-type="bibr" rid="B24">Thang et&#x20;al., 2011</xref>), a density-based DBSCAN algorithm was used for detecting the network traffic outliers of electricity meters, which dataset may include multiple traffic types with different characteristics. It has a high level of outliers detection performance, but there are difficulties in finding its parameters (epsilon and minpts) when the multidimensional feature data is taken into account. In literature (<xref ref-type="bibr" rid="B10">Li et&#x20;al., 2018</xref>), the isolation forest (IF) algorithm was proposed to detect the outliers, and the backpropagation neural network (BPNN) algorithm was used for predicting and repairing the outliers. However, IF algorithm is usually suitable for detecting global outliers, but not for detecting local outliers.</p>
<p>Traditionally, researchers have concentrated more on a basic and easy to repair statistical estimate method for outliers repairing. Still, mining a deep relationship between data is difficult, and the repairing results are not ideal (<xref ref-type="bibr" rid="B25">Waal et&#x20;al., 2001</xref>). By contrast, machine learning (also includes deep learning) methods is a very effective technology, which could easily recognize the outliers through the linear or nonlinear pattern relationships and the repairing results could more accurately. For instance, in literature (<xref ref-type="bibr" rid="B19">Qu et&#x20;al., 2016</xref>), a hierarchical clustering algorithm based on the clustering using representatives (CURE) was proposed for the outliers detection the repairing, which could confirm the normal value boundary samples from historical data. However, when the volume of data is large, the time complexity is poor and precision is low with the hierarchical clustering algorithm, which make a challenge to determine the ideal boundary sample number. In literature (<xref ref-type="bibr" rid="B9">Hu et&#x20;al., 2021</xref>) a data recovery method based on generative adversarial networks (GANs) was proposed for safe and efficient operation in the pipeline network, which could accurately recover incomplete pressure data caused by the device or communication aspect. But there are still some difficulties when the complete data pairs is no provided in the training process. On the other hand, to produce good repairing results, a metric learning and a cost functional model are proposed to estimate data repairing efficiency while taking sample distances into account (Li et&#x20;al., 2019). Distance is a term that describes the dissimilarity of two input samples. Among them, the most frequently used technique is the Euclidean distance. However, the Euclidean distance takes neither the correlation of the features nor the different weights of features into account, which may not reflect the real nature of the problem, and distorts the true dissimilarity between samples. To address this issue, in literature (<xref ref-type="bibr" rid="B14">Maesschalck et&#x20;al., 2000</xref>), the Mahalanobis distance idea was defined to use the similarity metric as a substitution to perform better. In another example (<xref ref-type="bibr" rid="B29">Yan et&#x20;al., 2020</xref>), the adoption of the Mahalanobis distance improves the classic k-nearest neighbor (KNN) outliers identification method, resulting in increased accuracy and a lower false detection rate. However, the repair of outliers has not been considered in this model. Furthermore, most of the models stated above focus on specific application scenarios and do not process real-time data from the distribution network system. Moreover, most of the outliers cleaning methods aforementioned presumed the underlying population distribution before the step of data cleaning. However, in real-word data, a hypothesis about an underlying population is a statement that may be true or&#x20;false.</p>
<p>In power grids, a huge amount of historical measurement data from various distribution stations is available, which could provide valuable information for detecting and repairing outliers. Furthermore, the association rules learning is a popular and well data mining method for discovering relations between variable features. Our motivation is to investigate how to capitalize on the historical data for outliers cleaning, including outliers detection and repairing, to achieve expected performance. An association rules-based method for outliers cleaning is given in this work to mine the information which whereas the assumptions of underlying population about the data is not required. In outliers detection, we adopt the density-based spatial clustering of application with noise (DBSCAN), k-means and Apriori technique to generate the association rules. After the outliers detection, the distance-based model is designed with Mahalanobis distance to repair to outliers. Various tests are carried out on data sets with simulated errors to evaluate the good performance of the proposed method. The test results indicate that the proposed method can effectively identify outliers in the distribution network&#x2019;s measurement data while achieving accurate data repairing. The proposed method detects the outliers with a F1-Score (a metric combine precision and recall) of 96%, even in the condition with a high anomaly rate. The F1-Score indicates how well precision and memory are balanced. Furthermore, the correlation between features of measurement data is also computed to detect and repair the outliers, thus improving the method&#x2019;s accuracy. The main contributions of this work could be summarized as follows two aspects:<list list-type="simple">
<list-item>
<p>1) This paper introduces an outlier detection and repairing technique based on association rule. The proposed technique uses the information provided from historical measurement data, whereas the assumptions of underlying distribution about the measurement data is not required.</p>
</list-item>
<list-item>
<p>2) The distance-based model is adopt for outliers repairing, which describes the similarity between outlier and normal data by the Mahalanobis distance. It estimates the outliers according to the normal data within historical data, which is employed to improve the estimation accuracy.</p>
</list-item>
</list>
</p>
<p>The remainder of the paper is organized in the following manner. <italic>Problem Statement</italic> discusses the measurement problems in the distribution network and data anomalies. The proposed methodology has been presented in <italic>Preliminaries</italic> and <italic>The Proposed Model for Outliers Detection and Repairing</italic>. Simulation results are provided in <italic>Experiment and Analysis</italic>, and the concluding remarks are summarized in <italic>Conclusion</italic>.</p>
</sec>
<sec id="s2">
<title>Problem Statement</title>
<sec id="s2-1">
<title>Measurement Problems in Distribution Network</title>
<p>Through SCADA system, a large amount of operating data is continuously collected, uploaded, and formed into big data for the distribution network, which provides abundant data resources for big data analysis (<xref ref-type="bibr" rid="B31">Ye et&#x20;al., 2010</xref>; <xref ref-type="bibr" rid="B22">Song et&#x20;al., 2013</xref>). And the data collected by SCADA has the following characteristics: large amount, high dimensions, and complex data types. Then the most common problems encountered in measurement data are the absence of data (nulls values and zeros), change in level and spikes (points more than N times the standard deviations away from the series mean), and generally are called outliers. Therefore, in order to improve the accuracy of the analysis and decision-making results based on measurement data, how to clean and repair outliers form the measurement data in distribution network is a challenge faced by distribution system.</p>
</sec>
<sec id="s2-2">
<title>The Source of Abnormal Data in Distribution Network</title>
<p>The process of collecting measurement data from the distribution network involves many components such as metering equipment, metering centers, and communication systems. However, if a malfunction occurs in any measurement channel, it can lead to data anomalies (<xref ref-type="bibr" rid="B3">Chen et&#x20;al., 2010</xref>). For example, the failure of smart electricity meters, noise interference, data transmission errors, and abnormal power consumption will cause these collected data to become outliers or data missing. Generally, there are three potential sources of data anomalies in distribution network measurement data:<list list-type="simple">
<list-item>
<p>1) Metering equipment. The measuring equipment from abnormal operating conditions may lead to errors in measurement (<xref ref-type="bibr" rid="B30">Yan et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B3">Chen et&#x20;al., 2010</xref>). In particular, the magnetic bias phenomenon in potential transformers (PT) and current transformers (CT) equipment would cause measurement errors (<xref ref-type="bibr" rid="B15">Mccamish et&#x20;al., 2016</xref>). Also, the non-synchronous problem on data collection could cause errors since the sampling time of some devices is asynchronous (<xref ref-type="bibr" rid="B11">Liu et&#x20;al., 2020</xref>). In particular, all forms of metering and communication equipment are constantly exposed to unknown conditions. They are vulnerable to the effects of real-world circumstances, which typically have a high failure rate. Meanwhile, the operation in the monitoring and communication equipment can not be carried out smoothly when a fault occurs. In that situation, erroneous or missing data will be recorded.</p>
</list-item>
<list-item>
<p>2) Distribution network. Control operations and faults in the distribution network have a significant impact on the accuracy of measurement data. Temporary inrush current interference caused by switchgear such as circuit breakers may cause temporary outliers to appear in some measurements when adjusting the operation of the distribution network. In any fault event, the metering equipment may fail to function properly, resulting in measurement issues.</p>
</list-item>
<list-item>
<p>3) Communication systems. Due to the distribution network&#x2019;s complex topology and geographical environment, local communication links usually use low-power and lossy networks in power distribution networks. This type of network is prone to data packet loss. Also, the reliability of distribution network data transmission is affected by the communication links. The way of communication will also affect the reliability of data transmission in the distribution network. Due to cost constraints, most distribution topologies use communication methods such as distribution carrier waves, Zigbee wireless technology, and industrial wiring (<xref ref-type="bibr" rid="B18">Pei et&#x20;al., 2010</xref>). These communication methods are less reliable and often break codes when the channel is exposed to heavy electromagnetic interference, resulting in missing&#x20;data.</p>
</list-item>
</list>
</p>
<p>All of the above issues may produce anomalous data, causing the quality of the data to be inconsistent and reduce usability. Therefore, it is necessary to clean the data before using it for analysis and utilization. The prominent data anomalies in the existing distribution network data are missing data and outliers. The term &#x201c;data missing&#x201d; applies when the collected value is null or contains an invalid value. In contrast, outliers occur when the collected value deviates from normal data. The value exceeds the acceptable range of change (data is too big or too small) and maintains a certain time pattern without repetition.</p>
<p>
<xref ref-type="fig" rid="F1">Figure&#x20;1</xref> illustrates the distribution of abnormal voltage data in the distribution network for various purposes. <xref ref-type="fig" rid="F1">Figure&#x20;1A</xref> shows the abnormality caused by the malfunction of the metering equipment. The characteristic feature of this phenomenon is that some observations are outliers or missing values, which do not last for a long time but occur frequently. <xref ref-type="fig" rid="F1">Figure&#x20;1B</xref> shows the data abnormality caused by the failure of a terminal monitoring point. It is characterized by continuous data anomalies or single-point data anomalies occurring at a single point of observation. <xref ref-type="fig" rid="F1">Figure&#x20;1C</xref> shows the data anomalies caused by excited inrush disturbances and automation equipment actions at controller monitoring points. This abnormality is characterized by short-term outliers in some observations, i.e.,&#x20;retained for a very small period of time. <xref ref-type="fig" rid="F1">Figure&#x20;1D</xref> shows data anomalies caused by faults in the communication system of the sub-stations. It is defined by a partial loss of temporal data at various intervals and is typically retained for a short period of&#x20;time.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>The distribution of abnormal voltage data for several reasons: <bold>(A)</bold> Abnormal voltage data cause by Metering equipment; <bold>(B)</bold> Abnormal voltage data cause by faults; <bold>(C)</bold> Abnormal voltage data cause by operation control; <bold>(D)</bold> Abnormal voltage data caused by the communication system.</p>
</caption>
<graphic xlink:href="fenrg-09-730058-g001.tif"/>
</fig>
</sec>
<sec id="s2-3">
<title>Outliers Cleaning in Distribution Network</title>
<p>The above issues might pollute the measurement data, which make it not suitable used directly for distribution system planning and operation. Therefore, the data preprocessing is an essential process before using data for analysis and decision-making, and the outliers detection is the most important part of the process. Generally speaking, a good outlier detection algorithm should be able to identify outliers correctly, and would no have any response to the normal data. As shown in <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>, a dataset of the voltage amplitude, which received from sensors and contains four outliers and one missing data, and highlight it in the figure. The aim of the outliers detection is to find the highlight points and mark it with lables, which is the kernel of the data cleaning.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>The data anomaly&#x20;plot.</p>
</caption>
<graphic xlink:href="fenrg-09-730058-g002.tif"/>
</fig>
<p>Association rules learning are a rule-based machine learning method, which is a research focus of the data mining and analysis. In an method for automatically generating association rules, it mainly includes three important steps: data denoising, data discretization and rule mining. Furthermore, how to select the sub-algorithm is a critical step. In this paper, the density-based algorithm, DBSCAN, is chosen in the data denoising step, which has a excellent result in denoising and high scalability. Then in the data discretization step, the distance-based algorithm, K-means is used since its precise classification result and high computation efficiency. And in the rule mining step, Apriori algorithm is selected because of its high stability and flexible extension ability. With that in mind, we presents an outliers cleaning method based on association rules, which could found the implicit relationship between features from the historical measurement data and pick up the valuable information on outliers detection and repairing. For the outliers detection, the DBSCAN, K-means and Apriori algorithm are chosen for generating the association rules from historical data, which make the detector more flexible and accurate. For the outliers repairing, the repairing cost is chosen with a distance-based model. And the Mahalanobis distance is chosen to use for constructing a data repairing cost function, which could reduce the errors.</p>
</sec>
</sec>
<sec id="s3">
<title>Preliminaries</title>
<sec id="s3-1">
<title>DBSCAN Clustering Algorithm</title>
<p>DBSCAN is an unsupervised machine learning clustering algorithm that could be used for data classification with a nonlinear density structure (<xref ref-type="bibr" rid="B4">Chen et&#x20;al., 2021</xref>; <xref ref-type="bibr" rid="B6">Chipade et&#x20;al., 2021</xref>). The algorithm treats the data as points in space and clusters them based on density magnitude, allowing clusters of arbitrary shapes to be found in a noisy space. The basic idea of DBSCAN is to introduce neighborhood and density connectivity concepts, explore the data points, and use density connectivity to grow clusters until outliers split them. The DBSCAN clustering algorithm can be adapted to any form of clustering. It can filter out noisy outliers in space, making it ideal for outliers detection in the distribution network.</p>
<p>Assume a historical measurements datasets <inline-formula id="inf1">
<mml:math id="m1">
<mml:mrow>
<mml:mi mathvariant="bold-italic">D</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> be the numerical attributes of observations with rows <inline-formula id="inf2">
<mml:math id="m2">
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi mathvariant="normal">&#x3f5;</mml:mi>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mn>1,2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. For any observation <inline-formula id="inf3">
<mml:math id="m3">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="bold-italic">D</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> has a timestamp. And it contains <italic>m</italic> features, as given by <xref ref-type="disp-formula" rid="e1">Equation 1</xref>.<disp-formula id="e1">
<mml:math id="m4">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>where <italic>f</italic>
<sub>
<italic>ij</italic>
</sub> represents the <italic>j</italic>th feature of the <italic>i</italic>th&#x20;data.</p>
<p>Meanwhile, let <inline-formula id="inf4">
<mml:math id="m5">
<mml:mi>&#x3b5;</mml:mi>
</mml:math>
</inline-formula> be the neighborhood distance parameter of DBSCAN. For each point <inline-formula id="inf5">
<mml:math id="m6">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="bold-italic">D</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, the <inline-formula id="inf6">
<mml:math id="m7">
<mml:mi>&#x3b5;</mml:mi>
</mml:math>
</inline-formula> - neighborhood set is defined using <xref ref-type="disp-formula" rid="e2">Eq. 2</xref>.<disp-formula id="e2">
<mml:math id="m8">
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mi>&#x3b5;</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mi>&#x3b5;</mml:mi>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="italic">dist</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>&#x3b5;</mml:mi>
</mml:mrow>
</mml:mrow>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>
</p>
<p>Then, for any point <inline-formula id="inf7">
<mml:math id="m9">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="bold-italic">D</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, the core point should be satisfied <italic>via</italic> <xref ref-type="disp-formula" rid="e3">Eq. 3</xref>.<disp-formula id="e3">
<mml:math id="m10">
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mi>&#x3b5;</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
<mml:mo>&#x2265;</mml:mo>
<mml:mi mathvariant="italic">minpts</mml:mi>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>where, <inline-formula id="inf8">
<mml:math id="m11">
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mi>&#x3b5;</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the count of elements in the <inline-formula id="inf9">
<mml:math id="m12">
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mi>&#x3b5;</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. And <italic>minpts</italic> is the minimum number of points in <inline-formula id="inf10">
<mml:math id="m13">
<mml:mi>&#x3b5;</mml:mi>
</mml:math>
</inline-formula> -neighborhood.</p>
<p>If a point <inline-formula id="inf11">
<mml:math id="m14">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mi>&#x3b5;</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> satisfy <xref ref-type="disp-formula" rid="e2">Eq. (2)</xref>, <inline-formula id="inf12">
<mml:math id="m15">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is directly-density reachable from the point <inline-formula id="inf13">
<mml:math id="m16">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. As shown in <xref ref-type="fig" rid="F3">Figure&#x20;3</xref>, points that are outside the range of clustering are considered outliers. For clarity, <xref ref-type="other" rid="alg1">Algorithm 1</xref> explains the step-by-step procedure of the DBSACN clustering algorithm.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>The process of DBSCAN clustering.</p>
</caption>
<graphic xlink:href="fenrg-09-730058-g003.tif"/>
</fig>
<p>
<statement content-type="algorithm" id="alg1">
<label>Algorithm 1</label>
<p>: Density-based spatial clustering of applications with noise(DBSCAN).</p>
<p>
<inline-graphic xlink:href="fenrg-09-730058-fx1.tif"/>
</p>
</statement>
</p>
</sec>
<sec id="s3-2">
<title>Association Rules Mining</title>
<p>Association rule learning is a rule-based machine learning method used to mine frequent patterns, correlations, or causal structures between itemsets. It is intended to determine valuable rules based on the frequency of occurrence between itemsets in a database (<xref ref-type="bibr" rid="B20">Rauch, 2005</xref>; <xref ref-type="bibr" rid="B5">Chengyu et&#x20;al., 2016</xref>). In data cleaning, the association rules algorithm is applied for mining the relationship between various features of the measurement data in the power system.</p>
<p>At first, the obtained features should be discretized to improve the robustness of association rules to outliers. After data discretization, a datasets <bold>D</bold> with size <inline-formula id="inf14">
<mml:math id="m17">
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, which is a 2-dimensional real-valued matrix, is converted to a Boolean matrix <inline-formula id="inf15">
<mml:math id="m18">
<mml:mrow>
<mml:mover accent="true">
<mml:mi mathvariant="bold-italic">D</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> with size <inline-formula id="inf16">
<mml:math id="m19">
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>m</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. When a feature is in the specified interval, the value in the Boolean matrix is labelled as 1. If not, it is labelled as 0. <xref ref-type="fig" rid="F4">Figure&#x20;4</xref> illustrates a portion of the datasets before and after data discretization.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>The process of data discretization.</p>
</caption>
<graphic xlink:href="fenrg-09-730058-g004.tif"/>
</fig>
<p>Let <inline-formula id="inf17">
<mml:math id="m20">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> be the itemsets of <inline-formula id="inf18">
<mml:math id="m21">
<mml:mrow>
<mml:mover accent="true">
<mml:mi mathvariant="bold-italic">D</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula>. Each observation <inline-formula id="inf19">
<mml:math id="m22">
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mi mathvariant="bold-italic">D</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> may contain one or more items. The aim is to search for the most frequent patterns of items from the datasets for generating association rules. With this in mind, an association rule between <italic>F</italic>
<sub>
<italic>A</italic>
</sub> and <italic>F</italic>
<sub>
<italic>B</italic>
</sub>, can be defined as <xref ref-type="disp-formula" rid="e4">Eq. 4</xref>.<disp-formula id="e4">
<mml:math id="m23">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>A</mml:mi>
</mml:msub>
<mml:mo>&#x21d2;</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>where <inline-formula id="inf20">
<mml:math id="m24">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>A</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf21">
<mml:math id="m25">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are itemsets, and <inline-formula id="inf22">
<mml:math id="m26">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>A</mml:mi>
</mml:msub>
<mml:mo>&#x2282;</mml:mo>
<mml:mi>F</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>
<italic>,</italic> <inline-formula id="inf23">
<mml:math id="m27">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
<mml:mo>&#x2282;</mml:mo>
<mml:mi>F</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>Since the magnitude of support and confidence value is commonly used to assess the effectiveness of an association rule. The following <xref ref-type="disp-formula" rid="e5">Eq. 5</xref> represents a definition to support an association rule between <inline-formula id="inf24">
<mml:math id="m28">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>A</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf25">
<mml:math id="m29">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>:<disp-formula id="e5">
<mml:math id="m30">
<mml:mrow>
<mml:mi mathvariant="italic">Sup</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>A</mml:mi>
</mml:msub>
<mml:mo>&#x21d2;</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>B</mml:mi>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="italic">coun</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>A</mml:mi>
</mml:msub>
<mml:mo>&#x222a;</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>B</mml:mi>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>D</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>where <italic>count</italic> is the number of occurrences of an item in data <inline-formula id="inf26">
<mml:math id="m31">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>D</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula>; <inline-formula id="inf27">
<mml:math id="m32">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>A</mml:mi>
</mml:msub>
<mml:mo>&#x222a;</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>B</mml:mi>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the coexistence of <inline-formula id="inf28">
<mml:math id="m33">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>A</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf29">
<mml:math id="m34">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>; and <inline-formula id="inf30">
<mml:math id="m35">
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>D</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the total count of itemsets.</p>
<p>The confidence value indicates the reliability of the association rule. The confidence of an association rule between and can be defined as follows:<disp-formula id="e6">
<mml:math id="m36">
<mml:mrow>
<mml:mi mathvariant="italic">Con</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>A</mml:mi>
</mml:msub>
<mml:mo>&#x21d2;</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>B</mml:mi>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="italic">coun</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>A</mml:mi>
</mml:msub>
<mml:mo>&#x21d2;</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>B</mml:mi>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">Sup</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>A</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>
</p>
<p>In this work, we use Apriori algorithm is employed to search for the itemsets frequency in the complete transaction set. In this approach, the ones with more than the minimum support and the minimum confidence are used as strong association rules. These rules give high confidence and strong support greater than or equal to a user-specified minimum confidence threshold and a minimum support threshold. The process of the algorithm is explained in <xref ref-type="other" rid="alg2">Algorithm&#x20;2</xref>.</p>
<p>
<statement content-type="algorithm" id="alg2">
<label>Algorithm 2</label>
<p>: Association Rule Mining with Apriori.</p>
<p>
<inline-graphic xlink:href="fenrg-09-730058-fx2.tif"/>
</p>
</statement>
</p>
</sec>
<sec id="s3-3">
<title>Mahalanobis Distance</title>
<p>The Euclidean distance is the most common metric of distance in data science, which describes the straight-line distance between two points in Euclidean space. Consider the case where two or more variables are linked. The axes are no longer at right angles in this situation, and measurements are no longer possible. The Euclidean distance cannot represent the real distance between feature vectors. To mitigate this issue, Mahalanobis distance is implemented. The Mahalanobis distance measures the distance between two points in multivariate space (<xref ref-type="bibr" rid="B14">Maesschalck et&#x20;al., 2000</xref>). It calculates distances between points, including correlated points for multiple variables.</p>
<p>For a given sample set, in order to integrate the correlation between the sample points, the distance between numerical data features can be calculated using the Mahalanobis distance metric to measure the similarity between samples (<xref ref-type="bibr" rid="B13">Liu, et&#x20;al., 2020</xref>). For a feature, the variation between two observations can be specified by <xref ref-type="disp-formula" rid="e7">Eq. 7</xref>.<disp-formula id="e7">
<mml:math id="m37">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold">M</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msqrt>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:msup>
<mml:mi mathvariant="bold">S</mml:mi>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msqrt>
<mml:mo>&#x3d;</mml:mo>
<mml:msqrt>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:mi mathvariant="bold">M</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>where <italic>S</italic> is the covariance matrix, and <italic>M</italic>&#x20;&#x3d; <italic>S</italic>
<sup>&#x2212;1</sup>, if the two samples have similar or identical characteristics, the martingale distance should be small or even&#x20;zero.</p>
<p>
<xref ref-type="fig" rid="F5">Figure&#x20;5</xref> illustrates the transformation in the two-dimension space. Here, the blue and pink dots represent the original and transformed data. For presented data points, the correlation of the two features causes the oval shape of the original distribution. If we apply the Euclidean distance, it will not reflect the real dissimilarity of the data. While calculating Mahalanobis distance, the eclipse is first transformed to a standardized circle with a radius equal to 1, and then the Euclidean distance in the transformed space is calculated. Meanwhile, computing the distance, the influence of correlation is offset by the transformation.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>An illustration of Mahalanobis distance.</p>
</caption>
<graphic xlink:href="fenrg-09-730058-g005.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<title>The Proposed Model for Outliers Detection and Repairing</title>
<p>Since historical measurement data from SCADAS is required processing, using this information, the proposed model generates a list of association rules to evaluate the correlation between distinct variables. The overall flowchart of the proposed method is shown in <xref ref-type="fig" rid="F6">Figure&#x20;6</xref>. The proposed model consists of three stages: data preparation, outliers detection and outliers repairing.<list list-type="simple">
<list-item>
<p>1) In the data preparation process, it has a task with DBSACN clustering algorithm to eliminate ineffective information, which purposes is to find the obvious abnormal data like null or missing values and so on. Then prepared dataset (<bold>D</bold>) will be discretized with k-means clustering algorithm, which purposes is to get the discrete dataset <inline-formula id="inf31">
<mml:math id="m38">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>D</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> to generate association&#x20;rules.</p>
</list-item>
<list-item>
<p>2) In the outliers detection process, it is necessary to mine the association rules in historical data for detecting outliers. The association rules will be mined from the dataset <inline-formula id="inf32">
<mml:math id="m39">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>D</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> with Apriori algorithm. Correspondingly, the intervals of the features will be define by the association rules, which the features are distributed. The newly obtained observation will be compared with the intervals which defined from the association rules. If these real-time observations out of the intervals (i.e.,&#x20;identified by the association rules), they will be flagged as outliers. In this paper, we use 1 as the label for outlier, and 0 as the label for normal&#x20;data.</p>
</list-item>
<list-item>
<p>3) In the outliers repairing process, all the sample points would be mapped into a feature space, which determining by the association rules. Subsequently, a novel cost function is constructed and used for data repairing, and the outlier will be repaired with the value which have the minimum repair cost. In this paper, the distance metric is formed with Mahalanobis distance, and similar normal data related to the query outliers are retrieved.</p>
</list-item>
</list>
</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>The schema of the proposed method.</p>
</caption>
<graphic xlink:href="fenrg-09-730058-g006.tif"/>
</fig>
<sec id="s4-1">
<title>Data Preparation</title>
<p>In practice, there are some anomalies in the historical measurement data from the distribution network, which may bring invalid/incorrect information. Therefore, it is necessary to process the historical data before mining the association rules. We employ the DBSCAN clustering algorithm as the noise detector and design a procedure to dispose of the anomalies in the historical measurement data. The DBSCAN algorithm divides these observations into several clusters and outliers. The parameter of and <italic>minpts</italic> is chosen based on the silhouette coefficient. Next, the data points in each cluster are labelled as 0, while the outliers are labelled as 1. Furthermore, to generate the association rules, we use k-means for data discretization.</p>
</sec>
<sec id="s4-2">
<title>Outliers Detection</title>
<p>The results from the association rules are used for detecting the outliers. In any case, the outliers detection of each feature is based on comparisons between the new real-time observations and the association rules generated from all historical measurement data. In the final analysis, if a real-time measurement mismatches the interval defined by the associated rule, it will be flagged as an outlier.</p>
<p>For example, assume association rules are generated as follows:<disp-formula id="equ1">
<mml:math id="m40">
<mml:mrow>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>A</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mn>1</mml:mn>
<mml:mi mathvariant="italic">min</mml:mi>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mn>1</mml:mn>
<mml:mi mathvariant="italic">max</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x21d2;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>B</mml:mi>
<mml:mn>1</mml:mn>
<mml:mi mathvariant="italic">min</mml:mi>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>B</mml:mi>
<mml:mn>1</mml:mn>
<mml:mi mathvariant="italic">max</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
</p>
<p>For a new observation <inline-formula id="inf33">
<mml:math id="m41">
<mml:mrow>
<mml:msub>
<mml:mi>O</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> which contains the same features <inline-formula id="inf34">
<mml:math id="m42">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>A</mml:mi>
</mml:msub>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>and</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>:<disp-formula id="equ2">
<mml:math id="m43">
<mml:mrow>
<mml:mi mathvariant="normal">If</mml:mi>
<mml:mtext>&#x2009;</mml:mtext>
<mml:msub>
<mml:mi>O</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>A</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mn>1</mml:mn>
<mml:mi mathvariant="italic">min</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mn>1</mml:mn>
<mml:mi mathvariant="italic">max</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="normal">while</mml:mi>
<mml:mtext>&#x2009;</mml:mtext>
<mml:msub>
<mml:mi>O</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>B</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2209;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>B</mml:mi>
<mml:mn>1</mml:mn>
<mml:mi mathvariant="italic">min</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>B</mml:mi>
<mml:mn>1</mml:mn>
<mml:mi mathvariant="italic">max</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
</p>
<p>According to the association rule, the new observation is compared to previous observations that have the same features. The <inline-formula id="inf35">
<mml:math id="m44">
<mml:mrow>
<mml:msub>
<mml:mi>O</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> stays out of the intervals, which signify that <inline-formula id="inf36">
<mml:math id="m45">
<mml:mrow>
<mml:msub>
<mml:mi>O</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is an outlier, and the current observation would be labeled as&#x20;1.</p>
</sec>
<sec id="s4-3">
<title>Outliers Repairing</title>
<p>As shown in <xref ref-type="fig" rid="F7">Figure&#x20;7</xref>, after outliers detection, for any point <inline-formula id="inf37">
<mml:math id="m46">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mi>D</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, it would be allocated in a feature space divided by the association rules, which we called &#x201c;data binding&#x201d;. When an observation of one feature is marked as an outlier, it is necessary to calculate the estimated value of the outlier. For an abnormal observation in the &#x201c;rule box space (RBS)&#x201d;, the point with the highest similarity to its attribute should fall in the same sub-RBS.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>The RBS with association&#x20;rules.</p>
</caption>
<graphic xlink:href="fenrg-09-730058-g007.tif"/>
</fig>
<p>For example, assume the presence of outliers for the voltage magnitude, and the following Rule1 are generated with the highest confidence:<disp-formula id="equ3">
<mml:math id="m47">
<mml:mrow>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mtext>Current</mml:mtext>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mn>0.2272</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>0.2408</mml:mn>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mtext>Active</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>Power</mml:mtext>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mn>38.5260</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>39.9682</mml:mn>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mtext>Reactive</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>Power</mml:mtext>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mn>42.1306</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>44.8896</mml:mn>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
<mml:mo>&#x21d2;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mtext>Voltage</mml:mtext>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mn>142.3809</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>144.0914</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
</p>
<p>According to this rule, the correct value of voltage magnitude should be in the range of [142.3809,144.0914). It means that after the outliers are repaired, the point is still in the original sub-RBS. This is because data repairing can be translated into searching for normal data with the highest similarity to it within the&#x20;RBS.</p>
<p>The Mahalanobis distance is used as a metric to account for the distributional differences between attributes. The repair results within this RBS are not unique, and each result has a corresponding repair cost. As shown in <xref ref-type="fig" rid="F8">Figure&#x20;8</xref>, in outlier repairing, the objective is to minimize the repairing cost function as <xref ref-type="disp-formula" rid="e8">Eq. 8</xref>.<disp-formula id="e8">
<mml:math id="m48">
<mml:mrow>
<mml:mi mathvariant="bold-italic">cost</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mi>x</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">M</mml:mi>
<mml:mi>d</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mi>x</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(8)</label>
</disp-formula>where <italic>x</italic>&#x2032; is the repairing result of the corresponding&#x20;<italic>x</italic>.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>The repairing cost of one outlier.</p>
</caption>
<graphic xlink:href="fenrg-09-730058-g008.tif"/>
</fig>
<p>However, in some cases, the minimum value of the cost function is more than one. In these cases, the abnormal observation will be replaced by <inline-formula id="inf38">
<mml:math id="m49">
<mml:mrow>
<mml:msubsup>
<mml:mi>O</mml:mi>
<mml:mi>t</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>, which calculated by the following <xref ref-type="disp-formula" rid="e9">Eq. 9</xref>.<disp-formula id="e9">
<mml:math id="m50">
<mml:mrow>
<mml:msubsup>
<mml:mi mathvariant="bold-italic">O</mml:mi>
<mml:mi>t</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">mean</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">D</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>min</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(9)</label>
</disp-formula>where <inline-formula id="inf39">
<mml:math id="m51">
<mml:mrow>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>min</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="bold-italic">D</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> that corresponds to <inline-formula id="inf40">
<mml:math id="m52">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi mathvariant="italic">min</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, and <inline-formula id="inf41">
<mml:math id="m53">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi mathvariant="italic">min</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the set with the minimize repairing&#x20;cost.</p>
<p>In addition, we consider the outliers repairing with the data feedback. When an outlier is cleaned to a normal value after outliers repairing, the observation would be updated in the RBS for a new outlier that needs repair. The accuracy of outliers repairing could be improved with data feedback.</p>
</sec>
</sec>
<sec id="s5">
<title>Experiment and Analysis</title>
<sec id="s5-1">
<title>The Metrics Used for Evaluating</title>
<p>Outliers detection of measurement data is an unbalanced binary classification problem. Data are classified as normal or abnormal. In this way, accuracy is not an appropriate metric for evaluating the performance of a method. The detection results could be classified into four types according to the label between actual and prediction values: true positive (TP), true negative (TN), false positive (FP) and false negative (FN). The confusion matrix of outliers detection is shown in <xref ref-type="table" rid="T1">Table&#x20;1</xref>. In general, the <italic>Precision</italic>, <italic>Recall</italic> and <italic>F1-Score</italic> are used as the metrics for evaluating of classification problem. Among them, <italic>the Precision</italic> is a metrics reflects the reliability of the detection results, while <italic>the Recall</italic> is a metrics reflects how many truly detection results are returned. And the <italic>F1-Score</italic> is the harmonic mean of precision and recall.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Confusion matrix of outliers detection.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Actual label</th>
<th colspan="2" align="center">Detection results</th>
</tr>
<tr>
<th align="center">Outlier/1</th>
<th align="center">Normal/0</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Outlier/1</td>
<td align="center">TP</td>
<td align="center">FN</td>
</tr>
<tr>
<td align="left">Normal/0</td>
<td align="center">FP</td>
<td align="center">TN</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>According to the confusion matrix of outliers detection, the <italic>Precision</italic>, <italic>Recall</italic> and <italic>F1-Score</italic> could be calculated by <xref ref-type="disp-formula" rid="e10">Eqs 10</xref>&#x2013;<xref ref-type="disp-formula" rid="e12">12</xref>.<disp-formula id="e10">
<mml:math id="m54">
<mml:mrow>
<mml:mi mathvariant="italic">Precision</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="italic">TP</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">TP</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="italic">FP</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(10)</label>
</disp-formula>
<disp-formula id="e11">
<mml:math id="m55">
<mml:mrow>
<mml:mi mathvariant="italic">Recall</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="italic">TP</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">TP</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="italic">FN</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(11)</label>
</disp-formula>
<disp-formula id="e12">
<mml:math id="m56">
<mml:mrow>
<mml:mi mathvariant="italic">F</mml:mi>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="italic">Score</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mi mathvariant="italic">Precision</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi mathvariant="italic">Recall</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">Precision</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="italic">Recall</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(12)</label>
</disp-formula>where TP is the count of outlier detected as an outlier, FP is the count of normal data detected as an outlier, FN is the count of outlier detected as normal&#x20;data.</p>
<p>In addition, two metrics used for outliers repairing: the mean absolute error (MAE) and the root mean square error (RSME). They are defined as follow <xref ref-type="disp-formula" rid="e13">Eqs 13</xref>, <xref ref-type="disp-formula" rid="e14">14</xref> .<disp-formula id="e13">
<mml:math id="m57">
<mml:mrow>
<mml:mi mathvariant="bold-italic">MAE</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mi mathvariant="bold">1</mml:mi>
<mml:mi mathvariant="bold-italic">N</mml:mi>
</mml:mfrac>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>x</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mi>j</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msubsup>
</mml:mrow>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
<label>(13)</label>
</disp-formula>
<disp-formula id="e14">
<mml:math id="m58">
<mml:mrow>
<mml:mi mathvariant="bold-italic">RMSE</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mi mathvariant="bold">1</mml:mi>
<mml:mi mathvariant="bold-italic">N</mml:mi>
</mml:mfrac>
<mml:msqrt>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>x</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:math>
<label>(14)</label>
</disp-formula>where <italic>N</italic> is the size of data, <inline-formula id="inf42">
<mml:math id="m59">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>x</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the <italic>i</italic>th actual value (without contaminated) of outliers, <inline-formula id="inf43">
<mml:math id="m60">
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> is the estimation of the <italic>i</italic>th outlier.</p>
</sec>
<sec id="s5-2">
<title>The Simulated Dataset</title>
<p>Unfortunately, the measurement datasets from real-world distribution networks are unlabeled. It means that it is not appropriate to use as a dataset for evaluating the proposed methods. Hence, we used the simulated datasets with artificial error. To verify the correctness and effectiveness of the proposed method, a test system (from a region in southwest China) is modeled in PSCAD/EMTDC to collect simulation data, as shown in <xref ref-type="fig" rid="F9">Figure&#x20;9</xref>. The operational datasets contain 4000 samples (with a sampling rate of 40 frames per second) and four features (voltage magnitude V, current magnitude I, active power P, reactive power Q) for the distribution network. There are no outliers in these datasets. We added some synthetic errors to the simulated measurement data, in which outliers are generated and injected into the datasets using a normal-distributed random function as <italic>z</italic>&#x20;&#x3d; <italic>G</italic>(<italic>x</italic>). <xref ref-type="table" rid="T2">Table&#x20;2</xref> shows that each bus data has 5&#x2013;15% noise injected into it. As an example, if a data point has a voltage magnitude feature of 110<italic>k</italic>V, the noise is calculated as 110&#x2a;105% &#x2b; G(x). For each sample, three-fourths of the data is taken as input, and the trained algorithm predicts the rest of the real-time&#x20;value.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>The test system from a region in southwest China.</p>
</caption>
<graphic xlink:href="fenrg-09-730058-g009.tif"/>
</fig>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>The noises injection of simulated dataset.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Anomaly rate</th>
<th align="center">The outliers calculation in each feature</th>
<th align="center">Anomaly proportion</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Noise 5%</td>
<td align="char" char="+">1.p.u &#x2a;105% &#x2b; G(x)</td>
<td align="char" char="/">569/4000</td>
</tr>
<tr>
<td align="left">Noise 10%</td>
<td align="char" char="+">1.p.u &#x2a;105% &#x2b; G(x)</td>
<td align="char" char="/">1091/4000</td>
</tr>
<tr>
<td align="left">Noise 15%</td>
<td align="char" char="+">1.p.u &#x2a;105% &#x2b; G(x)</td>
<td align="char" char="/">1529/4000</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5-3">
<title>Outliers Detection</title>
<p>The discretization of the pre-processed datasets is performed using k-means clustering. The numerical attributes voltage magnitude, current magnitude, active power and reactive power are clustered into 8, 5, 6, and 5 categories, respectively. The clusters are selected based on the quality metric that is finally estimated. After data discretization, the Apriori algorithm generates the association rules in the test datasets with confidence greater than 60%. For each posterior feature, the association rules are generated separately. Then the rules are generated for prediction individuals. Some of these rules are shown in <xref ref-type="table" rid="T3">Table&#x20;3</xref>. If the observation is not within the interval determined by the rules, it will be marked as an outlier.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>The part of association rules in different anomaly rate (voltage as posterior).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Anomaly rate (%)</th>
<th align="center">Prior</th>
<th align="center">Posterior</th>
<th align="center">
<italic>Con</italic>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="6" align="left">5</td>
<td align="center">{Current &#x3d; [0.2408, 0.2521), Active Power &#x3d; [38.5260, 39.9682), Reactive Power &#x3d; [45.6303, 46,6837)}</td>
<td align="center">{Voltage &#x3d; [142.3809, 144.0914)}</td>
<td align="char" char=".">0.7692</td>
</tr>
<tr>
<td align="center">{Current &#x3d; [0.2272, 0.2408], Active Power &#x3d; [38.5260, 39.9682), Reactive Power &#x3d; [41.1306, 43.8896]}</td>
<td align="center">{Voltage &#x3d; [134.5451, 137.8775]}</td>
<td align="char" char=".">0.7222</td>
</tr>
<tr>
<td align="center">{Current &#x3d; [0.2272, 0.2408], Active Power &#x3d; [38.0519, 38.5260), Reactive Power &#x3d; [45.6303, 46,6837)}</td>
<td align="center">{Voltage &#x3d; [144.0914, 149.0609)}</td>
<td align="char" char=".">0.6818</td>
</tr>
<tr>
<td align="center">{Current &#x3d; [0.2272, 0.2408], Active Power &#x3d; [38.0519, 38.5260), Reactive Power &#x3d; [41.1306, 43.8896]}</td>
<td align="center">{Voltage &#x3d; [137.8775, 140.3127)}</td>
<td align="char" char=".">0.6800</td>
</tr>
<tr>
<td align="center">{Current &#x3d; [0.2408, 0.2521), Active Power &#x3d; [38.5260, 39.9682), Reactive Power &#x3d; [43.8896, 45.6303)}</td>
<td align="center">{Voltage &#x3d; [140.3127, 142.3809)}</td>
<td align="char" char=".">0.6800</td>
</tr>
<tr>
<td align="center">
<bold>...</bold>
</td>
<td align="center">
<bold>...</bold>
</td>
<td align="char" char=".">
<bold>...</bold>
</td>
</tr>
<tr>
<td rowspan="6" align="left">10</td>
<td align="center">{Current &#x3d; [0.2272, 0.2408], Active Power &#x3d; [36.9662, 38.0519], Reactive Power &#x3d; [46.6837, 48.0561)}</td>
<td align="center">{Voltage &#x3d; [144.0914, 149.0609)}</td>
<td align="char" char=".">0.7826</td>
</tr>
<tr>
<td align="center">{Current &#x3d; [0.2408, 0.2521), Active Power &#x3d; [38.0519, 38.5260), Reactive Power &#x3d; [45.6303, 46,6837)}</td>
<td align="center">{Voltage &#x3d; [142.3809, 144.0914)}</td>
<td align="char" char=".">0.7407</td>
</tr>
<tr>
<td align="center">{Current &#x3d; [0.2408, 0.2521), Active Power &#x3d; [38.0519, 38.5260), Reactive Power &#x3d; [45.6303, 46,6837)}</td>
<td align="center">{Voltage &#x3d; [137.8775, 140.3127)}</td>
<td align="char" char=".">0.7333</td>
</tr>
<tr>
<td align="center">{Current &#x3d; [0.2272, 0.2408], Active Power &#x3d; [38.5260, 39.9682), Reactive Power &#x3d; [45.6303, 46,6837)}</td>
<td align="center">{Voltage &#x3d; [134.5451, 137.8775]}</td>
<td align="char" char=".">0.7142</td>
</tr>
<tr>
<td align="center">{Current &#x3d; [0.2272, 0.2408], Active Power &#x3d; [39.9682, 42.3351), Reactive Power &#x3d; [41.1306, 43.8896]}</td>
<td align="center">{Voltage &#x3d; [140.3127, 142.3809)}</td>
<td align="char" char=".">0.6785</td>
</tr>
<tr>
<td align="center">
<bold>...</bold>
</td>
<td align="center">
<bold>...</bold>
</td>
<td align="char" char=".">
<bold>...</bold>
</td>
</tr>
<tr>
<td rowspan="6" align="left">15</td>
<td align="center">{Current &#x3d; [0.2272, 0.2408], Active Power &#x3d; [38.5260, 39.9682), Reactive Power &#x3d; [43.8896, 45.6303)}</td>
<td align="center">{Voltage &#x3d; [140.3127, 142.3809)}</td>
<td align="char" char=".">0.8181</td>
</tr>
<tr>
<td align="center">{Current &#x3d; [0.2272, 0.2408], Active Power &#x3d; [39.9682, 42.3351), Reactive Power &#x3d; [46.6837, 48.0561)}</td>
<td align="center">{Voltage &#x3d; [137.8775, 140.3127)}</td>
<td align="char" char=".">0.7368</td>
</tr>
<tr>
<td align="center">{Current &#x3d; [0.2272, 0.2408], Active Power &#x3d; [36.9662, 38.0519], Reactive Power &#x3d; [45.6303, 46,6837)}</td>
<td align="center">{Voltage &#x3d; [134.5451, 137.8775]}</td>
<td align="char" char=".">0.7333</td>
</tr>
<tr>
<td align="center">{Current &#x3d; [0.2408, 0.2521), Active Power &#x3d; [38.5260, 39.9682), Reactive Power &#x3d; [45.6303, 46,6837)}</td>
<td align="center">{Voltage &#x3d; [142.3809, 144.0914)}</td>
<td align="char" char=".">0.7000</td>
</tr>
<tr>
<td align="center">{Current &#x3d; [0.2408, 0.2521), Active Power &#x3d; [38.5260, 39.9682), Reactive Power &#x3d; [46.6837, 48.0561)}</td>
<td align="center">{Voltage &#x3d; [144.0914, 149.0609)}</td>
<td align="char" char=".">0.6957</td>
</tr>
<tr>
<td align="center">
<bold>...</bold>
</td>
<td align="center">
<bold>...</bold>
</td>
<td align="center">
<bold>...</bold>
</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>For evaluation, the proposed method is compared with other methods such as decision tree, k-neighbors, and SVM; all methods are using simulated datasets. The results are shown in <xref ref-type="table" rid="T4">Table&#x20;4</xref> and <xref ref-type="fig" rid="F10">Figure&#x20;10</xref>, and the comparison is based on <italic>Precision</italic>, <italic>Recall</italic> and <italic>F1-Score</italic>, respectively. For the <italic>Precision</italic>, considering the dataset with 5&#x2013;15% noise, the above method have similar results, which means that the change of anomaly rate has little effect on the Precision. For the <italic>Recall</italic>, with the anomaly rate increases, the above method have worse results, which means that the change of anomaly rate mainly affects the <italic>Recall</italic>, resulting in the change of <italic>F1-score</italic>.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>The results of the comparative analysis in different outliers detection method.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Anomaly rate (%)</th>
<th align="center">Method</th>
<th align="center">Precision</th>
<th align="center">Recall</th>
<th align="center">F1-score</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="4" align="left">5</td>
<td align="left">Decision Tree</td>
<td align="char" char=".">0.9649</td>
<td align="char" char=".">0.9649</td>
<td align="char" char=".">0.9649</td>
</tr>
<tr>
<td align="left">K-Neighbors</td>
<td align="char" char=".">1.00</td>
<td align="char" char=".">0.9298</td>
<td align="char" char=".">0.9636</td>
</tr>
<tr>
<td align="left">SVM</td>
<td align="char" char=".">1.00</td>
<td align="char" char=".">0.9123</td>
<td align="char" char=".">0.9541</td>
</tr>
<tr>
<td align="left">Proposed</td>
<td align="char" char=".">1.00</td>
<td align="char" char=".">0.9649</td>
<td align="char" char=".">0.9821</td>
</tr>
<tr>
<td rowspan="4" align="left">10</td>
<td align="left">Decision Tree</td>
<td align="char" char=".">0.9820</td>
<td align="char" char=".">0.9646</td>
<td align="char" char=".">0.9732</td>
</tr>
<tr>
<td align="left">K-Neighbors</td>
<td align="char" char=".">1.00</td>
<td align="char" char=".">0.9646</td>
<td align="char" char=".">0.9820</td>
</tr>
<tr>
<td align="left">SVM</td>
<td align="char" char=".">1.00</td>
<td align="char" char=".">0.9381</td>
<td align="char" char=".">0.9680</td>
</tr>
<tr>
<td align="left">Proposed</td>
<td align="char" char=".">1.00</td>
<td align="char" char=".">0.9823</td>
<td align="char" char=".">0.9911</td>
</tr>
<tr>
<td rowspan="4" align="left">15</td>
<td align="left">Decision Tree</td>
<td align="char" char=".">0.9630</td>
<td align="char" char=".">0.9420</td>
<td align="char" char=".">0.9524</td>
</tr>
<tr>
<td align="left">K-Neighbors</td>
<td align="char" char=".">1.00</td>
<td align="char" char=".">0.9275</td>
<td align="char" char=".">0.9624</td>
</tr>
<tr>
<td align="left">SVM</td>
<td align="char" char=".">1.00</td>
<td align="char" char=".">0.8551</td>
<td align="char" char=".">0.9219</td>
</tr>
<tr>
<td align="left">Proposed</td>
<td align="char" char=".">1.00</td>
<td align="char" char=".">0.9348</td>
<td align="char" char=".">0.9663</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F10" position="float">
<label>FIGURE 10</label>
<caption>
<p>The histogram of the comparative analysis in different outliers detection methods.</p>
</caption>
<graphic xlink:href="fenrg-09-730058-g010.tif"/>
</fig>
<p>According to <xref ref-type="table" rid="T4">Table&#x20;4</xref> and <xref ref-type="fig" rid="F10">Figure&#x20;10</xref>, the result of decision tree is not satisfactory, which may mistakenly treats the normal data as an outlier. And for the datasets with highly contaminated (more than 15%), the SVM is leaving much to be desired, the <italic>Recall</italic> of it even less than 90%. For SVM, the reason may be that the high anomaly rate makes the training data extremely unbalanced. Therefore, in the training stage, the type of data may not meet the requirements of SVM, which limits the application of SVM in outliers detection. Under the same conditions, the <italic>Precision</italic> and <italic>Recall</italic> of k-neighbors is small than our proposed method. From the <xref ref-type="fig" rid="F10">Figure&#x20;10</xref>, the proposed method play a good performance, whose <italic>F1-Score</italic> remains more than 96% for the datasets with different anomaly rate. The comparative case studies show that our proposed method outperformed the other three methods.</p>
</sec>
<sec id="s5-4">
<title>Outliers Repairing</title>
<p>We employ two other widely-used data repairing methods to make a comparison, which includes decision tree and gradient boosting decision tree (GBDT). <xref ref-type="table" rid="T5">Table&#x20;5</xref> and <xref ref-type="fig" rid="F11">Figure&#x20;11</xref> show the comparison result of the MAE and RSME at different anomaly rates. The accuracy indices of the decision tree and GBDT is close under different anomaly rates. And the GBDT algorithm is in a better position than the decision tree algorithm for each evaluation indices. The proposed method has been superior in all the accuracy indices in terms of aggregate, indicating model robustness.</p>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>The results of the comparative analysis in different outliers repairing methods.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Anomaly rate (%)</th>
<th align="center">Method</th>
<th align="center">MAE</th>
<th align="center">RSME</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="3" align="left">5</td>
<td align="left">Decision Tree</td>
<td align="char" char=".">1.1307</td>
<td align="char" char=".">2.1325</td>
</tr>
<tr>
<td align="left">GBDT</td>
<td align="char" char=".">1.0744</td>
<td align="char" char=".">2.1315</td>
</tr>
<tr>
<td align="left">Proposed</td>
<td align="char" char=".">0.8720</td>
<td align="char" char=".">1.0871</td>
</tr>
<tr>
<td rowspan="3" align="left">10</td>
<td align="left">Decision Tree</td>
<td align="char" char=".">1.7386</td>
<td align="char" char=".">3.5983</td>
</tr>
<tr>
<td align="left">GBDT</td>
<td align="char" char=".">1.6941</td>
<td align="char" char=".">3.5701</td>
</tr>
<tr>
<td align="left">Proposed</td>
<td align="char" char=".">1.0647</td>
<td align="char" char=".">1.4617</td>
</tr>
<tr>
<td rowspan="3" align="left">15</td>
<td align="left">Decision Tree</td>
<td align="char" char=".">1.5002</td>
<td align="char" char=".">2.9637</td>
</tr>
<tr>
<td align="left">GBDT</td>
<td align="char" char=".">1.4785</td>
<td align="char" char=".">2.9260</td>
</tr>
<tr>
<td align="left">Proposed</td>
<td align="char" char=".">0.7751</td>
<td align="char" char=".">0.9595</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F11" position="float">
<label>FIGURE 11</label>
<caption>
<p>The histogram of the comparative analysis in different outliers repairing methods.</p>
</caption>
<graphic xlink:href="fenrg-09-730058-g011.tif"/>
</fig>
<p>In these cases, we only showed the result for detection and repairing the voltage, but the proposed method can also work for other features. In general, the proposed method outperforms the other methods in most cases. However, our approach&#x2019;s performance may not be good in some situations, especially when used with little historical data. The reason for the problem is the insufficiency of association rules. The association rules can not include all the conditions, which cause by the value of confidence (more than 60%). One way to improve the accuracy of the proposed method is to increase the amount of historical data. So that more association rules can be generated to mine the correlation between features from the&#x20;data.</p>
</sec>
</sec>
<sec sec-type="conclusion" id="s6">
<title>Conclusion</title>
<p>In this paper, we developed a association rules-based method for outliers cleaning. To detect outliers, the association rules are generated from historical data in conjunction with DBSCAN, k-means and Apriori technique. For the outliers repairing, we took into account the repairing cost by a distance-based model. The Mahalanobis distance was used for constructing a data repairing cost function to reduce the errors. The proposed method achieves accurate detection as compared to decision tree, k-neighbors, and SVM algorithms. When outliers is taken into account, our model produces a smaller MAE and RSME, which has a better result than decision tree and GBDT. The results show that the work has a positive effect on improving data quality, which means our works could provide a reliable data base for distribution network planning and operation. Future work will focus on combining this approach with Spark parallel computing technology to improve the efficiency of the algorithm to satisfy the practical application needs of distribution network measurement outliers cleaning.</p>
</sec>
</body>
<back>
<sec id="s7">
<title>Data Availability Statement</title>
<p>The raw data supporting the conclusion of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="s8">
<title>Author Contributions</title>
<p>Conception and design of study: HK; Acquisition of data: MH, XH; Drafting the article: RQ, XH; Analysis and interpretation of data: HK, RQ, CG, and XM; Revising the article critically for important intellectual content: RQ, XM, and RD.</p>
</sec>
<sec id="s9">
<title>Funding</title>
<p>This work was supported by the Science and Technology Foundation of China Southern Power Grid (YNKJXM20191369).</p>
</sec>
<sec sec-type="COI-statement" id="s10">
<title>Conflict of Interest</title>
<p>Author HK is employed by Yunnan Power Grid Co., Ltd., China.</p>
<p>The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s11">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Alimardani</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Therrien</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Atanackovic</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Jatskevich</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Vaahedi</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Distribution System State Estimation Based on Nonsynchronized Smart Meters</article-title>. <source>IEEE Trans. Smart Grid</source> <volume>6</volume> (<issue>6</issue>), <fpage>2919</fpage>&#x2013;<lpage>2928</lpage>. <pub-id pub-id-type="doi">10.1109/TSG.2015.2429640</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cai</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xiao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Tian</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Fu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A Multi-Source Data Collection and Information Fusion Method for Distribution Network Based on Iot Protocol</article-title>. <source>IOP Conf. Ser. Earth Environ. Sci.</source> <volume>651</volume> (<issue>2</issue>), <fpage>022076</fpage>. <pub-id pub-id-type="doi">10.1088/1755-1315/651/2/022076</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Lau</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Automated Load Curve Data Cleansing in Power Systems</article-title>. <source>IEEE Trans. Smart Grid</source> <volume>1</volume> (<issue>2</issue>), <fpage>213</fpage>&#x2013;<lpage>221</lpage>. <pub-id pub-id-type="doi">10.1109/TSG.2010.2053052</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Hong</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>HTsort: Enabling Fast and Accurate Spike Sorting on Multi-Electrode Arrays</article-title>. <source>Front. Comput. Neurosci.</source> <volume>15</volume>, <fpage>657151</fpage>. <pub-id pub-id-type="doi">10.3389/fncom.2021.657151</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chengyu</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Ying</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Research and Improvement of Apriori Algorithm for Association Rules</article-title>. <source>Phys. Rev. A</source>, <fpage>1</fpage>&#x2013;<lpage>4</lpage>. <pub-id pub-id-type="doi">10.1103/PhysRevA.94.042311</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chipade</surname>
<given-names>V. S.</given-names>
</name>
<name>
<surname>Marella</surname>
<given-names>V. S. A.</given-names>
</name>
<name>
<surname>Panagou</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Aerial Swarm Defense by StringNet Herding: Theory and Experiments</article-title>. <source>Front. Robot. AI</source> <volume>8</volume>, <fpage>640446</fpage>. <pub-id pub-id-type="doi">10.3389/frobt.2021.640446</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Esmalifalak</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Nguyen</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Detecting Stealthy False Data Injection Using Machine Learning in Smart Grid</article-title>. <source>IEEE Syst. J.</source>, <volume>11</volume>, <fpage>1</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1109/JSYST.2014.2341597</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hayes</surname>
<given-names>B. P.</given-names>
</name>
<name>
<surname>Gruber</surname>
<given-names>J.&#x20;K.</given-names>
</name>
<name>
<surname>Prodanovic</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Multi&#x2010;nodal Short&#x2010;term Energy Forecasting Using Smart Meter Data</article-title>. <source>IET Generation, Transm. Distribution</source> <volume>12</volume> (<issue>12</issue>), <fpage>2988</fpage>&#x2013;<lpage>2994</lpage>. <pub-id pub-id-type="doi">10.1049/iet-gtd.2017.1599</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Hierarchical Pressure Data Recovery for Pipeline Network via Generative Adversarial Networks</article-title>. <source>IEEE Trans. Automat. Sci. Eng.</source> (<issue>99</issue>), <fpage>1</fpage>&#x2013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1109/TASE.2021.3069003</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Cai</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>W.</given-names>
</name>
</person-group> &#x201c;<article-title>Power Data Cleaning Method Based on Isolation Forest and LSTM Neural Network</article-title>,&#x201d; in <conf-name>International Conference on Cloud Computing and Security</conf-name>, <conf-loc>Haikou, China</conf-loc>, <conf-date>26 September 2018</conf-date>, <fpage>539</fpage>&#x2013;<lpage>550</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-00018-9_47</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Deng</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A Big Data Cleaning Method Based on Improved CLOF and Random Forest for Distribution Network</article-title>. <source>CSEE J. Power Energy Syst</source>. <comment>(Early Access)</comment>. <pub-id pub-id-type="doi">10.17775/CSEEJPES.2020.04080</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Ding</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yan</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>L.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Data-Driven Condition Monitoring of Data Acquisition for Consumers&#x27; Transformers in Actual Distribution Systems Using T-Statistics</article-title>. <source>IEEE Trans. Power Deliv.</source> <volume>34</volume> (<issue>4</issue>), <fpage>1578</fpage>&#x2013;<lpage>1587</lpage>. <pub-id pub-id-type="doi">10.1109/TPWRD.2019.2912267</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Feng</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Data-driven Transient Stability Assessment Model Considering Network Topology Changes via Mahalanobis Kernel Regression and Ensemble Learning</article-title>. <source>J.&#x20;Mod. Power Syst. Clean Energ.</source> <volume>8</volume> (<issue>6</issue>), <fpage>1080</fpage>&#x2013;<lpage>1091</lpage>. <pub-id pub-id-type="doi">10.35833/MPCE.2020.000341</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Maesschalck</surname>
<given-names>R. D.</given-names>
</name>
<name>
<surname>Jouan-Rimbaud</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Massart</surname>
<given-names>D. L.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>The Mahalanobis Distance</article-title>. <source>Chemometrics Intell. Lab. Syst.</source> <volume>50</volume> (<issue>1</issue>), <fpage>1</fpage>&#x2013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1016/S0169-7439(99)00047-7</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mccamish</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Meier</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Landford</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Bass</surname>
<given-names>R. B.</given-names>
</name>
<name>
<surname>Chiu</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Cotilla-Sanchez</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>A Backend Framework for the Efficient Management of Power System Measurements</article-title>. <source>Electric Power Syst. Res.</source> <volume>140</volume> (<issue>nov</issue>), <fpage>797</fpage>&#x2013;<lpage>805</lpage>. <pub-id pub-id-type="doi">10.1016/j.epsr.2016.05.003</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Nascimento</surname>
<given-names>R. M. D.</given-names>
</name>
<name>
<surname>Oening</surname>
<given-names>A. P.</given-names>
</name>
<name>
<surname>Marcilio</surname>
<given-names>D. C.</given-names>
</name>
<name>
<surname>Alexandre</surname>
<given-names>R. A.</given-names>
</name>
<name>
<surname>J&#xfa;nior</surname>
<given-names>E. P. R.</given-names>
</name>
<name>
<surname>Schiochet</surname>
<given-names>J.&#x20;M.</given-names>
</name>
</person-group> &#x201c;<article-title>&#x201c;Outliers&#x2019; Detection and Filling Algorithms for Smart Metering Centers&#x201d;</article-title>,&#x201d; in <conf-name>Proceedings of the 2012 IEEE PES Transmission and Distribution Conference and Exposition</conf-name>, <conf-loc>Orlando, Florida, USA</conf-loc>, <conf-date>May 2012</conf-date>, <fpage>7</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1109/tdc.2012.6281659</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nemati</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Laso</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Manana</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Sant&#x27;Anna</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Nowaczyk</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Stream Data Cleaning for Dynamic Line Rating Application</article-title>. <source>Energies</source> <volume>11</volume> (<issue>8</issue>). <pub-id pub-id-type="doi">10.3390/en1101200710.3390/en11082007</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pei</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Bhatt</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Next-generation Monitoring, Analysis, and Control for the Future Smart Control center</article-title>. <source>IEEE Trans. Smart Grid</source> <volume>1</volume> (<issue>2</issue>), <fpage>186</fpage>&#x2013;<lpage>192</lpage>. <pub-id pub-id-type="doi">10.1109/TSG.2010.2053855</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Qu</surname>
<given-names>Z. Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y. W.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Qu</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Yan</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>A Data Cleaning Model for Electric Power Big Data Based on Spark Framework</article-title>. <source>Adv. Sci. Technology</source> <volume>9</volume>, <fpage>137</fpage>&#x2013;<lpage>150</lpage>. <pub-id pub-id-type="doi">10.14257/astl.2016.121.74</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rauch</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Logic of Association Rules</article-title>. <source>Appl. Intelligence</source> <volume>22</volume> (<issue>1</issue>), <fpage>9</fpage>&#x2013;<lpage>28</lpage>. <pub-id pub-id-type="doi">10.1023/B:APIN.0000047380.15356.7a</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shi</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Qiu</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Ling</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Spatio-Temporal Correlation Analysis of Online Monitoring Data for Anomaly Detection and Location in Distribution Networks</article-title>. <source>IEEE Trans. Smart Grid</source> <volume>11</volume> (<issue>2</issue>), <fpage>995</fpage>&#x2013;<lpage>1006</lpage>. <pub-id pub-id-type="doi">10.1109/TSG.2019.2929219</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Song</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Present Status and Challenges of Big Data Processing in Smart Grid</article-title>. <source>Power Syst. Technology</source> <volume>37</volume> (<issue>4</issue>), <fpage>927</fpage>&#x2013;<lpage>938</lpage>. <pub-id pub-id-type="doi">10.3969/j.issn.1006-9402.2014.05.038</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Thams</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Venzke</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Eriksson</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Chatzivasileiadis</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Efficient Database Generation for Data-Driven Security Assessment of Power Systems</article-title>. <source>IEEE Trans. Power Syst.</source> <volume>35</volume> (<issue>1</issue>), <fpage>30</fpage>&#x2013;<lpage>41</lpage>. <pub-id pub-id-type="doi">10.1109/TPWRS.2018.2890769</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Thang</surname>
<given-names>T. M.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>J.</given-names>
</name>
</person-group> <article-title>The Anomaly Detection by Using DBSCAN Clustering with Multiple Parameters</article-title>.&#x201d; in <conf-name>Proceedings of the 2011 International Conference on Information Science and Applications</conf-name>, <conf-loc>Jeju, Korea (South)</conf-loc>, <conf-date>April 2011</conf-date>, <publisher-name>IEEE</publisher-name>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1109/ICISA.2011.5772437</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Waal</surname>
<given-names>T. D.</given-names>
</name>
<name>
<surname>Pannekoek</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Scholtus</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2011</year>). <source>Handbook of Statistical Data Editing and Imputation</source>. <publisher-loc>Hoboken, New Jersey, USA</publisher-loc>: <publisher-name>John Wiley &#x26; Sons</publisher-name>. <comment>ISBN:978-0-470-54280-4</comment>. </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Integrating Model-Driven and Data-Driven Methods for Power System Frequency Stability Assessment and Control</article-title>. <source>IEEE Trans. Power Syst.</source> <volume>34</volume> (<issue>6</issue>), <fpage>4557</fpage>&#x2013;<lpage>4568</lpage>. <pub-id pub-id-type="doi">10.1109/TPWRS.2019.2919522</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Tu</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Xiao</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Gui</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Reduced-order Aggregate Model for Large-Scale Converters with Inhomogeneous Initial Conditions in Dc Microgrids</article-title>. <source>IEEE Trans. Energ. Convers.</source> <volume>36</volume> (<issue>99</issue>), <fpage>2473</fpage>&#x2013;<lpage>2484</lpage>. <pub-id pub-id-type="doi">10.1109/TEC.2021.3050434</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Hong</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Kang</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Review of Smart Meter Data Analytics: Applications, Methodologies, and Challenges</article-title>. <source>IEEE Trans. Smart Grid</source>, <volume>10</volume>, <fpage>1</fpage>. <pub-id pub-id-type="doi">10.1109/TSG.2018.2818167</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Yan</surname>
<given-names>J.&#x20;Z.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Y. C.</given-names>
</name>
</person-group> &#x201c;<article-title>Water Quality Data Outlier Detection Method Based on Spatial Series Features</article-title>,&#x201d; in <conf-name>Proceedings of the The 6th International Conference on Fuzzy Systems and Data Mining (FSDM)</conf-name>, <conf-loc>Xiamen, China</conf-loc>, <conf-date>November 2020</conf-date>. <pub-id pub-id-type="doi">10.3233/FAIA200715</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yan</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Sheng</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Du</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>An Method for Anomaly Detection of State Information of Power Equipment Based on Big Data Analysis</article-title>. <source>Proc. Csee</source> <volume>35</volume> (<issue>1</issue>), <fpage>52</fpage>&#x2013;<lpage>59</lpage>. <pub-id pub-id-type="doi">10.13334/j.0258-8013.pcsee.2015.01.007</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ye</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Z. D.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Z. Y.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>J.&#x20;G.</given-names>
</name>
<name>
<surname>Zhai</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>An Estimation Method of Energy Loss for Distribution Network Planning</article-title>. <source>Power Syst. Prot. Control.</source> <volume>17</volume>, <fpage>82</fpage>&#x2013;<lpage>86</lpage>. <pub-id pub-id-type="doi">10.3969/j.issn.1674-3415.2010.17.016</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>