<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article article-type="methods-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Energy Res.</journal-id>
<journal-title>Frontiers in Energy Research</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Energy Res.</abbrev-journal-title>
<issn pub-type="epub">2296-598X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">885673</article-id>
<article-id pub-id-type="doi">10.3389/fenrg.2022.885673</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Energy Research</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Time-Adaptive Transient Stability Assessment Based on the Gating Spatiotemporal Graph Neural Network and Gated Recurrent Unit</article-title>
<alt-title alt-title-type="left-running-head">Liu et al.</alt-title>
<alt-title alt-title-type="right-running-head">Time Adaptive Transient Stability Assessment</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Liu</surname>
<given-names>Jianfeng</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1698571/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Yao</surname>
<given-names>Chenxi</given-names>
</name>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Chen</surname>
<given-names>Lele</given-names>
</name>
</contrib>
</contrib-group>
<aff>
<institution>College of Electrical Engineering</institution>, <institution>Shanghai University of Electric Power</institution>, <addr-line>Shanghai</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1442158/overview">Chengzong Pang</ext-link>, Wichita State University, United States</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1628949/overview">Qilin Wang</ext-link>, Wichita State University, United States</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1703438/overview">Cheng Qian</ext-link>, First Affiliated Hospital of Zhengzhou University, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Jianfeng Liu, <email>bansen@sina.com</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Smart Grids, a section of the journal Frontiers in Energy Research</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>19</day>
<month>04</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>10</volume>
<elocation-id>885673</elocation-id>
<history>
<date date-type="received">
<day>28</day>
<month>02</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>17</day>
<month>03</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Liu, Yao and Chen.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Liu, Yao and Chen</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>With the continuous expansion of the UHV AC/DC interconnection scale, online, high-precision, and fast transient stability assessment (TSA) is very important for the safe operation of power grids. In this study, a transient stability assessment method based on the gating spatiotemporal graph neural network (GSTGNN) is proposed. A time-adaptive method is used to improve the accuracy and speed of transient stability assessment. First, in order to reduce the impact of dynamic topology on TSA after fault removal, GSTGNN is used to extract and fuse the key features of topology and attribute information of adjacent nodes to learn the spatial data correlation and improve the evaluation accuracy. Then, the extracted features are input into the gated recurrent unit (GRU) to learn the correlation of data at each time. Fast and accurate evaluation results are output from the stability threshold. At the same time, in order to avoid the influence of the quality of training samples, an improved weighted cross entropy loss function with the K-nearest neighbor (KNN) idea is used to deal with the unbalanced training samples. Through the analysis of an example, it is proved from the data visualization that the TSA method can effectively improve the assessment accuracy and shorten the assessment time.</p>
</abstract>
<kwd-group>
<kwd>transient stability assessment</kwd>
<kwd>gating spatiotemporal graph neural network</kwd>
<kwd>data visualization</kwd>
<kwd>K-nearest neighbor</kwd>
<kwd>gated recurrent unit</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Power system transient stability refers to the ability of each generator to maintain synchronous operation after a power system is greatly disturbed. At present, with the continuous expansion of the power system scale and the continuous growth of power consumption, accidents occur frequently when the system load reaches the limit transmission capacity. It leads the system closer to the limit of safe and stable operation. As long as the disturbance is slightly increased, the system will produce obvious voltage and frequency offsets, which will further lead to more serious transient stability problems (<xref ref-type="bibr" rid="B12">Liu et al., 2007</xref>). Therefore, the real-time assessment of transient stability after disturbance has attracted much attention.</p>
<p>Traditionally, transient stability assessment (TSA) has been modeled by solving a set of high-order differential algebraic equations (DAEs). It uses a direct method (<xref ref-type="bibr" rid="B8">Kang et al., 2021</xref>) based on a simplified model and a time domain simulation method (<xref ref-type="bibr" rid="B23">Wu and Ding, 2010</xref>) based on a trajectory model to evaluate system stability. Since 1980s, the transient energy function used by the direct method has been used to evaluate system stability (<xref ref-type="bibr" rid="B6">Hiskens and Hill., 1989</xref>; <xref ref-type="bibr" rid="B13">Owusu-Mireku and Chiang., 2018</xref>). However, in the actual large-scale AC/DC hybrid system, the direct method is based on the composition equation of the second-order simplified model of the generator, which leads to inaccurate evaluation. The time domain simulation method requires complete power grid and disturbance information, which consumes a lot of computing time.</p>
<p>With the innovative application of synchronous vector measurement technology and high-performance calculation methods, system variables can be sampled in the form of synchronization. In order to use the abovementioned technology for accurate and fast TSA, <xref ref-type="bibr" rid="B16">Song et al. (2005</xref>), <xref ref-type="bibr" rid="B4">Gomez et al. (2011</xref>), and <xref ref-type="bibr" rid="B7">Huang et al. (2019</xref>) carried out some experiments on TSA using synchronous vector measurements. It can make accurate and fast TSA for specific power systems. By extracting the relationship between the historical transient data set and the stability condition, the constraints of the system from an unstable state to a stable state can be confirmed accurately. Among them, the shallow layer neural network (<xref ref-type="bibr" rid="B19">Tang et al., 2019</xref>) had many applications, such as the support vector machine (SVM) method mentioned in <xref ref-type="bibr" rid="B21">Tian et al. (2017</xref>), which is suitable for small-scale training samples with a short evaluation time. However, there is a certain randomness in manual parameter adjustment based on experience. For different research objects, the model should have different forms and parameters to improve the accuracy of online evaluation. In addition, the poor adaptability of topology also affects the accuracy of online evaluation. In recent years, a series of deep learning methods (<xref ref-type="bibr" rid="B18">Tan et al., 2019</xref>; <xref ref-type="bibr" rid="B20">Tian et al., 2020</xref>) have been developed, which are suitable for automatically extracting data features from large samples. This method is used in TSA processes, such as the long- and short-term memory (LSTM) networks (<xref ref-type="bibr" rid="B17">Sun et al., 2020</xref>) and the gated recurrent unit (GRU) network (<xref ref-type="bibr" rid="B1">Chen and Wang., 2021</xref>) in the recurrent neural network (RNN), which rely on the timeliness of data for rapid evaluation. However, this method only considers the independent time series data of each node. The significant impact of the time-varying topology on TSA is ignored. The graph neural network (GNN) model (<xref ref-type="bibr" rid="B15">Scarselli et al., 2009</xref>) solves the problem of topological structure influence, such as the graph convolutional neural (GCN) network in <xref ref-type="bibr" rid="B11">Li et al. (2020</xref>) and the graph attention network (GAT) in <xref ref-type="bibr" rid="B28">Zong et al. (2021</xref>). They are embedded in power grid topology. When this characteristic information is input, the spatial correlation information between nodes is extracted to improve the evaluation accuracy. However, the method outputs the evaluation results at all times. It results in a large amount of computational data and prolongs the evaluation time. It is not beneficial for the stable recovery of the system.</p>
<p>The abovementioned deep learning models only consider the time-varying or topological space-varying data. They have a limitation between the time and accuracy of coordinated evaluation. At present, this model has been proven to be superior in graph data structure and time series information analysis in many fields, including traffic flow prediction (<xref ref-type="bibr" rid="B27">Zhao et al., 2020</xref>). However, this model has not been applied in the field of power systems. The complex traffic road structure is similar to the power grid structure. It is a complex network structure connected by points and lines. Therefore, the combined model of RNN and GNN provides a new solution for quickly extracting accurate dynamic spatial topology and time information.</p>
<p>Therefore, based on the two models, a gating spatiotemporal graph neural network (GSTGNN) framework with embedded topology and time series information is proposed. An adaptive method (<xref ref-type="bibr" rid="B9">Li et al., 2018</xref>) was adopted to improve the accuracy and speed of the assessment. Compared with the GAT model using multiple attention heads equally, the GSTGNN model is used in this study. The GSTGNN model is used to extract topological information of nodes and improve the accuracy of TSA. At the same time, the time requirement of TSA for a large-scale power network emergency control center is not more than 0.04&#xa0;s (<xref ref-type="bibr" rid="B3">Ding, 2016</xref>). The general conventional model is to aggregate all fixed time data for TSA. Therefore, an adaptive TSA method is proposed, which uses the GRU model to aggregate less time data and get accurate evaluation results quickly. In addition, the improved weighted cross entropy loss function of the K-nearest neighbor (KNN) method (<xref ref-type="bibr" rid="B22">Wang and Ye, 2020</xref>) was used to improve the evaluation performance. Compared with several TSA models, the proposed model can extract the sampling information of nodes more accurately and evaluate the transient stability of AC/DC heterogeneous multisource networks.</p>
<p>The arrangement of this study is as follows: the second part describes the improvement of the GAT network and puts forward the GSTGNN network. The combination of GSTGNN and GRU networks is introduced in detail to form the evaluation model of this study. The third part describes the adaptive TSA process based on the model in detail, including off-line training and online evaluation. Finally, the fourth part compares the evaluation performance of the proposed method with other methods through experiments and draws a conclusion.</p>
</sec>
<sec id="s2">
<title>2 Neural Network Framework of the Gated Spatiotemporal Graph</title>
<sec id="s2-1">
<title>2.1 Graph Neural Network and Attention Mechanism</title>
<p>The graph neural network uses the <inline-formula id="inf1">
<mml:math id="m1">
<mml:mrow>
<mml:mtext>G</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>A</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> matrix to describe the topological relationship of the power system network, where <italic>x</italic> represents the collected power grid information feature vector and <italic>A</italic> represents the adjacency matrix of the topology. The application of the graph neural network in the power system is to aggregate the information of <inline-formula id="inf2">
<mml:math id="m2">
<mml:mrow>
<mml:mtext>G</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>A</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> having a topological relationship. It can aggregate the characteristics of nodes themselves and neighbors and generate a new feature <inline-formula id="inf3">
<mml:math id="m3">
<mml:mrow>
<mml:mtext>G</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>x</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>A</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. <inline-formula id="inf4">
<mml:math id="m4">
<mml:mrow>
<mml:mtext>G</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>x</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>A</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, containing original information and topological information, has a higher correlation with the output results. <inline-formula id="inf5">
<mml:math id="m5">
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>x</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> represents a new feature vector with higher correlation with the evaluation results, as shown in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Graph of the neural network algorithm process.</p>
</caption>
<graphic xlink:href="fenrg-10-885673-g001.tif"/>
</fig>
<p>The attention mechanism is introduced to weighted summation of the features of adjacent nodes. The weight is completely determined by the features of nodes, which is not affected by the dynamic topology. Each node in the model represents a monitoring node in the grid topology. The input of the attention layer of the graph is the node feature vector set <italic>x</italic>, as given below:<disp-formula id="e1">
<mml:math id="m6">
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2192;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2192;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2192;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2192;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mi>F</mml:mi>
</mml:msup>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>
</p>
<p>Here, <inline-formula id="inf6">
<mml:math id="m7">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2192;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> is the eigenvector of node <italic>i</italic>, <italic>N</italic> is the number of system nodes, and <italic>F</italic> is the characteristic number of nodes.</p>
<p>The output of each layer is a new node feature vector set <inline-formula id="inf7">
<mml:math id="m8">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>x</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula>, as given below:<disp-formula id="e2">
<mml:math id="m9">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>x</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo stretchy="true">&#x2192;</mml:mo>
</mml:mover>
<mml:mo>,</mml:mo>
<mml:mover accent="true">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo stretchy="true">&#x2192;</mml:mo>
</mml:mover>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mover accent="true">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo stretchy="true">&#x2192;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mover accent="true">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo stretchy="true">&#x2192;</mml:mo>
</mml:mover>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mi>F</mml:mi>
</mml:msup>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>
</p>
<p>Here, <inline-formula id="inf8">
<mml:math id="m10">
<mml:msup>
<mml:mi>F</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:math>
</inline-formula> is the feature number of the new node, similar to the feature extractor.</p>
<p>For <italic>N</italic> nodes, input node features predict output new node features. The operation of <inline-formula id="inf9">
<mml:math id="m11">
<mml:mrow>
<mml:mtext>G</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>x</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>A</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> obtained from <inline-formula id="inf10">
<mml:math id="m12">
<mml:mrow>
<mml:mtext>G</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>A</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> first needs to calculate the attention coefficient <inline-formula id="inf11">
<mml:math id="m13">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, as given below:<disp-formula id="e3">
<mml:math id="m14">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>s</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mtext>exp</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>&#x3b7;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>cos</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2192;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2192;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mtext>exp</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>&#x3b7;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>cos</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2192;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2192;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>
</p>
<p>Here, the matrix <italic>W</italic> is initialized, <italic>&#x3b2;</italic> is the training parameter, cos is the cosine similarity, and <inline-formula id="inf12">
<mml:math id="m15">
<mml:mi>&#x3b7;</mml:mi>
</mml:math>
</inline-formula> is the LeakyReLU nonlinear activation function.<inline-formula id="inf13">
<mml:math id="m16">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> is determined in <xref ref-type="disp-formula" rid="e4">Eq. 4</xref> through the adjacency matrix <italic>A</italic>
<disp-formula id="e4">
<mml:math id="m17">
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mtable columnalign="left">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msub>
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mn>0</mml:mn>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msub>
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>
</p>
<p>Different attention coefficients are generated during each training. Therefore, this study proposes a network structure using a new attention head mechanism. It sets the weight for the important attention head containing topology information, that is, 0 to 1. The model pays attention to the node information of the important attention head and improves the interpretability of the model.</p>
<p>For a GAT layer with <italic>k</italic> attention heads, each attention head contains a different set of parameters <italic>W</italic> and <inline-formula id="inf14">
<mml:math id="m18">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. <inline-formula id="inf15">
<mml:math id="m19">
<mml:mrow>
<mml:msub>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents a vector composed of <italic>k</italic> attention head weights. The working process is shown in <xref ref-type="fig" rid="F2">Figure 2</xref>.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Characteristic extraction process.</p>
</caption>
<graphic xlink:href="fenrg-10-885673-g002.tif"/>
</fig>
<p>Among them, node 1 has three attention heads in its neighborhood. Different arrow styles and colors represent independent attention heads. Different soft gates aggregate and control the features of each head to get more relevant feature vectors. The weight formula of attention head is shown below:<disp-formula id="e5">
<mml:math id="m20">
<mml:mtable columnalign="left">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>F</mml:mi>
<mml:msubsup>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>m</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2295;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>g</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2295;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(5)</label>
</disp-formula>
</p>
<p>Here, the combined maximum pool and average pool are used to construct the network. <inline-formula id="inf16">
<mml:math id="m21">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the adjacent node of node <italic>i</italic>, <inline-formula id="inf17">
<mml:math id="m22">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>g</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> means to map the features of neighbor nodes to the vector of <italic>g</italic> dimensions as small as possible, <inline-formula id="inf18">
<mml:math id="m23">
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the number of neighbor nodes of node <italic>i</italic>, <inline-formula id="inf19">
<mml:math id="m24">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:msubsup>
<mml:mi>C</mml:mi>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>&#x3b1;</mml:mi>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mi>x</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represents the single and full connection layers, <inline-formula id="inf20">
<mml:math id="m25">
<mml:mi>&#x3b1;</mml:mi>
</mml:math>
</inline-formula> is the sigmoid activation function, <inline-formula id="inf21">
<mml:math id="m26">
<mml:mo>&#x2295;</mml:mo>
</mml:math>
</inline-formula> is the linker, <italic>k</italic> is the number of attention heads, <inline-formula id="inf22">
<mml:math id="m27">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>m</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> maps connected features to the <italic>k</italic>-dimensional space, and <inline-formula id="inf23">
<mml:math id="m28">
<mml:mrow>
<mml:msubsup>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> is the weight of the <italic>k</italic>th attention head of the <italic>i</italic>th node.</p>
<p>Each node <italic>i</italic> aggregates the new output features of the attention head topology information, as follows:<disp-formula id="e6">
<mml:math id="m29">
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mo>&#x2225;</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>K</mml:mi>
</mml:msubsup>
<mml:msubsup>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:munder>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mi mathvariant="italic">&#x3f5;</mml:mi>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msubsup>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mi>k</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:msup>
<mml:mi>W</mml:mi>
<mml:mi>k</mml:mi>
</mml:msup>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>
</p>
<p>Here, after <inline-formula id="inf24">
<mml:math id="m30">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> going by the GSTGNN layer, node <italic>i</italic> contains the output new features of the feature information of different spatial adjacent nodes. The attention head <italic>k</italic> satisfies inequality <inline-formula id="inf25">
<mml:math id="m31">
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2265;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>. &#x3c3; represents the GELU activation function (<xref ref-type="bibr" rid="B5">Hendrycks and Gimpel., 2016</xref>).</p>
<p>In order to prevent overfitting of the model, the random regularization idea is introduced to control the performance on the training set, as follows:<disp-formula id="e7">
<mml:math id="m32">
<mml:mtable columnalign="left">
<mml:mtr>
<mml:mtd>
<mml:mtext>GELU</mml:mtext>
<mml:mi>s</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>x</mml:mi>
<mml:mi>&#x3d5;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0.5</mml:mn>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:mfrac>
<mml:mn>2</mml:mn>
<mml:mi>&#x3c0;</mml:mi>
</mml:mfrac>
</mml:mrow>
</mml:msqrt>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>0.044715</mml:mn>
<mml:msup>
<mml:mi>x</mml:mi>
<mml:mn>3</mml:mn>
</mml:msup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(7)</label>
</disp-formula>
</p>
<p>Here, <inline-formula id="inf26">
<mml:math id="m33">
<mml:mrow>
<mml:mi>&#x3c6;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> uses N (0,1) normal distribution.</p>
<p>When <italic>x</italic> decreases, the output value will depend on the input value randomly according to the probability. It improves the generalization ability of the model.</p>
</sec>
<sec id="s2-2">
<title>2.2 Gated Recurrent Unit</title>
<p>As mentioned before, GSTGNN is introduced to extract power grid topology information to improve the evaluation accuracy. However, the evaluation time is not considered. The system studied in this study is a high-dimensional dynamic system. Not only each node has power grid topology information but also its physical quantity has time series characteristics. The GRU network has the characteristics of forgetting and selective memory. It learns and retains the timing characteristics of input data for later use. In transient assessment, the stability after fault removal can be deduced from the fault state. Therefore, the GRU layer is added to capture the historical sequence and current information of nodes and predict the future time stability. It does not need to input all time data. Thus, it can shorten the evaluation time.</p>
<p>In conclusion, in order to process the sequence information with a complex topological structure and time correlation, GSTGNN and GRU are combined to form a gated spatiotemporal GNN framework. At each time, the input <inline-formula id="inf27">
<mml:math id="m34">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and hidden state <inline-formula id="inf28">
<mml:math id="m35">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> get new features <inline-formula id="inf29">
<mml:math id="m36">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>x</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf30">
<mml:math id="m37">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>h</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> with topology information through the GSTGNN layer, as follows:<disp-formula id="e8">
<mml:math id="m38">
<mml:mtable columnalign="left">
<mml:mtr>
<mml:mtd>
<mml:mo>&#xa0;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>x</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mtext>G</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>A</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>h</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mtext>G</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>A</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(8)</label>
</disp-formula>
</p>
<p>Here, two gating states <inline-formula id="inf31">
<mml:math id="m39">
<mml:mrow>
<mml:msub>
<mml:mi>z</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf32">
<mml:math id="m40">
<mml:mrow>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are obtained by the state transmitted from the previous node <inline-formula id="inf33">
<mml:math id="m41">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> and the input of the current node <inline-formula id="inf34">
<mml:math id="m42">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula>; <inline-formula id="inf35">
<mml:math id="m43">
<mml:mrow>
<mml:msub>
<mml:mi>z</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the gating of control update, which decides the unit to update its active content to reduce the risk of gradient disappearance; <inline-formula id="inf36">
<mml:math id="m44">
<mml:mrow>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the gating of control reset, which determines the degree of combining the new input information with the previous memory state features. The expression is shown as follows:<disp-formula id="e9">
<mml:math id="m45">
<mml:mtable columnalign="left">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mi>z</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>z</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>z</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>r</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>r</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>U</mml:mi>
<mml:mi>r</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>r</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(9)</label>
</disp-formula>
</p>
<p>Here, <italic>&#x3b1;</italic> is the sigmoid activation function with a range of (0,1).</p>
<p>Through <inline-formula id="inf37">
<mml:math id="m46">
<mml:mrow>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, the new memory will store the information related to the past in the following:<disp-formula id="e10">
<mml:math id="m47">
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="true">&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mtext>tanh</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>h</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>U</mml:mi>
<mml:mi>h</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x2299;</mml:mo>
<mml:msub>
<mml:mi>U</mml:mi>
<mml:mi>h</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>h</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(10)</label>
</disp-formula>
</p>
<p>Here, <inline-formula id="inf38">
<mml:math id="m48">
<mml:mo>&#x2299;</mml:mo>
</mml:math>
</inline-formula> means the Hadamard product and multiplies the corresponding elements in the matrix to identify retained and forgotten previous information. <inline-formula id="inf39">
<mml:math id="m49">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="true">&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> remembers the state of previous time by resetting.</p>
<p>The memory state of current time step <inline-formula id="inf40">
<mml:math id="m50">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> needs <inline-formula id="inf41">
<mml:math id="m51">
<mml:mrow>
<mml:msub>
<mml:mi>z</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> to forget and select memory at the same time as follows:<disp-formula id="e11">
<mml:math id="m52">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>z</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2299;</mml:mo>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>z</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x2299;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="true">&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(11)</label>
</disp-formula>Here, <inline-formula id="inf42">
<mml:math id="m53">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>z</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2299;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> means selective &#x201c;forgetting&#x201d; of the original hidden state and removing unimportant information. <inline-formula id="inf43">
<mml:math id="m54">
<mml:mrow>
<mml:msub>
<mml:mi>z</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x2299;</mml:mo>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="true">&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> selects some information in <inline-formula id="inf44">
<mml:math id="m55">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="true">&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula>, which means forgetting the state information of the past moment and adding some state information input by the current node. In order to improve the generalization ability and evaluation accuracy of the model, topology information is included in the transmission of the abovementioned state information.</p>
<p>Finally, new features aggregated are input into the softmax separator for classification. The output prediction value <inline-formula id="inf45">
<mml:math id="m56">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> at the last time is obtained in <xref ref-type="disp-formula" rid="e12">Eq. 12</xref>. The value range is (0,1)<disp-formula id="e12">
<mml:math id="m57">
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>s</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>o</mml:mi>
</mml:msub>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>o</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(12)</label>
</disp-formula>
</p>
<p>Here, <inline-formula id="inf46">
<mml:math id="m58">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>o</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf47">
<mml:math id="m59">
<mml:mrow>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>o</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are the parameters learned by gradient back propagation.</p>
</sec>
</sec>
<sec id="s3">
<title>3 Adaptive Transient Stability Assessment</title>
<sec id="s3-1">
<title>3.1 Principle</title>
<p>In order to start the emergency control immediately after the fault is removed, TSA needs to be performed quickly and accurately. In <xref ref-type="bibr" rid="B17">Sun et al. (2020</xref>) and <xref ref-type="bibr" rid="B20">Tian et al. (2020</xref>), most of the existing TSA methods use a fixed length observation window; that is, the evaluation time is constant. However, this static evaluation time may not be able to cope with fast transient instability. The system models with different fault degrees need different observation window lengths. After the fault is removed, the dynamic data of the system are observed in the dynamic time window. The stability of the system is evaluated in the future time window, and the observation time window is gradually adjusted. As long as the evaluation system loses stability in the future, the emergency control will be started immediately.</p>
</sec>
<sec id="s3-2">
<title>3.2 Off-Line Training</title>
<sec id="s3-2-1">
<title>3.2.1 Generate Data Set</title>
<p>The purpose of the evaluation model is to obtain the evaluation results from the sampling data. The input data need to fully reflect the dynamic behavior of the system. In this study, the transient stability of the power system after fault removal is studied. <xref ref-type="bibr" rid="B14">Rajapakse et al. (2009)</xref> proposed to use PMU to sample the generator voltage amplitude after fault as input. Due to the inertia of the rotor, it takes a long time to display the change of generator rotor speed and angle after fault. In contrast, the generator voltage amplitude reflects the fault faster than the rotor angle variable. It has been verified that the voltage amplitude can accurately evaluate the transient stability of the system. <xref ref-type="bibr" rid="B4">Gomez et al. (2011)</xref> further proved that the use of voltage amplitude has a higher evaluation accuracy than mechanical variables (rotor angle and angular velocity). Therefore, this study selects voltage amplitude as the input variable. In addition, other node dynamic variables are selected to form the input data together with the voltage amplitude to improve the evaluation accuracy. <xref ref-type="bibr" rid="B10">Li et al. (2021)</xref> selected the initial characteristics reflecting the system dynamics to construct the input characteristics of TSA. The input characteristics include voltage amplitude, phase angle, injected active power, and injected reactive power of each node on all buses.</p>
<p>Analog measurements are used by 50&#xa0;Hz sampling. In order to better simulate the system response after various faults are removed, different contingencies are simulated to generate the whole data set. Input <italic>x</italic> is expressed in the form of time as follows:<disp-formula id="e13">
<mml:math id="m60">
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>T</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mrow>
<mml:mn>1,1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mo>&#x2026;</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mo>&#x2026;</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mo>&#x2026;</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mo>&#x2026;</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mn>1,1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mo>&#x2026;</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mo>&#x2026;</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mo>&#x2026;</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mo>&#x2026;</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mn>1,1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mo>&#x2026;</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mo>&#x2026;</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mo>&#x2026;</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mo>&#x2026;</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mrow>
<mml:mn>1,1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mo>&#x2026;</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mo>&#x2026;</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mo>&#x2026;</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mo>&#x2026;</mml:mo>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(13)</label>
</disp-formula>
</p>
<p>Here, <inline-formula id="inf48">
<mml:math id="m61">
<mml:mrow>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf49">
<mml:math id="m62">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, and <inline-formula id="inf50">
<mml:math id="m63">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf51">
<mml:math id="m64">
<mml:mrow>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are the node voltage amplitude, phase angle, and active and reactive power, respectively; <italic>T</italic> is the observation time window, which affects the accuracy and complexity of GRU evaluation; and <italic>N</italic> is the number of nodes, which is determined by the grid topology.</p>
<p>The transient stability index (TSI) of the power angle after fault removal is used to judge the sample stability. The expression is shown as follows:<disp-formula id="e14">
<mml:math id="m65">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x227b;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>U</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2264;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1,2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>...</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>T</mml:mi>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(14)</label>
</disp-formula>
</p>
<p>Here, <italic>t</italic> satisfies the formula <inline-formula id="inf52">
<mml:math id="m66">
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>360</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>360</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula>. <inline-formula id="inf53">
<mml:math id="m67">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the maximum power angle difference of any two synchronous generators at the end of the simulation, and <inline-formula id="inf54">
<mml:math id="m68">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the label of the real category (<xref ref-type="bibr" rid="B24">Xie et al., 2021</xref>). A complete data set is established.</p>
</sec>
<sec id="s3-2-2">
<title>3.2.2 Improved Weighted Cross Entropy Loss Function</title>
<p>In a really large power grid, because the number of stable samples is much larger than the number of unstable samples, that is, &#x201c;imbalance&#x201d;, some important unstable situations may be misjudged as stable. In practical applications, more attention is paid to the accurate evaluation of unstable samples. Conventional methods dealing with data imbalance only consider the imbalance of the number of two types of samples, while ignoring the spatial distribution information of the number of two types of samples (<xref ref-type="bibr" rid="B2">Chen et al., 2017</xref>).</p>
<p>The KNN method (<xref ref-type="bibr" rid="B22">Wang and Ye, 2020</xref>) can obtain the spatial distribution information of each sample so as to solve the problem of data imbalance. First, the distance between each sample and the nearest sample of the opposite category is calculated. It is the reference position of the sample, which is used to divide the region. Then, the number of unstable and stable data of a series of location regions is obtained, which is the spatial distribution of the sample.</p>
<p>The study sets the total space to have <italic>a</italic> areas. The formula is shown as follows:<disp-formula id="e15">
<mml:math id="m69">
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>A</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>B</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>A</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>B</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>B</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>A</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>A</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>B</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(15)</label>
</disp-formula>
</p>
<p>Here, <inline-formula id="inf55">
<mml:math id="m70">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>A</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf56">
<mml:math id="m71">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>B</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are the number of unstable and stable samples in <italic>i</italic>th region. <inline-formula id="inf57">
<mml:math id="m72">
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>A</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> (<italic>i</italic> &#x3d; 1,2, &#x2026;, <italic>a</italic>) and <inline-formula id="inf58">
<mml:math id="m73">
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>B</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> (<italic>i</italic> &#x3d; 1,2, &#x2026;, <italic>a</italic>) are the weights of samples in the <italic>i</italic>th region.</p>
<p>In this study, an improved weighted cross entropy loss function is used to increase the cost of misjudgment, as follows:<disp-formula id="e16">
<mml:math id="m74">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>N</mml:mi>
</mml:mfrac>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>A</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>l</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>B</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(16)</label>
</disp-formula>
</p>
<p>Here, <italic>N</italic> is the number of single training samples, <inline-formula id="inf59">
<mml:math id="m75">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the label of real class, and <inline-formula id="inf60">
<mml:math id="m76">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the probability of evaluation class.</p>
<p>The purpose of off-line training is to get the optimal <italic>w</italic> and <italic>b</italic> by using the Adam optimizer to train the model under the condition of minimum loss function <italic>L</italic>. The abovementioned methods are used to reduce the impact of unbalanced data.</p>
</sec>
</sec>
<sec id="s3-3">
<title>3.3 Online Evaluation</title>
<p>The stability threshold is used to evaluate the evaluation results. Because the GRU layer is used in this study, <inline-formula id="inf61">
<mml:math id="m77">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mtext>y</mml:mtext>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> only focuses on the stability index generated at time <italic>T</italic>. When <italic>i</italic> &#x3d; 1, 2,..., <italic>T</italic> &#x2212; 1 is ignored, the rule is shown as follows:<disp-formula id="e17">
<mml:math id="m78">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>y</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mtable columnalign="left">
<mml:mtr>
<mml:mtd>
<mml:mi>S</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2265;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3e;</mml:mo>
<mml:mi>&#x3b4;</mml:mi>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mi>U</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3c;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3b4;</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1,2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>...</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>T</mml:mi>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mi>U</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>k</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>w</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>O</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>s</mml:mi>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(17)</label>
</disp-formula>
</p>
<p>Here, <inline-formula id="inf62">
<mml:math id="m79">
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>0.5</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the stability threshold.</p>
<p>It is necessary to search and set the appropriate threshold <inline-formula id="inf63">
<mml:math id="m80">
<mml:mi>&#x3b4;</mml:mi>
</mml:math>
</inline-formula> to balance the accuracy of TSA with the average evaluation time. An adaptive TSA process is proposed in <xref ref-type="fig" rid="F3">Figure 3</xref>.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Adaptive TSA process.</p>
</caption>
<graphic xlink:href="fenrg-10-885673-g003.tif"/>
</fig>
<p>In this study, <inline-formula id="inf64">
<mml:math id="m81">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is input to the model according to the observation time window after fault removal. The evaluation results are predicted moment by moment. When it is within the judgment range, the results are directly output. Otherwise, the hidden state <inline-formula id="inf65">
<mml:math id="m82">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> at this time and the next time <inline-formula id="inf66">
<mml:math id="m83">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are input to the next neuron for TSA. If the evaluation time exceeds the observation time window, the sliding time window method is adopted until the reliable evaluation result. <inline-formula id="inf67">
<mml:math id="m84">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> is obtained, or the maximum allowable evaluation time <inline-formula id="inf68">
<mml:math id="m85">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is reached. If <inline-formula id="inf69">
<mml:math id="m86">
<mml:mrow>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> has not been determined, it is regarded as unstable. In this process, <inline-formula id="inf70">
<mml:math id="m87">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf71">
<mml:math id="m88">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are needed to go through the GSTGNN layer to get <inline-formula id="inf72">
<mml:math id="m89">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf73">
<mml:math id="m90">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> with topological relationship and inputed to the next evaluation time. It is found that <inline-formula id="inf74">
<mml:math id="m91">
<mml:mrow>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is generally set to 10 cycles.</p>
</sec>
<sec id="s3-4">
<title>3.4 TSA Performance Comparison</title>
<p>With the continuous development of machine learning, for the classification problem, the index based on the confusion matrix in <xref ref-type="table" rid="T1">Table 1</xref> can better evaluate the applicability of the classification model than only using the accuracy.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Confusion matrix for TSA.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Evaluation quantity</th>
<th align="center">Real stability</th>
<th align="center">Real unstability</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Prediction stability</td>
<td>True positive (<italic>TP</italic>)</td>
<td>False negative (<italic>FN</italic>)</td>
</tr>
<tr>
<td align="left">Prediction unstability</td>
<td>False positive (<italic>FP</italic>)</td>
<td>True negative (<italic>TN</italic>)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In the power system transient stability analysis, researchers pay more attention to whether the instability is classified correctly. As a result, the following evaluation indexes are mainly based on whether the instability can be correctly judged. In <xref ref-type="table" rid="T1">Table 1</xref>, <italic>TP</italic> is the number of stable samples correctly predicted, <italic>FP</italic> is the number of unstable samples incorrectly predicted, <italic>FN</italic> is the number of mispredicted stable samples, and <italic>TN</italic> is the number of unstable samples correctly predicted. In the confusion matrix, the more the number of <italic>TP</italic> and <italic>TN</italic> is, the better the number of <italic>FP</italic> and <italic>FN</italic> is. However, there are a large number of sample data. It is difficult to measure the model only by the number. Therefore, in order to comprehensively evaluate the TSA performance, four indexes are obtained as follows to calculate the accuracy (<italic>ACC</italic>), miscalculation rate (<italic>recall</italic>), <italic>precision</italic>, and comprehensive evaluation index <inline-formula id="inf75">
<mml:math id="m92">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> of the model instability.</p>
<p>
<italic>ACC</italic> is the proportion of the number of correct assessments to the total number of assessments, defined as follows:<disp-formula id="e18">
<mml:math id="m93">
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>C</mml:mi>
<mml:mi>C</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>100</mml:mn>
<mml:mo>%</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(18)</label>
</disp-formula>
</p>
<p>Once an unstable situation is misjudged as a stable situation, it may lead to a large area power outage of the system. Thus, the unstability recall is used to express the misjudgment rate. The closer the <italic>recall</italic> is to 1, the lower the possibility of misjudgment of the unstability situation. The formula of <italic>recall</italic> is shown as follows:<disp-formula id="e19">
<mml:math id="m94">
<mml:mrow>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>l</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>100</mml:mn>
<mml:mo>%</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(19)</label>
</disp-formula>
</p>
<p>Unstability <italic>precision</italic> is the proportion of correctly predicted unstable samples, defined as follows:<disp-formula id="e20">
<mml:math id="m95">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(20)</label>
</disp-formula>
</p>
<p>The higher the <italic>precision</italic> is, the more accurate it is to predict the unstable situation. However, <italic>recall</italic> and <italic>precision</italic> are inversely proportional and cannot reach a high ratio at the same time. Therefore, <inline-formula id="inf76">
<mml:math id="m96">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, the weighted harmonic average of <italic>precision</italic> and <italic>recall</italic>, is introduced in <xref ref-type="disp-formula" rid="e21">Eq. 21</xref>. The two performance indicators are comprehensively considered. <inline-formula id="inf77">
<mml:math id="m97">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is distributed between [0,1]. The closer it is to 1, the stronger the feature extraction ability of the model and the better the evaluation performance<disp-formula id="e21">
<mml:math id="m98">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(21)</label>
</disp-formula>
</p>
<p>The accuracy mainly includes <italic>ACC</italic> and <inline-formula id="inf78">
<mml:math id="m99">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. Although the deep learning algorithm studied in this study can reach 1 (or 100%) in theory, it will lead to overfitting of the model. The data model does not have enough prediction ability. Therefore, this study introduces the idea of stochastic regularization, which is as close as 100% when the prediction requirements are met.</p>
<p>Besides the accuracy, the average response time (<italic>ART</italic>) in <xref ref-type="disp-formula" rid="e22">Eq. 22</xref> is also an important index to evaluate the performance of TSA<disp-formula id="e22">
<mml:math id="m100">
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>R</mml:mi>
<mml:mi>T</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:mo>/</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(22)</label>
</disp-formula>
</p>
<p>Here, <inline-formula id="inf79">
<mml:math id="m101">
<mml:mrow>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the maximum allowable evaluation time (10 cycles). <inline-formula id="inf80">
<mml:math id="m102">
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the total number of classified instances in the current evaluation cycle. <italic>ART</italic> refers to the average TSA time after fault removal.</p>
</sec>
<sec id="s3-5">
<title>3.5 Total Evaluation Process</title>
<p>The total TSA process is shown in <xref ref-type="fig" rid="F4">Figure 4</xref>:</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>General evaluation flow chart.</p>
</caption>
<graphic xlink:href="fenrg-10-885673-g004.tif"/>
</fig>
<p>1) Off-line training: In this study, the three-phase permanent short-circuit fault is set under different fault points of different lines. It can realize the diversity of sample data. The time domain simulation method generates data, including voltage amplitude, phase angle, injected active power, injected reactive power, and the maximum power angle difference <inline-formula id="inf81">
<mml:math id="m103">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> at the end of the simulation. <inline-formula id="inf82">
<mml:math id="m104">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> obtains the stability index <italic>t</italic>; when <italic>t</italic> &#x3e; 0, it indicates the system transient stability, and the label is &#x201c;1"; on the contrary, the system is unstable and the label is &#x201c;0". This study selects the dynamic data of each node on the bus as the input data to determine the system stability in order to be limited to a certain range to reduce the difference of the data. The training data and test data are generally normalized. The weights under instability and stability are obtained from the training data so as to form a weighted cross entropy loss function. The model parameters are continuously adjusted. The Adam method is used to minimize the loss function until the optimal evaluation model is obtained.</p>
<p>2) Online evaluation: Test data and actual data are added, and performance is evaluated by using the best trained model.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Numerical Example</title>
<sec id="s4-1">
<title>4.1 Generation of Sample Data Sets</title>
<p>This study uses a representative New England 10-machine 39-node power system in <xref ref-type="fig" rid="F5">Figure 5</xref> to evaluate the performance of the method. The reference frequency is 50&#xa0;Hz. The time domain simulation data of PSA-BPA software is used, similar to the data generated by PMU in real time. Python is used to program.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>New England 10-machine 39-bus power system.</p>
</caption>
<graphic xlink:href="fenrg-10-885673-g005.tif"/>
</fig>
<p>The sampling step is set as 0.02&#xa0;s. Various faults are simulated in various scenarios to obtain complete data sets. A total of 11 load levels of 75&#x2013;125% (step size 5%) are set. Generator output is changed accordingly to ensure power flow convergence. The Three-phase permanent short-circuit fault is set for the fault type. Fault points are set for 0, 25, 50, 75, and 100% of lines. The fault clearing time is set for 0.2&#xa0;s. The simulation time is 5&#xa0;s. Fault samples are selected considering the whole wiring system and N-1 accident. The fault samples are 8855. A total of 6000 training set samples are randomly selected, and test set samples are selected according to 4:1, including 5211 stable samples and 2289 unstable samples.</p>
<p>In order to train the evaluation model with the best performance, these control parameters are defined later. In terms of the model structure, this study sets up two GSTGNN input layers to avoid excessive smoothing caused by too many layers. One GSTGNN middle layer and four attention heads are set to extract and aggregate the features with the topological structure of the power grid. The GRU layer is set to two layers, with 128 batch sizes in each layer, which further improves the evaluation speed. The last layer is the dense layer, which uses the nonlinear activation function to fit the nonlinear problem to improve the evaluation accuracy. In terms of training parameters, the dropout is set to 0.05. The learning rate is set to 0.001. The training iterations are set to 300. The aforementioned settings can improve the performance of the evaluation model.</p>
<p>In the test, each evaluation cycle of the adaptive TSA method based on the sampling step is 0.02&#xa0;s. The maximum evaluation time is set to 0.2&#xa0;s (<inline-formula id="inf83">
<mml:math id="m105">
<mml:mrow>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mrow>
<mml:mi>max</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> &#x3d; 10). The observation window <italic>T</italic> and <italic>&#x3b4;</italic> are set to 8 and 0.62, respectively. The specific setting experiment is described later.</p>
</sec>
<sec id="s4-2">
<title>4.2 Handling Unbalanced Data</title>
<p>In this study, the idea of KNN is introduced to calculate the weight of each sample. The weight is added to the training parameters in the loss function to obtain the best model. A total of 6000 samples with or without imbalance are taken as the training data. The number of iterations is 300. When <italic>T</italic> &#x3d; 8, <italic>&#x3b4;</italic> &#x3d; 62; the training process is shown in <xref ref-type="fig" rid="F6">Figure 6</xref>.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Relationship between iteration times and accuracy.</p>
</caption>
<graphic xlink:href="fenrg-10-885673-g006.tif"/>
</fig>
<p>After the fault occurs, with the increase of iteration times, the accuracy (<italic>ACC</italic>) increases rapidly until it reaches 50 times. The accuracy tends to increase gently. Finally, the evaluation accuracy of balanced processing can reach 99.237%. Therefore, data quality has an important influence on the training high-performance model.</p>
<p>For the same set of data, the performance of this method is improved compared with other deep learning models, as shown in <xref ref-type="table" rid="T2">Table 2</xref>.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Model performance comparison.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Model</th>
<th align="center">ACC/%</th>
<th align="center">Recall</th>
<th align="center">Precision</th>
<th align="center">F<sub>1</sub>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">This study</td>
<td align="char" char=".">99.82</td>
<td align="char" char=".">0.998</td>
<td align="char" char=".">0.995</td>
<td align="char" char=".">0.996</td>
</tr>
<tr>
<td align="left">GAT-GRU</td>
<td align="char" char=".">99.02</td>
<td align="char" char=".">0.983</td>
<td align="char" char=".">0.989</td>
<td align="char" char=".">0.988</td>
</tr>
<tr>
<td align="left">GAT</td>
<td align="char" char=".">98.26</td>
<td align="char" char=".">0.987</td>
<td align="char" char=".">0.988</td>
<td align="char" char=".">0.987</td>
</tr>
<tr>
<td align="left">SVM</td>
<td align="char" char=".">95.06</td>
<td align="char" char=".">0.897</td>
<td align="char" char=".">0.990</td>
<td align="char" char=".">0.941</td>
</tr>
<tr>
<td align="left">GRU</td>
<td align="char" char=".">97.80</td>
<td align="char" char=".">0.949</td>
<td align="char" char=".">0.986</td>
<td align="char" char=".">0.967</td>
</tr>
<tr>
<td align="left">LSTM</td>
<td align="char" char=".">97.51</td>
<td align="char" char=".">0.947</td>
<td align="char" char=".">0.991</td>
<td align="char" char=".">0.969</td>
</tr>
<tr>
<td align="left">CNN</td>
<td align="char" char=".">93.23</td>
<td align="char" char=".">0.945</td>
<td align="char" char=".">0.993</td>
<td align="char" char=".">0.968</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>According to <xref ref-type="table" rid="T2">Table 2</xref>, the <italic>recall</italic> of SVM which belongs to the shallow model is less than 0.9, so it cannot extract features accurately and is prone to misjudgment. GAT, LSTM, and GRU models all have higher <italic>ACC</italic>. However, when the scale of system topology expands, the <inline-formula id="inf84">
<mml:math id="m106">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> value of the GAT model is higher than that of LSTM and GRU models. It indicates the importance of topological relationship between nodes for TSA. Because the CNN model does not take into account the topological structure and time series information of the power grid, <italic>ACC</italic> is relatively low. Compared with the previous model, the <italic>ACC</italic> of the framework established in this study is closer to 100%, and the <inline-formula id="inf85">
<mml:math id="m107">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> value is relatively close to 1. It still has strong generalization performance for high-dimensional data. The framework is relatively more stable. Its performance is better than that of GAT-GRU, GAT, SVM, GRU, LSTM, and CNN models. This indicates that this framework can better mine the essential characteristics of data. Its training parameter sharing overcomes the complex problems of the traditional adaptive evaluation system.</p>
</sec>
<sec id="s4-3">
<title>4.3 Average Evaluation Time and Training Time</title>
<p>The following table shows the comparison of art and training time in <xref ref-type="table" rid="T3">Table 3</xref>.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Average evaluation time versus training time.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Method</th>
<th align="center">
<italic>ART</italic>/Cycles</th>
<th align="center">Training time/s</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">This study</td>
<td align="char" char=".">1.010</td>
<td align="char" char=".">95.3</td>
</tr>
<tr>
<td align="left">GAT-GRU</td>
<td align="char" char=".">1.360</td>
<td align="char" char=".">84.6</td>
</tr>
<tr>
<td align="left">LSTM</td>
<td align="char" char=".">1.672</td>
<td align="char" char=".">41.5</td>
</tr>
<tr>
<td align="left">GRU</td>
<td align="char" char=".">1. 555</td>
<td align="char" char=".">32.7</td>
</tr>
<tr>
<td align="left">ELM</td>
<td align="char" char=".">2.111</td>
<td align="char" char=".">156.2</td>
</tr>
<tr>
<td align="left">SVM</td>
<td align="char" char=".">4</td>
<td align="char" char=".">20.2</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The <italic>ART</italic> of the GRU model and LSTM model are basically the same (<xref ref-type="bibr" rid="B25">Yu et al., 2018</xref>). Because the parameters of the GRU model are less than those of the LSTM model, the training time is relatively short (<xref ref-type="bibr" rid="B1">Chen and Wang. 2021</xref>). The extreme learning machine (ELM) network (<xref ref-type="bibr" rid="B26">Zhang et al., 2015</xref>) needs to train 10 classifiers whose parameters are not shared. A large number of parameters lead to the longest training time. A simple SVM network structure cannot effectively extract features, resulting in the longest <italic>ART</italic>. This method extracts the topological relationship and time series information of important attention heads. This method feature extraction ability is higher than that of the GAT-GRU model, which only extracts the topological relationship of attention heads on average. Thus, the output results are easier to reach the stability threshold, and the TSA evaluation speed is faster. In the case of ensuring a high accuracy, the <italic>ART</italic> of this study is 1.01 cycles, which meets the requirements of emergency control. At the same time, although the model layers of this study and GAT-GRU are more and the offline training time is longer than those of the GRU model and LSTM model, it does not affect the effect of online evaluation. This model is still practical.</p>
</sec>
<sec id="s4-4">
<title>4.4 Observation Time Window <italic>T</italic> and Stability Threshold <italic>&#x3b4;</italic> Test</title>
<p>In this study, these two parameters <italic>T</italic> and <italic>&#x3b4;</italic> affect the evaluation performance. The following figures in <xref ref-type="fig" rid="F7">Figure 7</xref> and <xref ref-type="fig" rid="F8">Figure 8</xref> are the histograms of <italic>ACC</italic> and <italic>ART</italic>, respectively.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>
<italic>T</italic> on the relationship of accuracy.</p>
</caption>
<graphic xlink:href="fenrg-10-885673-g007.tif"/>
</fig>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>
<italic>T</italic> on the relationship of <italic>ART</italic>.</p>
</caption>
<graphic xlink:href="fenrg-10-885673-g008.tif"/>
</fig>
<p>When <italic>T &#x3d;</italic> 8, it means that after the fault is cleared, the input data of the first eight cycles are used for training and continuous time series data are used for testing. The maximum accuracy is about 99.27%, which shows that the input data affect the system performance. With the increase of <italic>T</italic>, the overall trend of <italic>ACC</italic> increases (there is a small fluctuation), and the TSA time is longer. The shorter T may damage the integrity of the input data, reducing the accuracy. Therefore, in the case of <italic>T</italic> as small as possible to ensure a high accuracy, <italic>T</italic> &#x3d; 8 is the best choice.</p>
<p>When <italic>T &#x3d;</italic> 8, <italic>&#x3b4;</italic> for the histogram of <italic>ACC</italic> and <italic>ART</italic> is obtained in <xref ref-type="fig" rid="F9">Figure 9</xref> and <xref ref-type="fig" rid="F10">Figure 10</xref>.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>
<italic>&#x3b4;</italic> on the relationship of accuracy.</p>
</caption>
<graphic xlink:href="fenrg-10-885673-g009.tif"/>
</fig>
<fig id="F10" position="float">
<label>FIGURE 10</label>
<caption>
<p>
<italic>&#x3b4;</italic> on the relationship of <italic>ART</italic>.</p>
</caption>
<graphic xlink:href="fenrg-10-885673-g010.tif"/>
</fig>
<p>As with the observation time window, the size of <italic>&#x3b4;</italic> also affects <italic>ACC</italic> and <italic>ART</italic>. The smaller <italic>&#x3b4;</italic> may lead to earlier evaluation. However, the accuracy rate may be reduced, while the larger <italic>&#x3b4;</italic> produces more accurate results at the expense of evaluation speed. As can be seen from <xref ref-type="fig" rid="F9">Figures 9</xref>, <xref ref-type="fig" rid="F10">10</xref>, the increase of <italic>&#x3b4;</italic> leads to the reduction of <italic>ACC</italic> and the increase of <italic>ART</italic>. Small <italic>ART</italic> and high <italic>ACC</italic> are the expected results, so the stability threshold of 0.62 is the best.</p>
<p>To sum up, <italic>T</italic> is set to 8. <italic>&#x3b4;</italic> is set to 0.62 to get the fastest evaluation time without reducing the accuracy.</p>
</sec>
<sec id="s4-5">
<title>4.5 Test of Topology Change</title>
<p>After the fault is removed, the adaptability of the model to topology changes is tested. Four changes are listed in the following table, and 1500 test sets are selected for the performance test in <xref ref-type="table" rid="T4">Table 4</xref>.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Results of TSA under topological structure change.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Topological condition</th>
<th align="center">ACC/%</th>
<th align="center">Recall</th>
<th align="center">Precision</th>
<th align="center">F<sub>1</sub>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Line 25&#x2013;26 cutoff</td>
<td align="char" char=".">99.01</td>
<td align="char" char=".">0.989</td>
<td align="char" char=".">0.996</td>
<td align="char" char=".">0.992</td>
</tr>
<tr>
<td align="left">Generator 2 shutdown</td>
<td align="char" char=".">98.95</td>
<td align="char" char=".">0.987</td>
<td align="char" char=".">0.986</td>
<td align="char" char=".">0.986</td>
</tr>
<tr>
<td align="left">Node load 28 resection</td>
<td align="char" char=".">99.42</td>
<td align="char" char=".">0.995</td>
<td align="char" char=".">0.991</td>
<td align="char" char=".">0.993</td>
</tr>
<tr>
<td align="left">Line 25&#x2013;26 cut off, generator 2 shutdown, and node load 28 resection</td>
<td align="char" char=".">99.25</td>
<td align="char" char=".">0.990</td>
<td align="char" char=".">0.994</td>
<td align="char" char=".">0.991</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>It is found from <xref ref-type="table" rid="T4">Table 4</xref> that the accuracy of TSA caused by generator shutdown is the lowest, but it is still more than 99% when the load is cut off. In general, the TSA accuracy of this model is still about 99% in the case of topology changes. The results show that the features selected in this study are enough to reveal the system state of the power grid in the case of large disturbance. The selected model can fully mine the transient variation law and extract data features. The model with strong robustness is suitable for complex topology.</p>
</sec>
<sec id="s4-6">
<title>4.6 Data Visualization</title>
<p>t-Distributed stochastic neighbor embedding (t-SNE) for the extracted feature data is conducive to visually verify the effectiveness of the algorithm. The following is the feature map of different layers of the GSTGNN framework in <xref ref-type="fig" rid="F11">Figure 11</xref>, <xref ref-type="fig" rid="F12">Figure 12</xref>, and <xref ref-type="fig" rid="F13">Figure 13</xref>.</p>
<fig id="F11" position="float">
<label>FIGURE 11</label>
<caption>
<p>Input layer.</p>
</caption>
<graphic xlink:href="fenrg-10-885673-g011.tif"/>
</fig>
<fig id="F12" position="float">
<label>FIGURE 12</label>
<caption>
<p>GRU layer.</p>
</caption>
<graphic xlink:href="fenrg-10-885673-g012.tif"/>
</fig>
<fig id="F13" position="float">
<label>FIGURE 13</label>
<caption>
<p>Dense layer.</p>
</caption>
<graphic xlink:href="fenrg-10-885673-g013.tif"/>
</fig>
<p>t-SNE is used to establish linear projection. The mapping relationship between approximate high-dimensional data space and low-dimensional embedded space is obtained. The high-dimensional input data are mapped to the two-dimensional space. The stability classification is explained intuitively. <xref ref-type="fig" rid="F11">Figure 11</xref>, <xref ref-type="fig" rid="F12">Figure 12</xref>, and <xref ref-type="fig" rid="F13">Figure 13</xref> show the characteristic diagrams of the input layer, GRU layer, and full dense layer of the model. With the deepening of the layer, the boundary between the two categories is more and more obvious, and the spatial overlapping data are less and less. It shows that the GSTGNN framework plays an important role in deep structure level extraction. Finally, it achieves an obvious classification effect.</p>
</sec>
</sec>
<sec id="s5">
<title>5 Conclusion</title>
<p>In this study, a power system adaptive TSA method based on GSTGNN is proposed. It learns spatial and temporal correlation of data to balance assessment accuracy and assessment time. The New England 10-machine 39-bus system is used for verification, compared with a variety of verification algorithms to indicate the following:<list list-type="simple">
<list-item>
<p>1) This study introduces the graph attention depth learning network and considers the topological relationship between nodes. It also proposes a new attention head mechanism, which considers the feature correlation of adjacent nodes from important attention heads. In the process of aggregation, the new characteristics of nodes will change according to the changes of topology and attention coefficient. Therefore, the use of GNN can improve the evaluation accuracy and is more suitable for the changing and complex power grid structure.</p>
</list-item>
<list-item>
<p>2) This study uses adaptive TSA and GRU layers to capture the characteristic quantity of input data every moment. As long as the output data reach the threshold, the result is directly output. Thus, we can achieve rapid and accurate assessment with less data.</p>
</list-item>
<list-item>
<p>3) In addition to TSA performance and the response time test, two basic parameters are introduced for the preliminary sensitivity test: stability threshold and training observation window length. The simulation results show that the parameter configuration has good performance to promote the improvement of evaluation performance.</p>
</list-item>
</list>
</p>
<p>In conclusion, the composite framework is specially used to model sequence data with complex topology and time correlation. It reduces the training time and response time without sacrificing the evaluation accuracy. It is suitable for a more complex large-scale power grid research. However, the actual noise interference will affect the accuracy of the evaluation results, which needs to be considered in the future.</p>
</sec>
</body>
<back>
<sec id="s6">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/Supplementary Materials, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>Conceptualization, JL; methodology, CY and LC; investigation, CY; software, CY and LC; resources, JL; data curation, LC; writing&#x2014;original draft preparation, CY and LC; writing&#x2014;review and editing, CY; visualization, LC; project administration, JL. All authors have read and agreed to the published version of the manuscript.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This study is supported by the National Natural Science Foundation of China (51807114).</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations or those of the publisher, the editors, and the reviewers. Any product that may be evaluated in this article or claim that may be made by its manufacturer is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Time-adaptive Transient Stability Assessment Based on Gated Recurrent Unit</article-title>. <source>Int. J. Electr. Power Energ. Syst.</source> <volume>133</volume> (<issue>133</issue>), <fpage>107156</fpage>. <pub-id pub-id-type="doi">10.1016/j.ijepes.2021.107156</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Kirschen</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Model-Free Renewable Scenario Generation Using Generative Adversarial Networks</article-title>. <source>IEEE Trans. Power Syst.</source> <volume>33</volume> (<issue>99</issue>), <fpage>3265</fpage>&#x2013;<lpage>3275</lpage>. <pub-id pub-id-type="doi">10.1109/TPWRS.2018.2794541</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ding</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2016</year>). <source>Research on Fast Algorithm of Power System Transient Stability</source>. <publisher-loc>Zhejiang, China</publisher-loc>: <publisher-name>Zhejiang University</publisher-name>. </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gomez</surname>
<given-names>F. R.</given-names>
</name>
<name>
<surname>Rajapakse</surname>
<given-names>A. D.</given-names>
</name>
<name>
<surname>Annakkage</surname>
<given-names>U. D.</given-names>
</name>
<name>
<surname>Fernando</surname>
<given-names>I. T.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Support Vector Machine-Based Algorithm for Post-fault Transient Stability Status Prediction Using Synchronized Measurements</article-title>. <source>IEEE Trans. Power Syst. A Publ. Power Eng. Soc.</source> <volume>26</volume> (<issue>03</issue>), <fpage>1474</fpage>&#x2013;<lpage>1483</lpage>. <pub-id pub-id-type="doi">10.1109/pes.2011.6038936</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Hendrycks</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Gimpel</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2016</year>). <source>Gaussian Error Linear Unit (GELUs)</source>. <comment>arXiv:1606.08415</comment>. </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hiskens</surname>
<given-names>I. A.</given-names>
</name>
<name>
<surname>Hill</surname>
<given-names>D. J.</given-names>
</name>
</person-group> (<year>1989</year>). <article-title>Energy Functions, Transient Stability and Voltage Behaviour in Power Systems with Nonlinear Loads</article-title>. <source>IEEE Trans. Power Syst.</source> <volume>4</volume> (<issue>4</issue>), <fpage>1525</fpage>&#x2013;<lpage>1533</lpage>. <pub-id pub-id-type="doi">10.19783/j.cnki.pspc.200384</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>S. Y.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y. C.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Online Assessment for Transient Stability Based on Response Time Series of Wide-Area Measurement System</article-title>. <source>Power Syst. Techn.</source> <volume>43</volume> (<issue>03</issue>), <fpage>286</fpage>&#x2013;<lpage>295</lpage>. <pub-id pub-id-type="doi">10.13335/j.1000-3673.pst.2018.0960</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kang</surname>
<given-names>Z. R.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>M. Q.</given-names>
</name>
<name>
<surname>Gan</surname>
<given-names>D. Q.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Research on Network Voltage Analysis Algorithm Suitable for Power System Transient Stability Analysis</article-title>. <source>Power Syst. Prot. Control.</source> <volume>49</volume> (<issue>03</issue>), <fpage>32</fpage>&#x2013;<lpage>38</lpage>. <pub-id pub-id-type="doi">10.19783/j.cnki.pspc.200384</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Pan</surname>
<given-names>F. L.</given-names>
</name>
<name>
<surname>Mei</surname>
<given-names>Q. J.</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>P. Y.</given-names>
</name>
<name>
<surname>Pen</surname>
<given-names>Y. H.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>X. F.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <source>A Time-Aadaptive Method for On-Line Transient Stability Assessment of Power Systems</source>. <publisher-loc>China</publisher-loc>: <publisher-name>The People&#x27;s Republic of China. State Intellectual Property Office</publisher-name>. <comment>Patent CN 107993012A</comment>. </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Lei</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y. L.</given-names>
</name>
<name>
<surname>Xiao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yan</surname>
<given-names>B. J.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Transient Stability Assessment Method for Power System Based on Deep Forest</article-title>. <source>Electr. Meas. Instrumentation</source> <volume>58</volume> (<issue>02</issue>), <fpage>53</fpage>&#x2013;<lpage>58</lpage>. <pub-id pub-id-type="doi">10.19753/j.issn1001-1390.2021.02.009</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Xiao</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Extraction of Power Lines and Pylons from LiDAR Point Clouds Using a GCN-Based Method</article-title>,&#x201d; in <source>IGARSS 2020 - 2020 IEEE International Geoscience and Remote Sensing Symposium</source>. <publisher-loc>Waikoloa, HI, USA</publisher-loc>: <publisher-name>IEEE (Institute of Electrical and Electronics Engineers)</publisher-name>, <fpage>2767</fpage>&#x2013;<lpage>2770</lpage>. <pub-id pub-id-type="doi">10.1109/IGARSS39084.2020.9323218</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>J. Q.</given-names>
</name>
<name>
<surname>Tao</surname>
<given-names>J. Q.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>X. W.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>H. P.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>A Survey on Research of Load Model for Stability Analysis in Foreign Countries</article-title>. <source>Power Syst. Techn.</source> <volume>31</volume> (<issue>04</issue>), <fpage>11</fpage>&#x2013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1002/jrs.1570</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Owusu-Mireku</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Chiang</surname>
<given-names>H.-D.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>A Direct Method for the Transient Stability Analysis of Transmission Switching Events</article-title>,&#x201d; in <source>IEEE Power &#x26; Energy Society General Meeting</source>. (<publisher-loc>Portland, OR, United States</publisher-loc>: <publisher-name>IEEE (Institute of Electrical and Electronics Engineers)</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1109/PESGM.2018.8586242</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Rajapakse</surname>
<given-names>A. D.</given-names>
</name>
<name>
<surname>Gomez</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Nanayakkara</surname>
<given-names>O. M. K. K.</given-names>
</name>
<name>
<surname>Crossley</surname>
<given-names>P. A.</given-names>
</name>
<name>
<surname>Terzija</surname>
<given-names>V. V.</given-names>
</name>
</person-group> (<year>2009</year>). &#x201c;<article-title>Rotor Angle Stability Prediction Using Post-disturbance Voltage Trajectory Patterns</article-title>,&#x201d; in <source>IEEE Power &#x26; Energy Society General Meeting IEEE</source>, <fpage>1</fpage>&#x2013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1109/PES.2009.5275270</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Scarselli</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Gori</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Tsoi</surname>
<given-names>A. C.</given-names>
</name>
<name>
<surname>Hagenbuchner</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Monfardini</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>The Graph Neural Network Model</article-title>. <source>IEEE Trans. Neural Netw.</source> <volume>20</volume> (<issue>1</issue>), <fpage>61</fpage>&#x2013;<lpage>80</lpage>. <pub-id pub-id-type="doi">10.1109/TNN.2008.2005605</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Song</surname>
<given-names>F. F.</given-names>
</name>
<name>
<surname>Bi</surname>
<given-names>T. S.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Q. X.</given-names>
</name>
</person-group> (<year>2005</year>). &#x201c;<article-title>Study on Wide Area Measurement System Based Transient Stability Control for Power System</article-title>,&#x201d; in <conf-name>International Power Engineering Conference</conf-name>. (<publisher-loc>Singapore</publisher-loc>: <publisher-name>IEEE (Institute of Electrical and Electronics Engineers)</publisher-name>), <fpage>757</fpage>&#x2013;<lpage>760</lpage>. <pub-id pub-id-type="doi">10.1109/IPEC.2005.207008</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname>
<given-names>L. X.</given-names>
</name>
<name>
<surname>Bai</surname>
<given-names>J. T.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Z. Y.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>C. Y.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Transient Stability Assessment of Power System Based on Bi-directional Long-Short-Term Memory Network</article-title>. <source>Automation Electric Power Syst.</source> <volume>44</volume> (<issue>13</issue>), <fpage>64</fpage>&#x2013;<lpage>72</lpage>. <pub-id pub-id-type="doi">10.7500/AEPS20191225003</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Tan</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Xiao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>A Novel Temporal Feature Selection for Time-Adaptive Transient Stability Assessment</article-title>,&#x201d; in <source>IEEE PES Innovative Smart Grid Technologies Europe (ISGT-Europe)</source>. (<publisher-loc>Bucharest, Romania</publisher-loc>: <publisher-name>IEEE (Institute of Electrical and Electronics Engineers)</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1109/ISGTEurope.2019.8905487</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Cui</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Review on Artificial Intelligence in Power System Transient Stability Analysis</article-title>. <source>Chin. J. Electr. Eng.</source> <volume>39</volume> (<issue>01</issue>), <fpage>2</fpage>&#x2013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.13334/j.0258-8013.pcsee.180706</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tian</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>X. X.</given-names>
</name>
<name>
<surname>Shi</surname>
<given-names>D. Y.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>Y. H.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Z. H.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A Preventive Control Method of Power System Transient Stability Based on a Convolutional Neural Network</article-title>. <source>Power Syst. Prot. Control.</source> <volume>48</volume> (<issue>18</issue>), <fpage>1</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.19783/j.cnki.pspc.191310</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tian</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>X. X.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Z. H.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Power System Transient Stability Assessment Based on Comprehensive SVM Classification Model and Key Sample Set</article-title>. <source>Power Syst. Prot. Control.</source> <volume>45</volume> (<issue>22</issue>), <fpage>1</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.7667/PSPC161864</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Ye</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Transient Stability Evaluation Model Based on SSDAE with Imbalanced Correction</article-title>. <source>IET Generation, Transm. Distribution</source> <volume>14</volume> (<issue>11</issue>), <fpage>2209</fpage>&#x2013;<lpage>2216</lpage>. <pub-id pub-id-type="doi">10.1049/iet-gtd.2019.1388</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>H. B.</given-names>
</name>
<name>
<surname>Ding</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Newton Method with Variable Step Size for Power System Transient Stability Simulation</article-title>. <source>Chin. J. Electr. Eng.</source> <volume>30</volume> (<issue>7</issue>), <fpage>36</fpage>&#x2013;<lpage>41</lpage>. <pub-id pub-id-type="doi">10.13334/j.0258-8013.pcsee.2010.07.006</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xie</surname>
<given-names>P. Y.</given-names>
</name>
<name>
<surname>Yuan</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Y. G.</given-names>
</name>
<name>
<surname>Pan</surname>
<given-names>F. L.</given-names>
</name>
<name>
<surname>Ye</surname>
<given-names>W. H.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Transient Stability Assessment Method in Power System Based on Active Learning</article-title>. <source>Electr. Meas. Instrumentation.</source> <volume>58</volume> (<issue>05</issue>), <fpage>86</fpage>&#x2013;<lpage>91</lpage>. <pub-id pub-id-type="doi">10.19753/j.issn1001-1390.2021.05.012</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hill</surname>
<given-names>D. J.</given-names>
</name>
<name>
<surname>Lam</surname>
<given-names>A. Y. S.</given-names>
</name>
<name>
<surname>Gu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>V. O. K.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Intelligent Time-Adaptive Transient Stability Assessment System</article-title>. <source>IEEE Trans. Power Syst.</source> <volume>33</volume> (<issue>01</issue>), <fpage>1049</fpage>&#x2013;<lpage>1058</lpage>. <pub-id pub-id-type="doi">10.1109/tpwrs.2017.2707501</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Dong</surname>
<given-names>Z. Y.</given-names>
</name>
<name>
<surname>Wong</surname>
<given-names>K. P.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Post&#x2010;disturbance Transient Stability Assessment of Power Systems by a Self&#x2010;adaptive Intelligent System</article-title>. <source>IET Generation, Transm. Distribution</source> <volume>9</volume> (<issue>3</issue>), <fpage>296</fpage>&#x2013;<lpage>305</lpage>. <pub-id pub-id-type="doi">10.1049/iet-gtd.2014.0264</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>T.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>T-GCN: A Temporal Graph Convolutional Network for Traffic Prediction</article-title>. <source>IEEE Trans. Intell. Transport. Syst.</source> <volume>21</volume> (<issue>9</issue>), <fpage>3848</fpage>&#x2013;<lpage>3858</lpage>. <pub-id pub-id-type="doi">10.1109/tits.2019.2935152</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zong</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Guan</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>Y. S.</given-names>
</name>
<name>
<surname>Yao</surname>
<given-names>H. C.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>J. Y.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>M. X.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Power System Transient Stability Assessment Based on Graph Attention Deep Network</article-title>. <source>Power Syst. Techn.</source> <volume>45</volume> (<issue>06</issue>), <fpage>2122</fpage>&#x2013;<lpage>2130</lpage>. <pub-id pub-id-type="doi">10.13335/j.1000-3673.pst.2020.0897</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>