<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Energy Res.</journal-id>
<journal-title>Frontiers in Energy Research</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Energy Res.</abbrev-journal-title>
<issn pub-type="epub">2296-598X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">772508</article-id>
<article-id pub-id-type="doi">10.3389/fenrg.2021.772508</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Energy Research</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A Hybrid Model for Power Consumption Forecasting Using VMD-Based the Long Short-Term Memory Neural Network</article-title>
<alt-title alt-title-type="left-running-head">Ruan et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Load Forecasting with VMD-LSTM</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Ruan</surname>
<given-names>Yingjun</given-names>
</name>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Wang</surname>
<given-names>Gang</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/1476182/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Meng</surname>
<given-names>Hua</given-names>
</name>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Qian</surname>
<given-names>Fanyue</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1471030/overview"/>
</contrib>
</contrib-group>
<aff>
<institution>School of Mechanical Engineering, Tongji University</institution>, <addr-line>Shanghai</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1303271/overview">Jian Chai</ext-link>, Xidian University, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1479451/overview">Zheng Qian</ext-link>, Beihang University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/806405/overview">Gabriel Mendon&#xe7;a Paiva</ext-link>, Universidade Federal de Goi&#xe1;s, Brazil</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Fanyue Qian, <email>qianfanyue91@163.com</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Smart Grids, a section of the journal Frontiers in Energy Research</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>07</day>
<month>01</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>9</volume>
<elocation-id>772508</elocation-id>
<history>
<date date-type="received">
<day>08</day>
<month>09</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>06</day>
<month>12</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Ruan, Wang, Meng and Qian.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Ruan, Wang, Meng and Qian</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Energy consumption prediction is a popular research field in computational intelligence. However, it is difficult for general machine learning models to handle complex time series data such as building energy consumption data, and the results are often unsatisfactory. To address this difficulty, a hybrid prediction model based on modal decomposition was proposed in this paper. For data preprocessing, the variational mode decomposition (VMD) technique was used to used to decompose the original sequence into more robust subsequences. In the feature selection, the maximum relevance minimum redundancy (mRMR) algorithm was chosen to analyse the correlation between each component and the individual features while eliminating the redundancy between individual features. In the forecasting module, the long short-term memory (LSTM) neural network model was used to predict power consumption. In order to verify the performance of the proposed model, three categories of contrast methods were applied: 1) Comparing the hybrid model to a single predictive model, 2) Comparing the hybrid model with the backpropagation neural network (BPNN) to the hybrid model with the LSTM and 3) Comparing the hybrid model using mRMR and the hybrid model using mutual information maximization (MIM). The experimental results on the measured data of an office building in Qingdao show that the proposed hybrid model can improve the prediction accuracy and has better robustness compared to VMD-MIM-LSTM. In the three control groups mentioned above, the <italic>R</italic>
<sup>2</sup> value of the hybrid model improved by 10, 3 and 3%, respectively, the values of the mean absolute error (MAE) decreased by 48.9, 41.4 and 35.6%, respectively, and the root mean square error (RMSE) decreased by 54.7, 35.5 and 34.1%, respectively.</p>
</abstract>
<kwd-group>
<kwd>load forecasting</kwd>
<kwd>variational mode decomposition</kwd>
<kwd>feature selection</kwd>
<kwd>machine learning</kwd>
<kwd>deep learning</kwd>
</kwd-group>
<contract-num rid="cn001">No.2020YFD1100504-05</contract-num>
<contract-sponsor id="cn001">National Key Research and Development Program of China<named-content content-type="fundref-id">10.13039/501100012166</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Energy is critical in modern society, and energy consumption is a major issue that has long plagued humanity. Increasing demand for energy is gradually drawing attention to energy conservation issues around the world. Among energy sources, building electricity consumption accounts for a large proportion of total social energy consumption. From a global perspective, building energy consumption accounts for about 40% of the global energy consumption, and this proportion is likely to increase in the future.</p>
<p>Scientists have explored various methods for predicting building electricity consumption, aiming to achieve intelligent energy management and energy-saving building reconstruction based on predicted energy consumption. However, building electricity forecasting continues to be a challenging effort due to the variety of factors that affect energy consumption, such as building structure, equipment, weather conditions, and energy-use behaviours of the building occupants.</p>
<p>Building electricity consumption predictions can be divided into three methods according to the type of data input and processing method used: White-box physics-based models, grey-box reduced-order models and black-box data-driven models.</p>
<p>White-box physics-based models rely on thermodynamic rules for detailed energy modelling and analysis. The construction of the physical model requires a large number of physical parameters related to the building and a detailed setting of the system operation. Its accuracy depends on the input parameters and the selected simulation software. Zhu et&#x20;al. (<xref ref-type="bibr" rid="B38">Zhu et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B27">Said, 2016</xref>) compared the Dest, Energy Plus and DOE-2 simulation software calculation methods, and their research results showed that the difference of load between the simulation results of Dest and Energy Plus was less than 10%. However, some detailed architectural data may not be readily available to researchers, resulting in an inability to provide accurate inputs and thus leading to poor predictive performance.</p>
<p>Grey-box modelling approaches offer a combination of physical and data-driven prediction models, leveraging the advantages and minimizing the disadvantages of both approaches. In grey-box models, some internal parameters and equations are physically interpretable (<xref ref-type="bibr" rid="B5">Eom et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B1">Amasyali and El-Gohary, 2018</xref>). Grey-box models may also show better performance compared to black-box and white-box models. For example, Dong et&#x20;al. (<xref ref-type="bibr" rid="B3">Dong et&#x20;al., 2016</xref>) developed a hybrid model which coupled a data-driven model and a thermal network model for predicting the total energy consumption of residential areas and compared its prediction performance to artificial neural networks (ANN), support vector machines (SVM) and least square support vector machine (LSSVM)-based models.</p>
<p>Unlike physical models, black-box data-driven models do not require detailed building data, but rather they learn from the available historical data to make predictions. Common machine learning algorithms include SVM and ANN. These algorithms have a wide range of applications in the field of energy consumption prediction. Currently, about 47% of studies use ANN to predict energy consumption (<xref ref-type="bibr" rid="B19">Liu et&#x20;al., 2019</xref>). For example, Mansoor et&#x20;al. (<xref ref-type="bibr" rid="B21">Muhammad et&#x20;al., 2020</xref>) compared two different neural network models, feed-forward neural networks (FFNN) and echo state networks (ESN) for electrical load forecasting in real commercial buildings; their results indicated that the ESN model generally performed slightly better than the FFNN model Katarina. Liu et&#x20;al. (<xref ref-type="bibr" rid="B18">Liu et&#x20;al., 2020</xref>) proposed a hybrid forecasting model that combined the Jaya algorithm and SVM. In this model, the representative features of the input data were selected and the hyper-parameters of SVM were optimized by using the Jaya optimization algorithm to efficiently improve the forecasting accuracy of wind speed. Mendon&#xe7;a et&#x20;al. (<xref ref-type="bibr" rid="B2">de Paiva et&#x20;al., 2020</xref>) investigated the application of machine learning models for solar radiation intensity prediction. They evaluated multigene genetic programming (MGGP) and the multilayer perceptron (MLP) ANN. The results showed that MGGP produced better results in the case of a single prediction, while ANN presented more accurate results for ensemble forecasting. Anderso et&#x20;al. (<xref ref-type="bibr" rid="B20">Marcello Anderson et&#x20;al., 2017</xref>) applied portfolio theory to solar and wind energy forecasting to improve resource forecasting for specific solar and wind energy conditions in the Brazilian region. Their study showed that the optimal combination of 30% solar and 70% wind resources generated the smallest calculated standard deviation.</p>
<p>However, the original time series were often unstable due to the disturbance of uncertainty. For this type of data, a single model did not produce excellent results (<xref ref-type="bibr" rid="B8">He et&#x20;al., 2018</xref>). To improve the prediction accuracy, the segregation of these series with different frequencies from the energy data was considered as a possible solution.</p>
<p>Empirical mode decomposition (EMD) was proposed by Dr Norden E. Huang in 1998 (<xref ref-type="bibr" rid="B11">Huang Norden et&#x20;al., 1998</xref>) as a method for processing nonstationary signals; it is an adaptive time-frequency localization analysis method, the number of decomposed IMFs depends on the data itself. Liu et&#x20;al. (<xref ref-type="bibr" rid="B16">Liu et&#x20;al., 2012</xref>) proposed a standard hybridization of EMD with the backpropagation neural network (BPNN) method. In this study, all intrinsic mode functions (IMFs) and the residue were forecasted with BPNN models. Similarly, Guo et&#x20;al. (<xref ref-type="bibr" rid="B6">Guo et&#x20;al., 2011</xref>) proposed a modified EMD&#x2013;FFNN model in the form of an EMD-based FFNN ensemble learning paradigm. This study showed that the first IMF containing high-frequency components was mostly unsymmetrical and disordered, which led to the generation of large forecasting disturbances. The simplest combinations of hybrid EMD&#x2013;SVM models are presented in the literature (<xref ref-type="bibr" rid="B15">Lin and Peng, 2011</xref>; <xref ref-type="bibr" rid="B36">Zhang et&#x20;al., 2015</xref>). These models decomposed wind data into a series of components (IMFs) using EMD, and then different models were built with various kernel functions and parameters for each component using the SVM&#x20;model.</p>
<p>However, the IMF components obtained by EMD often exhibit mode mixing, resulting in inaccurate IMF components. To solve this problem, many scholars have proposed improved algorithms. Wu and Huang (<xref ref-type="bibr" rid="B33">Wu and Huang, 2009</xref>) suggested the ensemble empirical mode decomposition (EEMD) method. Numerous articles in distinct research areas have claimed the superior performance of the EEMD method over hybrid EMD models. The hybrid EEMD&#x2013;SVM model has been used in the literature and has achieved better prediction accuracy than other models (<xref ref-type="bibr" rid="B10">Hu et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B31">Wu et&#x20;al., 2018</xref>). In one work (<xref ref-type="bibr" rid="B31">Wu et&#x20;al., 2018</xref>), wind speed data was decomposed into seven IMF components with EEMD and then the IMFs were predicted using the appropriate SVM models. Elsewhere, a similar approach was used in which the first IMF (IMF1) was removed from the prediction analysis and all remaining IMFs were forecasted with SVM models (<xref ref-type="bibr" rid="B10">Hu et&#x20;al., 2013</xref>). Yu et&#x20;al. (<xref ref-type="bibr" rid="B31">Wu et&#x20;al., 2018</xref>) proposed a novel model based on EEMD and LSTM for crude oil price forecasting. In this study, a method to select the same number of proper inputs in various decomposition scenarios was developed. To extract features from the selected components more adequately, LSTM was introduced as a forecasting method to predict price movement directly. Dragomiretskiy and Zosso (<xref ref-type="bibr" rid="B4">Dragomiretskiy and Zosso, 2014</xref>) introduced the variational mode decomposition (VMD) method in 2014. The VMD algorithm is more robust in that it inherits the advantages of the EMD algorithm while solving the mode mixing problem of the EMD algorithm. In recent years, the VMD algorithm has been successfully applied in many fields, such as fault diagnosis research (<xref ref-type="bibr" rid="B35">Zhang et&#x20;al., 2017</xref>) and forecast research (<xref ref-type="bibr" rid="B17">Liu et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B22">Niu et&#x20;al., 2020</xref>). The studies of He (<xref ref-type="bibr" rid="B7">He et&#x20;al., 2019</xref>) and Li (<xref ref-type="bibr" rid="B14">Li et&#x20;al., 2018</xref>) have shown that the combination model based on &#x201c;decomposition-prediction&#x201d; can achieve high prediction accuracy in heating and cooling seasons. He et&#x20;al. developed a VMD-LSTM forecasting model for electricity load forecasting in Hubei province. They divided the 1-year data into four parts, corresponding to four seasons. The results show that the proposed forecasting model has high forecasting accuracy on all four data sets. The lowest prediction accuracy is found in summer, attributed to the higher fluctuation and uncertainty of load in summer.</p>
<p>Studies using signal decomposition methods have some shortcomings. Firstly, some literature uses different prediction methods for different IMF frequencies while ignoring the feature selection variability of IMFs. Secondly, it is difficult to provide a reasonable explanation for the physical meaning of each component using signal decomposition methods. To address these inadequacies, a hybrid system was developed that comprises three modules to predict the electricity load of public buildings in Qingdao. Compared with existing studies on short-term load forecasting, the main contributions of this paper are as follows:<list list-type="simple">
<list-item>
<p>1) A novel deep learning-based method for predicting building electricity consumption is proposed. The idea of &#x2018;&#x2018;decomposition&#x2013;reconstruction&#x2013;integration&#x201d; results in a feasible and efficient method to model and forecast nonlinear, non-stationary, complex time series.</p>
</list-item>
<list-item>
<p>2) Due to the volatility and uncertainty of the load data, VMD is used to decompose the raw load into more stable series. Most of the literature does not detail the determination of the number of VMD components (<xref ref-type="bibr" rid="B29">Sun et&#x20;al., 2019b</xref>). In this paper, the mean value of the instantaneous frequency of each component is used to determine the number of&#x20;K.</p>
</list-item>
<list-item>
<p>3) Most of the literature does not provide a reasonable interpretation of the components decomposed by the modal decomposition algorithm. In this paper, the highly volatile load is decomposed into several subsequences by VMD. The redundancy between features is removed by the mRMR algorithm so that each subsequence has a suitable feature. With the features selected by mRMR, this paper attempts to analyse the physical meaning of each subsequence.</p>
</list-item>
</list>
</p>
<p>This study is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> outlines the principles of the methods related to the proposed hybrid system. In addition, a case study is presented in <xref ref-type="sec" rid="s3">Section 3</xref>. Finally, the study&#x2019;s conclusions and avenues for future work are presented in <xref ref-type="sec" rid="s4">Section 4</xref>. Note that the data decomposition and feature selection were performed on a laptop with an Intel(R) i5-7400 CPU with MATLAB 2020a installed, and the deep learning model was performed on a laptop with an Intel(R) i5-7400 CPU with Python 3.8 installed.</p>
</sec>
<sec id="s2">
<title>2 Methodology</title>
<p>The main contents of this section introduce the algorithm used in this paper: Variational mode decomposition (VMD), Max-Relevance and Min-Redundancy (mRMR) and Long Short-Term Memory Neural Network (LSTM). The flow chart of the hybrid model is shown in <xref ref-type="fig" rid="F1">Figure&#x20;1</xref>.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Hybrid model flow chart.</p>
</caption>
<graphic xlink:href="fenrg-09-772508-g001.tif"/>
</fig>
<sec id="s2-1">
<title>2.1 Principle of Variational Mode Decomposition</title>
<sec id="s2-1-1">
<title>2.1.1 VMD Principle</title>
<p>In order to solve the modal mixing problem existing in EMD, Dragomiretskiy et&#x20;al. (<xref ref-type="bibr" rid="B4">Dragomiretskiy and Zosso, 2014</xref>) proposed the VMD algorithm, which is essentially a set of adaptive Wiener filter sets. The decomposition number K of VMD is determined artificially. Theoretically, if the value of <italic>K</italic> is more reasonable, it can effectively suppress the modal mixing phenomenon. the main process of VMD is divided into five steps:<list list-type="simple">
<list-item>
<p>1) Suppose <inline-formula id="inf1">
<mml:math id="m1">
<mml:mrow>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the <italic>K</italic>th order mode of the original signal <inline-formula id="inf2">
<mml:math id="m2">
<mml:mi>f</mml:mi>
</mml:math>
</inline-formula> and <inline-formula id="inf3">
<mml:math id="m3">
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is a Dirac distribution. The analytic signal of the mode <inline-formula id="inf4">
<mml:math id="m4">
<mml:mrow>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is calculated by Hilbert transform, then its unilateral frequency spectrum can be expressed&#x20;as:</p>
</list-item>
</list>
<disp-formula id="e1">
<mml:math id="m5">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mi>j</mml:mi>
<mml:mrow>
<mml:mi>&#x3c0;</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mtext>&#x2a;</mml:mtext>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>
<list list-type="simple">
<list-item>
<p>2) Adding a pre-estimated center frequency to the resolved signal of the mode, the frequency of the mode can be modulated to the corresponding baseband:</p>
</list-item>
</list>
<disp-formula id="e2">
<mml:math id="m6">
<mml:mrow>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mi>j</mml:mi>
<mml:mrow>
<mml:mi>&#x3c0;</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mtext>&#x2a;</mml:mtext>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>j</mml:mi>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>
<list list-type="simple">
<list-item>
<p>3) Calculating the bandwidth of each modal signal, the constrained optimization problem is expressed&#x20;as:</p>
</list-item>
</list>
<disp-formula id="e3">
<mml:math id="m7">
<mml:mrow>
<mml:munder>
<mml:mrow>
<mml:mtext>min</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:munder>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>k</mml:mi>
</mml:munderover>
<mml:mo>&#x2016;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mi>j</mml:mi>
<mml:mrow>
<mml:mi>&#x3c0;</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mtext>&#x2a;</mml:mtext>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>j</mml:mi>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msup>
<mml:msubsup>
<mml:mo>&#x2016;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>where the constraint of <xref ref-type="disp-formula" rid="e3">Eq. 3</xref> is: <inline-formula id="inf5">
<mml:math id="m8">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>k</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>.<list list-type="simple">
<list-item>
<p>4) The Lagrangian function <inline-formula id="inf6">
<mml:math id="m9">
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> and quadratic penalty factor <inline-formula id="inf7">
<mml:math id="m10">
<mml:mi>&#x3b1;</mml:mi>
</mml:math>
</inline-formula> are introduced to solve the optimal solution of the constrained problem and transform the constrained optimization problem into an unconstrained optimization problem.</p>
</list-item>
</list>
<disp-formula id="e4">
<mml:math id="m11">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>k</mml:mi>
</mml:munderover>
<mml:mo>&#x2016;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mi>j</mml:mi>
<mml:mrow>
<mml:mi>&#x3c0;</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mtext>&#x2a;</mml:mtext>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>j</mml:mi>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msup>
<mml:msubsup>
<mml:mo>&#x2016;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mo>&#x2016;</mml:mo>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>k</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:msubsup>
<mml:mo>&#x2016;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mo>&#x2329;</mml:mo>
<mml:mi>&#x3bb;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>k</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x232a;</mml:mo>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>
<list list-type="simple">
<list-item>
<p>5) Use the multiplicative operator alternating direction method to update <inline-formula id="inf8">
<mml:math id="m12">
<mml:mrow>
<mml:msubsup>
<mml:mi>u</mml:mi>
<mml:mi>k</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf9">
<mml:math id="m13">
<mml:mrow>
<mml:msubsup>
<mml:mi>w</mml:mi>
<mml:mi>k</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf10">
<mml:math id="m14">
<mml:mrow>
<mml:msup>
<mml:mi>&#x3bb;</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> alternately in both directions until the following iteration conditions are satisfied:</p>
</list-item>
</list>
<disp-formula id="e5">
<mml:math id="m15">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>k</mml:mi>
</mml:munder>
<mml:mo>&#x2016;</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>u</mml:mi>
<mml:mo>&#x2227;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>k</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>u</mml:mi>
<mml:mo>&#x2227;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mo>&#x2016;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
<mml:mo>/</mml:mo>
<mml:mo>&#x2016;</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>u</mml:mi>
<mml:mo>&#x2227;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mo>&#x2016;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>&#x3b5;</mml:mi>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>Where <inline-formula id="inf11">
<mml:math id="m16">
<mml:mrow>
<mml:mi>&#x3b5;</mml:mi>
<mml:mo>&#x3e;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf12">
<mml:math id="m17">
<mml:mrow>
<mml:msubsup>
<mml:mi>u</mml:mi>
<mml:mi>k</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf13">
<mml:math id="m18">
<mml:mrow>
<mml:msubsup>
<mml:mi>w</mml:mi>
<mml:mi>k</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf14">
<mml:math id="m19">
<mml:mrow>
<mml:msup>
<mml:mi>&#x3bb;</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> are denoted as:<disp-formula id="e6">
<mml:math id="m20">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>u</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>k</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>w</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>f</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>&#x3c9;</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>u</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>k</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mstyle>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>&#x3c9;</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3e;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>u</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>k</mml:mi>
<mml:mi>n</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>&#x3c9;</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>&#x3bb;</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:msup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>w</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:mfrac>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mi>&#x3b1;</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>&#x3c9;</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mi>&#x3c9;</mml:mi>
<mml:mi>k</mml:mi>
<mml:mi>n</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>
<disp-formula id="e7">
<mml:math id="m21">
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3c9;</mml:mi>
<mml:mi>k</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mstyle displaystyle="true">
<mml:mo>&#x222b;</mml:mo>
</mml:mstyle>
<mml:mn>0</mml:mn>
<mml:mi>&#x221e;</mml:mi>
</mml:msubsup>
<mml:mi>&#x3c9;</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>u</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>k</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>w</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mi>d</mml:mi>
<mml:mi>&#x3c9;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mstyle displaystyle="true">
<mml:mo>&#x222b;</mml:mo>
</mml:mstyle>
<mml:mn>0</mml:mn>
<mml:mi>&#x221e;</mml:mi>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>u</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>k</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>w</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mi>d</mml:mi>
<mml:mi>&#x3c9;</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>
<disp-formula id="e8">
<mml:math id="m22">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>&#x3bb;</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>&#x3c9;</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>&#x3bb;</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:msup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>&#x3c9;</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3c4;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>f</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>w</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>u</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>k</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>&#x3c9;</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(8)</label>
</disp-formula>where <inline-formula id="inf15">
<mml:math id="m23">
<mml:mi>&#x3c4;</mml:mi>
</mml:math>
</inline-formula> is the updated noise parameter.</p>
</sec>
<sec id="s2-1-2">
<title>2.1.2 VMD parameter determination</title>
<p>
<list list-type="simple">
<list-item>
<p>1) Modal Number</p>
</list-item>
</list>
</p>
<p>The number of modalities <italic>K</italic> should be determined before the VMD is used to decompose. Too large or too small a value of <italic>K</italic> will affect the accuracy of the model. In this paper, the mean value of instantaneous frequency of each component is used to determine the number of <italic>K</italic>. When the value of <italic>k</italic> is too large, and the high-frequency component will be broken. It means that the instantaneous frequency at the break of the high-frequency component is 0. As a result, the high-frequency component breaks lead to a decrease in the average instantaneous frequency. <xref ref-type="fig" rid="F2">Figure&#x20;2</xref> shows the mean values of instantaneous frequencies for the nine cases of VMD components. It can be seen from the figure that the number of VMD components increases to a certain number, and the curve has an obvious bending phenomenon. To sum up, the value of <italic>K</italic> is chosen as 4.<list list-type="simple">
<list-item>
<p>2) Penalty Factor</p>
</list-item>
</list>
</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Average instantaneous frequency when <italic>K</italic> from 1 to 9</p>
</caption>
<graphic xlink:href="fenrg-09-772508-g002.tif"/>
</fig>
<p>The penalty factor changes the constrained variational problem into a non-constrained variational problem. According to Ref. (<xref ref-type="bibr" rid="B32">Wu, 2016</xref>), when the value of the penalty factor is set to 2000 has strong adaptability and can ensure a certain convergence&#x20;speed.</p>
</sec>
</sec>
<sec id="s2-2">
<title>2.2 Principle of Max-Relevance and Min-Redundancy</title>
<p>Peng et&#x20;al.(<xref ref-type="bibr" rid="B26">Peng et&#x20;al., 2005</xref>) proposed a feature selection method based on Mutual Information, which uses Mutual Information to measure the dependency between two variables while taking into account the redundancy between features.</p>
<sec id="s2-2-1">
<title>2.2.1&#x20;Max-Relevance</title>
<p>The maximum correlation criterion solution can be expressed as the average of the mutual information between the feature <inline-formula id="inf16">
<mml:math id="m24">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and the target variable <inline-formula id="inf17">
<mml:math id="m25">
<mml:mi>y</mml:mi>
</mml:math>
</inline-formula>:<disp-formula id="e9">
<mml:math id="m26">
<mml:mrow>
<mml:mtext>max</mml:mtext>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>J</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mi>J</mml:mi>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
<mml:munder>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mi>J</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mi>I</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(9)</label>
</disp-formula>where <inline-formula id="inf18">
<mml:math id="m27">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the characteristic; <inline-formula id="inf19">
<mml:math id="m28">
<mml:mi>y</mml:mi>
</mml:math>
</inline-formula> represents the target variable; <inline-formula id="inf20">
<mml:math id="m29">
<mml:mi>J</mml:mi>
</mml:math>
</inline-formula> is the set containing the <inline-formula id="inf21">
<mml:math id="m30">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>; <inline-formula id="inf22">
<mml:math id="m31">
<mml:mrow>
<mml:mi>I</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represents the mutual information between the feature <inline-formula id="inf23">
<mml:math id="m32">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and the target variable <inline-formula id="inf24">
<mml:math id="m33">
<mml:mi>y</mml:mi>
</mml:math>
</inline-formula>. The expression is as follows:<disp-formula id="e10">
<mml:math id="m34">
<mml:mrow>
<mml:mi>I</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:mrow>
<mml:mo>&#x222b;</mml:mo>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mrow>
<mml:mo>&#x222b;</mml:mo>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
<mml:mtext>log</mml:mtext>
<mml:mfrac>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
<mml:mi>d</mml:mi>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mi>d</mml:mi>
<mml:mi>y</mml:mi>
</mml:mrow>
</mml:math>
<label>(10)</label>
</disp-formula>where <inline-formula id="inf25">
<mml:math id="m35">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> are the edge probability density functions of <inline-formula id="inf26">
<mml:math id="m36">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf27">
<mml:math id="m37">
<mml:mi>y</mml:mi>
</mml:math>
</inline-formula>, respectively <inline-formula id="inf28">
<mml:math id="m38">
<mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the joint probability density function of of <inline-formula id="inf29">
<mml:math id="m39">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and&#x20;<inline-formula id="inf30">
<mml:math id="m40">
<mml:mi>y</mml:mi>
</mml:math>
</inline-formula>.</p>
</sec>
<sec id="s2-2-2">
<title>2.2.2 Minimum Redundancy</title>
<p>The overlapping information between any two feature variables is called redundancy information. The features selected according to <xref ref-type="disp-formula" rid="e9">Eq. 9</xref> only consider the degree of correlation and do not consider the existence of redundancy between features. The input of redundant features increases the number of input features, and decreases the accuracy of the prediction model. The Minimum Redundancy expression is as follows:<disp-formula id="e11">
<mml:math id="m41">
<mml:mrow>
<mml:mi>min</mml:mi>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>J</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mi>J</mml:mi>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:mfrac>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mi>J</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mi>J</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mrow>
<mml:mi>I</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
<label>(11)</label>
</disp-formula>mRMR can be expressed by <xref ref-type="disp-formula" rid="e9">Eqs 9</xref>, <xref ref-type="disp-formula" rid="e11">11</xref> as:<disp-formula id="e12">
<mml:math id="m42">
<mml:mrow>
<mml:mi>max</mml:mi>
<mml:mi>&#x3c8;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>&#x3c8;</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>R</mml:mi>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(12)</label>
</disp-formula>
</p>
</sec>
</sec>
<sec id="s2-3">
<title>2.3 Prediction Model</title>
<p>The prediction part uses Long Short-Term Memory Neural Network (LSTM) model, which was proposed by Hochreiter and Schmidhuber (<xref ref-type="bibr" rid="B9">Hochreiter and Schmidhuber, 1997</xref>) to learn long-term dependence information. It can handle more complex problems, and has more mature applications in the field of load prediction (<xref ref-type="bibr" rid="B28">Sun et&#x20;al., 2019a</xref>). Long short-term memory neural network is a special form of the recurrent neural network. LSTM is composed of cells with the same structure (<xref ref-type="fig" rid="F3">Figure&#x20;3</xref>). In this model, the data of the next moment is predicted each time by the previous data and historical data, which is processed by the cells. Each cell has three input parameters: Historically stored information <inline-formula id="inf31">
<mml:math id="m43">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, historical data <inline-formula id="inf32">
<mml:math id="m44">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf33">
<mml:math id="m45">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, which represent the prediction results of the last cell and input parameter to the cell. Each cell contains four parts, including the forgotten gate, the input gate, the update gate and the output&#x20;gate.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>LSTM structure.</p>
</caption>
<graphic xlink:href="fenrg-09-772508-g003.tif"/>
</fig>
<p>The data <inline-formula id="inf34">
<mml:math id="m46">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> that processed by the previous cell, and the input data of the current time <inline-formula id="inf35">
<mml:math id="m47">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are linked by a matrix and obtain <inline-formula id="inf36">
<mml:math id="m48">
<mml:mrow>
<mml:msubsup>
<mml:mi>X</mml:mi>
<mml:mi>t</mml:mi>
<mml:mtext>&#x27;</mml:mtext>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>,<disp-formula id="e13">
<mml:math id="m49">
<mml:mrow>
<mml:msub>
<mml:msup>
<mml:mi>X</mml:mi>
<mml:mtext>&#x27;</mml:mtext>
</mml:msup>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(13)</label>
</disp-formula>
</p>
<p>In the forgotten gate, LSTM can decide what information to discard from the cell. After the sigmoid function processing, <inline-formula id="inf37">
<mml:math id="m50">
<mml:mrow>
<mml:msubsup>
<mml:mi>X</mml:mi>
<mml:mi>t</mml:mi>
<mml:mtext>&#x27;</mml:mtext>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> can get <inline-formula id="inf38">
<mml:math id="m51">
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. LSTM can remember large amounts of historical data by <inline-formula id="inf39">
<mml:math id="m52">
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> filtered data.<disp-formula id="e14">
<mml:math id="m53">
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>f</mml:mi>
</mml:msub>
<mml:mo>&#x22C5;</mml:mo>
<mml:msup>
<mml:mi>X</mml:mi>
<mml:mtext>&#x27;</mml:mtext>
</mml:msup>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>f</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(14)</label>
</disp-formula>
</p>
<p>In the input gate, LSTM acquires the new data, After the sigmoid function processing, <inline-formula id="inf40">
<mml:math id="m54">
<mml:mrow>
<mml:msubsup>
<mml:mi>X</mml:mi>
<mml:mi>t</mml:mi>
<mml:mtext>&#x27;</mml:mtext>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> can get <inline-formula id="inf41">
<mml:math id="m55">
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. The <inline-formula id="inf42">
<mml:math id="m56">
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> decides the useful data in <inline-formula id="inf43">
<mml:math id="m57">
<mml:mrow>
<mml:msubsup>
<mml:mi>X</mml:mi>
<mml:mi>t</mml:mi>
<mml:mtext>&#x27;</mml:mtext>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>. Moreover, <inline-formula id="inf44">
<mml:math id="m58">
<mml:mrow>
<mml:msubsup>
<mml:mi>X</mml:mi>
<mml:mi>t</mml:mi>
<mml:mtext>&#x27;</mml:mtext>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> is processed by the tanh function to calculate <inline-formula id="inf45">
<mml:math id="m59">
<mml:mrow>
<mml:msubsup>
<mml:mi>C</mml:mi>
<mml:mi>t</mml:mi>
<mml:mo>&#x27;</mml:mo>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>,<disp-formula id="e15">
<mml:math id="m60">
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x22C5;</mml:mo>
<mml:msup>
<mml:mi>X</mml:mi>
<mml:mtext>&#x27;</mml:mtext>
</mml:msup>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(15)</label>
</disp-formula>
<disp-formula id="e16">
<mml:math id="m61">
<mml:mrow>
<mml:msub>
<mml:msup>
<mml:mi>C</mml:mi>
<mml:mtext>&#x27;</mml:mtext>
</mml:msup>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>tan</mml:mi>
<mml:mtext>h</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
<mml:mo>&#x22C5;</mml:mo>
<mml:msup>
<mml:mi>X</mml:mi>
<mml:mtext>&#x27;</mml:mtext>
</mml:msup>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(16)</label>
</disp-formula>
<disp-formula id="e17">
<mml:math id="m62">
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:msup>
<mml:mi>C</mml:mi>
<mml:mtext>&#x27;</mml:mtext>
</mml:msup>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
<label>(17)</label>
</disp-formula>
<inline-formula id="inf46">
<mml:math id="m63">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is updated in the update gate. To obtain the historical data, <inline-formula id="inf47">
<mml:math id="m64">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf48">
<mml:math id="m65">
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are multiplied by the matrix, In order to keep more accurate rules in the cell for accurate prediction, <inline-formula id="inf49">
<mml:math id="m66">
<mml:mrow>
<mml:mtext>&#xa0;</mml:mtext>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is added to the equation to get output <inline-formula id="inf50">
<mml:math id="m67">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>,<disp-formula id="e18">
<mml:math id="m68">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
<label>(18)</label>
</disp-formula>
</p>
<p>LSTM outputs the result in the output gate. After the sigmoid function processing, <inline-formula id="inf51">
<mml:math id="m69">
<mml:mrow>
<mml:msubsup>
<mml:mi>X</mml:mi>
<mml:mi>t</mml:mi>
<mml:mtext>&#x27;</mml:mtext>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> can get <inline-formula id="inf52">
<mml:math id="m70">
<mml:mrow>
<mml:msub>
<mml:mi>O</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. <inline-formula id="inf53">
<mml:math id="m71">
<mml:mrow>
<mml:msub>
<mml:mi>O</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> decides which <inline-formula id="inf54">
<mml:math id="m72">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> needs to be retained as the result. In addition, <inline-formula id="inf55">
<mml:math id="m73">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is processed by the tanh function to get <inline-formula id="inf56">
<mml:math id="m74">
<mml:mrow>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mo>&#x27;</mml:mo>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf57">
<mml:math id="m75">
<mml:mrow>
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
<mml:mo>&#x27;</mml:mo>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf58">
<mml:math id="m76">
<mml:mrow>
<mml:msub>
<mml:mi>O</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are multiplied to obtain the final data <inline-formula id="inf59">
<mml:math id="m77">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>,<disp-formula id="e19">
<mml:math id="m78">
<mml:mrow>
<mml:msub>
<mml:mi>O</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>o</mml:mi>
</mml:msub>
<mml:mo>&#x22C5;</mml:mo>
<mml:msub>
<mml:msup>
<mml:mi>X</mml:mi>
<mml:mtext>&#x27;</mml:mtext>
</mml:msup>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>o</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(19)</label>
</disp-formula>
<disp-formula id="e20">
<mml:math id="m79">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>O</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(20)</label>
</disp-formula>
<disp-formula id="e21">
<mml:math id="m80">
<mml:mrow>
<mml:mi>tan</mml:mi>
<mml:mtext>h</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(21)</label>
</disp-formula>
<disp-formula id="e22">
<mml:math id="m81">
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mi>x</mml:mi>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mi>x</mml:mi>
</mml:msup>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(22)</label>
</disp-formula>
</p>
<p>In order to compare the prediction results of different models, three evaluation metrics will be used in this paper: Decision factor: R-square (<italic>R</italic>
<sup>2</sup>), mean absolute error (MAE), and root mean square error (RMSE). The specific calculation of these metrics is described as follows:<disp-formula id="e23">
<mml:math id="m82">
<mml:mrow>
<mml:msub>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>Q</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>Q</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>Q</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
<label>(23)</label>
</disp-formula>
<disp-formula id="e24">
<mml:math id="m83">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>E</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>n</mml:mi>
</mml:mfrac>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>Q</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
<label>(24)</label>
</disp-formula>
<disp-formula id="e25">
<mml:math id="m84">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>M</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>E</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msqrt>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>n</mml:mi>
</mml:mfrac>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>Q</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:math>
<label>(25)</label>
</disp-formula>Where <inline-formula id="inf60">
<mml:math id="m85">
<mml:mrow>
<mml:msub>
<mml:mi>Q</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the recorded value of building power consumption at time <italic>i</italic>, and <inline-formula id="inf61">
<mml:math id="m86">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the predicted value of building power consumption at time <italic>i</italic>. These three criteria describe the close-ness of the predicted data to the actual data in three different ways, The value of <italic>R</italic>
<sup>2</sup> is between 0 and 1, with 0 indicating wo<sub>r</sub>se than the mean and 1 indicating perfect prediction, And for MAE and RMSE, the smaller the value, the better the prediction result of the&#x20;model.</p>
</sec>
</sec>
<sec id="s3">
<title>3 Case Study</title>
<sec id="s3-1">
<title>3.1 Data Introduction</title>
<p>The building electricity consumption data obtained in this article was obtained from the Qingdao civil building energy consumption monitoring platform. Raw data was selected from three summer cooling months (June, July and August) with a time granularity of 1&#xa0;hour. The maximum and minimum values of the original data are 803.5&#xa0;KW and 65&#xa0;KW; the difference between the maximum and minimum values is 738.5&#xa0;KW, which demonstrates the volatility of the data. The mean and standard deviation of this data are 333.48&#xa0;KW and 199.97&#xa0;KW, respectively, which shows the large dispersion of the data. In <xref ref-type="fig" rid="F4">Figure&#x20;4</xref>, which illustrates the sequence of the original data, it can be seen that the raw load fluctuates considerably.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Original total energy series.</p>
</caption>
<graphic xlink:href="fenrg-09-772508-g004.tif"/>
</fig>
</sec>
<sec id="s3-2">
<title>3.2 Comparison of Decomposition by EEMD and VMD</title>
<p>
<xref ref-type="fig" rid="F5">Figure&#x20;5A</xref> shows how EEMD decomposes the original load into 11 intrinsic mode functions (IMFs) and a residual, and <xref ref-type="fig" rid="F5">Figure&#x20;5B</xref> shows the spectrum after passing the Fourier transform. Even with the improved EMD algorithm, the phenomenon of modal mixing is still evident. Modal mixing occurs when one modal component is decomposed into multiple components. In the figure, the frequency band of IMF4 overlaps with the frequency bands of IMF3 and IMF5. Modal mixing is a defect of the EEMD algorithm and leads to degradation of the model accuracy, so it is important to avoid this phenomenon.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>
<bold>(A)</bold> The decomposition results of EEMD. <bold>(B)</bold> The decomposition spectrogram of EEMD.</p>
</caption>
<graphic xlink:href="fenrg-09-772508-g005.tif"/>
</fig>
<p>The VMD algorithm solves the modal mixing problem inherent in the EEMD algorithm. In <xref ref-type="fig" rid="F6">Figure&#x20;6A</xref>, the VMD decomposition results are shown for <italic>K</italic>&#x20;&#x3d; 4 and the penalty parameter <italic>a</italic>&#x20;&#x3d; 2000, which were determined in <xref ref-type="sec" rid="s2-1-2">Section 2.1.2</xref>. u<sub>1</sub> is the lowest frequency component, <italic>u</italic>
<sub>
<italic>2</italic>
</sub> is the medium frequency component and <italic>u</italic>
<sub>
<italic>3</italic>
</sub> and <italic>u</italic>
<sub>
<italic>4</italic>
</sub> are the highest frequency components. According to the additional analysis supplied by the spectrogram in <xref ref-type="fig" rid="F6">Figure&#x20;6B</xref>, there is no overlap in the frequencies of the components, which indicates that the VMD algorithm solves the problem of modal mixing.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>
<bold>(A)</bold> The decomposition results of VMD. <bold>(B)</bold> The decomposition spectrogram of VMD.</p>
</caption>
<graphic xlink:href="fenrg-09-772508-g006.tif"/>
</fig>
<p>In summary, both the EEMD and VMD algorithms are capable of handling volatile raw data, and both algorithms decompose the load into several stable components. The EEMD algorithm is an improved algorithm based on EMD, but it is limited due to the phenomenon of modal confusion. The VMD algorithm overcomes this shortcoming. The VMD algorithm can sufficiently decompose the raw load data to obtain more physically meaningful components and improve the accuracy of the model prediction.</p>
</sec>
<sec id="s3-3">
<title>3.3 Feature Selection</title>
<p>Building electricity consumption is influenced by climate and historical load. However, the raw data contains only meteorological factors. To fully consider the independence of each component and research the physical significance of each component, a set of feature matrices are established in this paper. The appropriate feature set is selected by the mRMR algorithm for input into the prediction model. The established feature matrices and their representations are shown in <xref ref-type="table" rid="T1">Table&#x20;1</xref>.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Construction of feature matrix.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Feature name</th>
<th align="center">Representation</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">D_t</td>
<td align="left">D_1,D_2,D_3&#x2026;D_48</td>
</tr>
<tr>
<td align="left">T_t</td>
<td align="left">T_1,T_2,T_3&#x2026;T_48</td>
</tr>
<tr>
<td align="left">H_t</td>
<td align="left">H_1,H_2,H_3&#x2026;H_48</td>
</tr>
<tr>
<td align="left">Dp_t</td>
<td align="left">Dp_1,Dp_2,Dp_3&#x2026;Dp_48</td>
</tr>
<tr>
<td align="left">Wind speed</td>
<td align="left">Wind speed</td>
</tr>
<tr>
<td align="left">Temperature</td>
<td align="left">Temperature</td>
</tr>
<tr>
<td align="left">Humidity</td>
<td align="left">Humidity</td>
</tr>
<tr>
<td align="left">Dew point</td>
<td align="left">Dew point</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The time interval of the load data collected in this paper is 1&#xa0;h. In <xref ref-type="table" rid="T1">Table&#x20;1</xref>, D_t represents the load point for the previous 48&#xa0;h at time t. Similarly, T_t, H_t and Dp_t represent the temperature, humidity and dew point temperature, respectively, for the previous 48&#xa0;h at time t. The wind speed, temperature, humidity and dew point are the meteorological characteristics of the dataset.</p>
<p>After establishing the feature matrix, each component of the decomposition is used as the target variable <inline-formula id="inf62">
<mml:math id="m87">
<mml:mi>y</mml:mi>
</mml:math>
</inline-formula>, and <inline-formula id="inf63">
<mml:math id="m88">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the data point in the feature matrix. The mRMR values are calculated according to <xref ref-type="disp-formula" rid="e12">Eq. 12</xref> and the calculation results are sorted in descending order. The top 15 influencing factors are selected as the feature matrix of the input model. The final selection results are shown in <xref ref-type="table" rid="T2">Table&#x20;2</xref>.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Feature selection results.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Number</th>
<th colspan="2" align="center">u<sub>1</sub>
</th>
<th colspan="2" align="center">u<sub>2</sub>
</th>
<th colspan="2" align="center">u<sub>3</sub>
</th>
<th colspan="2" align="center">u<sub>4</sub>
</th>
</tr>
<tr>
<th align="center">mRMR</th>
<th align="center">MIM</th>
<th align="center">mRMR</th>
<th align="center">MIM</th>
<th align="center">mRMR</th>
<th align="center">MIM</th>
<th align="center">mRMR</th>
<th align="center">MIM</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">1</td>
<td align="center">D_8</td>
<td align="center">D_8</td>
<td align="center">D_1</td>
<td align="center">D-1</td>
<td align="center">D_2</td>
<td align="center">D-2</td>
<td align="center">D_1</td>
<td align="center">D-1</td>
</tr>
<tr>
<td align="left">2</td>
<td align="center">D_46</td>
<td align="center">D-15</td>
<td align="center">D_13</td>
<td align="center">D-2</td>
<td align="center">D_15</td>
<td align="center">D-3</td>
<td align="center">Humidity</td>
<td align="center">D-2</td>
</tr>
<tr>
<td align="left">3</td>
<td align="center">D_15</td>
<td align="center">D-16</td>
<td align="center">D_43</td>
<td align="center">D-6</td>
<td align="center">Humidity</td>
<td align="center">D-4</td>
<td align="center">D_12</td>
<td align="center">D-3</td>
</tr>
<tr>
<td align="left">4</td>
<td align="center">D_1</td>
<td align="center">D-5</td>
<td align="center">D_7</td>
<td align="center">D-13</td>
<td align="center">D_8</td>
<td align="center">D-1</td>
<td align="center">D_21</td>
<td align="center">D-4</td>
</tr>
<tr>
<td align="left">5</td>
<td align="center">D_36</td>
<td align="center">D-4</td>
<td align="center">D_27</td>
<td align="center">D-5</td>
<td align="center">D_23</td>
<td align="center">D-7</td>
<td align="center">1D_8</td>
<td align="center">D-5</td>
</tr>
<tr>
<td align="left">6</td>
<td align="center">D_26</td>
<td align="center">D-46</td>
<td align="center">D_36</td>
<td align="center">D-10</td>
<td align="center">D_45</td>
<td align="center">D-6</td>
<td align="center">D_2</td>
<td align="center">D-6</td>
</tr>
<tr>
<td align="left">7</td>
<td align="center">D_17</td>
<td align="center">D-30</td>
<td align="center">D_47</td>
<td align="center">D-7</td>
<td align="center">D_4</td>
<td align="center">D-5</td>
<td align="center">D_41</td>
<td align="center">D-8</td>
</tr>
<tr>
<td align="left">8</td>
<td align="center">D_30</td>
<td align="center">D-26</td>
<td align="center">D_17</td>
<td align="center">D-3</td>
<td align="center">D_13</td>
<td align="center">D-10</td>
<td align="center">D_47</td>
<td align="center">D-23</td>
</tr>
<tr>
<td align="left">9</td>
<td align="center">D_44</td>
<td align="center">D-10</td>
<td align="center">D_2</td>
<td align="center">D-9</td>
<td align="center">D_33</td>
<td align="center">D-9</td>
<td align="center">D_14</td>
<td align="center">D-24</td>
</tr>
<tr>
<td align="left">10</td>
<td align="center">D_4</td>
<td align="center">D-47</td>
<td align="center">D_10</td>
<td align="center">D-4</td>
<td align="center">D_1</td>
<td align="center">D-8</td>
<td align="center">D_3</td>
<td align="center">D-12</td>
</tr>
<tr>
<td align="left">11</td>
<td align="center">D_10</td>
<td align="center">D-48</td>
<td align="center">D_16</td>
<td align="center">D-12</td>
<td align="center">D_17</td>
<td align="center">D-13</td>
<td align="center">Wind speed</td>
<td align="center">D-25</td>
</tr>
<tr>
<td align="left">12</td>
<td align="center">D_48</td>
<td align="center">D-14</td>
<td align="center">D_6</td>
<td align="center">D-8</td>
<td align="center">D_7</td>
<td align="center">D-23</td>
<td align="center">D_10</td>
<td align="center">D-7</td>
</tr>
<tr>
<td align="left">13</td>
<td align="center">D_39</td>
<td align="center">D-28</td>
<td align="center">D_48</td>
<td align="center">D-27</td>
<td align="center">D_46</td>
<td align="center">D-12</td>
<td align="center">D_23</td>
<td align="center">D-9</td>
</tr>
<tr>
<td align="left">14</td>
<td align="center">D_16</td>
<td align="center">D-22</td>
<td align="center">D_38</td>
<td align="center">D-11</td>
<td align="center">D_3</td>
<td align="center">D-14</td>
<td align="center">D_4</td>
<td align="center">D-10</td>
</tr>
<tr>
<td align="left">15</td>
<td align="center">D_3</td>
<td align="center">D-39</td>
<td align="center">D_3</td>
<td align="center">D-28</td>
<td align="center">D_14</td>
<td align="center">D-11</td>
<td align="center">D_35</td>
<td align="center">D-22</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To illustrate the superiority of the mRMR algorithm, mutual information maximization (MIM) (<xref ref-type="bibr" rid="B23">Novakovic et&#x20;al., 2011</xref>) is used as a comparison in this paper. The MIM algorithm is based on the theory of mutual information, but unlike mRMR, the MIM algorithm only considers the correlation between features and target variables and does not consider the redundancy between features. The results of the MIM feature selection are shown in <xref ref-type="table" rid="T2">Table&#x20;2</xref>.</p>
<p>Consider <italic>u</italic>
<sub>
<italic>1</italic>
</sub> and <italic>u</italic>
<sub>
<italic>4</italic>
</sub> in <xref ref-type="table" rid="T2">Table&#x20;2</xref> as an example. The low-frequency components <italic>u</italic>
<sub>
<italic>1</italic>
</sub> and <italic>u</italic>
<sub>
<italic>2</italic>
</sub> are mainly influenced by D_t, which indicates that the <italic>u</italic>
<sub>
<italic>1</italic>
</sub> and <italic>u</italic>
<sub>
<italic>2</italic>
</sub> components are influenced more heavily by the historical load of the past 48&#xa0;h. It is further seen through <xref ref-type="fig" rid="F6">Figure&#x20;6A</xref> that although both <italic>u</italic>
<sub>
<italic>1</italic>
</sub> and <italic>u</italic>
<sub>
<italic>2</italic>
</sub> are strongly influenced by historical loads, <italic>u</italic>
<sub>
<italic>1</italic>
</sub> presents a load variation trend with a week as a period, while <italic>u</italic>
<sub>
<italic>2</italic>
</sub> presents a load variation trend with a 24-h period. In contrast, the high-frequency components <italic>u</italic>
<sub>
<italic>3</italic>
</sub> and <italic>u</italic>
<sub>
<italic>4</italic>
</sub> are not only influenced by the historical load but also by the weather factor. Weather factors are usually seen as uncertainty factors. The influence of humidity on <italic>u</italic>
<sub>
<italic>4</italic>
</sub> is ranked second among all the features. This explains why <italic>u</italic>
<sub>
<italic>4</italic>
</sub> is more volatile than <italic>u</italic>
<sub>
<italic>1</italic>
</sub>: <italic>u</italic>
<sub>
<italic>1</italic>
</sub> is mainly influenced by historical load and has a certain regularity, while <italic>u</italic>
<sub>
<italic>4</italic>
</sub> is influenced by uncertainties such as humidity. Thus, <italic>u</italic>
<sub>
<italic>4</italic>
</sub> is more irregular.</p>
<p>
<xref ref-type="table" rid="T3">Table&#x20;3</xref> shows that the results of the MIM feature selection method are similarly ranked, with the higher-ranked features all being historical loads at a given moment. This is especially apparent for the <italic>u</italic>
<sub>
<italic>3</italic>
</sub> and <italic>u</italic>
<sub>
<italic>4</italic>
</sub> components. The top five features selected using MIM have a high degree of overlap because the MIM algorithm only considers the maximum correlation between features and variables while ignoring the degree of redundancy between features. This is improved by using the mRMR algorithm. For <italic>u</italic>
<sub>
<italic>3</italic>
</sub> and <italic>u</italic>
<sub>
<italic>4</italic>
</sub>, the feature overlap selected using the mRMR algorithm is not high, and features that are not considered by MIM, such as wind speed and humidity, are taken into account by the mRMR algorithm.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>VMD component prediction result.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Model</th>
<th align="center">Subsequences</th>
<th align="center">MAE (kWh)</th>
<th align="center">RMSE (kWh)</th>
<th align="center">R<sub>2</sub>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="4" align="left">VMD &#x2b; mRMR &#x2b; LSTM</td>
<td align="center">u<sub>1</sub>
</td>
<td align="char" char=".">2.61</td>
<td align="char" char=".">3.34</td>
<td align="char" char=".">0.997</td>
</tr>
<tr>
<td align="center">u<sub>2</sub>
</td>
<td align="char" char=".">8.61</td>
<td align="char" char=".">11.23</td>
<td align="char" char=".">0.994</td>
</tr>
<tr>
<td align="center">u<sub>3</sub>
</td>
<td align="char" char=".">2.90</td>
<td align="char" char=".">4.02</td>
<td align="char" char=".">0.992</td>
</tr>
<tr>
<td align="center">u<sub>4</sub>
</td>
<td align="char" char=".">2.18</td>
<td align="char" char=".">3.07</td>
<td align="char" char=".">0.982</td>
</tr>
<tr>
<td rowspan="4" align="left">VMD &#x2b; MIM &#x2b; LSTM</td>
<td align="center">u<sub>1</sub>
</td>
<td align="char" char=".">8.83</td>
<td align="char" char=".">10.44</td>
<td align="char" char=".">0.977</td>
</tr>
<tr>
<td align="center">u<sub>2</sub>
</td>
<td align="char" char=".">25.63</td>
<td align="char" char=".">35.95</td>
<td align="char" char=".">0.943</td>
</tr>
<tr>
<td align="center">u<sub>3</sub>
</td>
<td align="char" char=".">7.08</td>
<td align="char" char=".">11.14</td>
<td align="char" char=".">0.960</td>
</tr>
<tr>
<td align="center">u<sub>4</sub>
</td>
<td align="char" char=".">2.11</td>
<td align="char" char=".">3.09</td>
<td align="char" char=".">0.980</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In summary, the mRMR algorithm considers not only the correlation between features and target variables but also the degree of redundancy between features. The selected features can better reflect some characteristics of the modal components and reduce the dimensionality of the feature matrix.</p>
</sec>
<sec id="s3-4">
<title>3.3 Model Predictions</title>
<p>In this paper, the LSTM model is used for predictions. The training and test sets are divided for a total of 2,208 data points from June 1, 2017 to August 31, 2017. Of this data, 80% is used to build the model and 20% is used to check the validity of the established models.</p>
<p>The number of layers of the LSTM model serves to remember important information, and theoretically, more hidden layers give the model an improved nonlinear fitting ability and a better learning effect. However, increasing the number of layers consumes a considerable amount of computation time. According to the literature (<xref ref-type="bibr" rid="B24">Pan, 2018</xref>; <xref ref-type="bibr" rid="B13">Li et&#x20;al., 2019</xref>), the number of implied layers generally does not exceed 3, so the number of implied layers in this paper has been determined to be&#x20;1.</p>
<p>The number of nodes in the hidden layer affects the performance of the model. If the number of nodes in the implicit layer is too small, less effective information is obtained in the prediction process. If the number of nodes in the implicit layer is too large, it may lead to a longer training time and overfitting problems. According to the literature (<xref ref-type="bibr" rid="B34">Xu et&#x20;al., 2020</xref>), the number of nodes in the hidden layer can be determined by <xref ref-type="disp-formula" rid="e26">Eq. (26)</xref>:<disp-formula id="e26">
<mml:math id="m89">
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msqrt>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msqrt>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:math>
<label>(26)</label>
</disp-formula>where <inline-formula id="inf64">
<mml:math id="m90">
<mml:mi>l</mml:mi>
</mml:math>
</inline-formula> is the number of nodes in the hidden layer, <inline-formula id="inf65">
<mml:math id="m91">
<mml:mi>m</mml:mi>
</mml:math>
</inline-formula> is the number of input nodes, <inline-formula id="inf66">
<mml:math id="m92">
<mml:mi>n</mml:mi>
</mml:math>
</inline-formula> is the number of output nodes, and <inline-formula id="inf67">
<mml:math id="m93">
<mml:mi>a</mml:mi>
</mml:math>
</inline-formula> is a constant from 1 to 10. By calculation, the number of nodes <inline-formula id="inf68">
<mml:math id="m94">
<mml:mi>l</mml:mi>
</mml:math>
</inline-formula> in the hidden layer is determined to be 8. The LSTM network is trained using the Adam optimization algorithm (<xref ref-type="bibr" rid="B30">Wang et&#x20;al., 2019</xref>). By referring to relevant literature (<xref ref-type="bibr" rid="B12">Kong et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B25">Pei et&#x20;al., 2020</xref>) and experimental measurements, the remaining parameters are set: The number of iterations of the neural network is set to 1,000, the learning rate is set to 0.01, and the expected error is set to 0.0004. The prediction results are shown in <xref ref-type="fig" rid="F7">Figure&#x20;7</xref> and <xref ref-type="table" rid="T4">Table&#x20;4</xref>.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>VMD component prediction figure.<bold>(A)</bold> u1, <bold>(B)</bold> u2, <bold>(C)</bold> u3, <bold>(D)</bold>&#x20;u4.</p>
</caption>
<graphic xlink:href="fenrg-09-772508-g007.tif"/>
</fig>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Evaluation metrics of each&#x20;model.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Model</th>
<th align="center">MAE (kWh)</th>
<th align="center">RMSE (kWh)</th>
<th align="center">R<sub>2</sub>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">LSTM</td>
<td align="char" char=".">41.77</td>
<td align="char" char=".">67.70</td>
<td align="char" char=".">0.87</td>
</tr>
<tr>
<td align="left">VMD &#x2b; mRMR &#x2b; BPNN</td>
<td align="char" char=".">36.42</td>
<td align="char" char=".">47.54</td>
<td align="char" char=".">0.94</td>
</tr>
<tr>
<td align="left">EEMD &#x2b; mRMR &#x2b; BPNN</td>
<td align="char" char=".">50.39</td>
<td align="char" char=".">62.21</td>
<td align="char" char=".">0.89</td>
</tr>
<tr>
<td align="left">VMD &#x2b; MIM &#x2b; LSTM</td>
<td align="char" char=".">33.15</td>
<td align="char" char=".">46.42</td>
<td align="char" char=".">0.94</td>
</tr>
<tr>
<td align="left">VMD &#x2b; mRMR &#x2b; LSTM</td>
<td align="char" char=".">21.36</td>
<td align="char" char=".">30.64</td>
<td align="char" char=".">0.97</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>
<xref ref-type="fig" rid="F7">Figure&#x20;7</xref> and <xref ref-type="table" rid="T3">Table&#x20;3</xref> demonstrate that the integrated model proposed in this paper achieves better results for the prediction of each component. In general, the prediction results for the low and medium frequency components (<italic>u</italic>
<sub>
<italic>1</italic>
</sub> and <italic>u</italic>
<sub>
<italic>2</italic>
</sub>, respectively) are better, with <italic>R</italic>
<sup>2</sup> values of 0.997 and 0.994, respectively. The u<sub>3</sub> component also achieved a better prediction, having an <italic>R</italic>
<sup>2</sup> value of 0.992. In contrast, the prediction results for the high-frequency component u<sub>4</sub> are slightly worse, with an <italic>R</italic>
<sup>2</sup> value of only&#x20;0.982.</p>
<p>
<xref ref-type="table" rid="T3">Table&#x20;3</xref> also shows the prediction results for each component obtained using the MIM feature selection method. The <italic>R</italic>
<sup>2</sup> values of the prediction results for all four components are lower than those obtained by the mRMR method, especially for the <italic>u</italic>
<sub>
<italic>2</italic>
</sub> and <italic>u</italic>
<sub>
<italic>3</italic>
</sub> components. The main reason for this result is because the MIM feature selection algorithm does not consider the redundancy among the features, which leads to a certain degree of repetitiveness of the selected features.</p>
</sec>
<sec id="s3-5">
<title>3.4 Model Comparison</title>
<p>The proposed model is compared and analysed alongside other models to verify its reliability. The other models are singular and include a model using EEMD decomposition (EEMD&#x2013;mRMR&#x2013;BPNN), a model using the MIM algorithm (VMD&#x2013;MIM&#x2013;LSTM) and a model using the BPNN algorithm (VMD&#x2013;mRMR&#x2013;BPNN). The prediction results and evaluation metrics of all models are shown in <xref ref-type="fig" rid="F8">Figure&#x20;8</xref> and <xref ref-type="table" rid="T4">Table&#x20;4</xref>.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>Prediction results of each&#x20;model.</p>
</caption>
<graphic xlink:href="fenrg-09-772508-g008.tif"/>
</fig>
<p>The predictions of the integrated model with the addition of the modal decomposition algorithm are more accurate compared to the single prediction model (LSTM), as shown in <xref ref-type="fig" rid="F8">Figure&#x20;8</xref>. This indicates that the modal decomposition algorithm can indeed handle more complex data and improve the accuracy of the model. In addition, the prediction model proposed in this paper has the highest prediction accuracy among the four models.</p>
<p>According to the evaluation metrics analysis in <xref ref-type="table" rid="T4">Table&#x20;4</xref>, the prediction error of the single LSTM model is larger than the prediction error of the integrated model. This is mainly due to the instability of the load data and the limitations of the input features. The modal decomposition algorithm can decompose the fluctuating data into several stable IMFs, and the mRMR algorithm can select suitable features for the model. Thus, the prediction results of the integrated model are better than those of the single LSTM model. In addition, the integrated model using the VMD modal decomposition method (VMD&#x2013;mRMR&#x2013;BPNN) predicts better results than the integrated model using EEMD (EEMD&#x2013;mRMR&#x2013;BPNN). The <italic>R</italic>
<sup>2</sup> is improved by 5.6% and the MAE and RMSE are reduced by 27.7 and 23.6%, respectively, because the VMD algorithm solves the problems of modal aliasing and elusive components.</p>
<p>Comparing the VMD&#x2013;MIM&#x2013;LSTM and VMD&#x2013;mRMR&#x2013;LSTM integrated models, the mRMR algorithm, which takes into account the redundancy between features, achieves better prediction accuracy for the feature selection algorithm. The <italic>R</italic>
<sup>2</sup> is improved by 3.0%, the value of the MAE is reduced by 35.6% and the value of RMSE is reduced by 34.1%. This is because the mRMR algorithm takes into account the redundancy between features and can select the appropriate feature matrix for each&#x20;IMF.</p>
<p>Comparing the VMD&#x2013;mRMR&#x2013;LSTM and VMD&#x2013;mRMR&#x2013;BPNN prediction models, the integrated model using LSTM outperforms the integrated model using BPNN. The <italic>R</italic>
<sup>2</sup> of the LSTM integrated model is improved by 3.0%, the value of MAE is reduced by 41.4% and the value of RMSE is reduced by 35.5%. The power load series is a sample of power load variation over time, and the BPNN model has shortcomings in analysing these types of time series. For the time series, the LSTM model better mines the relationship between the data points. In brief, the model proposed in this paper has the highest prediction accuracy.</p>
</sec>
<sec id="s3-6">
<title>3.5 Model Robustness</title>
<p>The experimental results show that the hybrid model proposed in this paper has high prediction accuracy. In this section, the robustness of the hybrid model is analysed by varying the number of input feature parameters and the number of neurons in the hidden layer. For simplicity, the VMD-decomposed u<sub>4</sub> has been selected as the target dataset.</p>
<sec id="s3-6-1">
<title>3.5.1 Number of Neurons in The Hidden Layers</title>
<p>In theory, with the increase of the number of neurons in the hidden layer and the more abstract features extracted by deep learning, the more accurate a time series will be, which is favourable for predictions (<xref ref-type="bibr" rid="B37">Zhang et&#x20;al., 2020</xref>). <xref ref-type="fig" rid="F9">Figure&#x20;9A</xref> shows the RMSE of the VMD&#x2013;mRMR&#x2013;LSTM and VMD&#x2013;MIM&#x2013;LSTM models when the number of neurons in the hidden layer is changed. When the number of hidden layer nodes is between 5 and 25, the RMSE shows a gradual decrease; when the number of hidden layer nodes is between 25 and 65, the RMSE has a large fluctuation; when the number of hidden layer nodes is between 65 and 100, the RMSE tends to be smooth, its value is mostly between 3 and 5 and the prediction error is relatively stable, which indicates that these two models are highly robust. In addition, the RMSE of VMD&#x2013;mRMR&#x2013;LSTM is lower than that of VMD&#x2013;MIM&#x2013;LSTM in most situations, which indicates that VMD&#x2013;mRMR&#x2013;LSTM has better prediction performance and more stable robustness.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>
<bold>(A)</bold> Effect of the number of hidden layers on RMSE of u4. <bold>(B)</bold> Effect of the number of input parameters on RMSE of&#x20;u4.</p>
</caption>
<graphic xlink:href="fenrg-09-772508-g009.tif"/>
</fig>
</sec>
<sec id="s3-6-2">
<title>3.5.2 Number of Input Parameters</title>
<p>The redundancy between features is theoretically taken into account by the mRMR algorithm so that more input parameters lead to a better prediction performance of the model. However, too many feature parameters can increase the complexity of the model and increase the computing cost. <xref ref-type="fig" rid="F9">Figure&#x20;9B</xref> illustrates the effect of the number of feature parameters on the accuracy of the model. When the number of input features is between 1 and 6, the RMSE of the models is decreasing and fluctuates. When the number of input features is between 6 and 15, the RMSE of both models decreases smoothly with values in the range of 3.5&#x2013;4.5. The prediction error of the VMD&#x2013;mRMR&#x2013;LSTM model is smaller than that of the VMD&#x2013;MIM&#x2013;LSTM model. Therefore, the VMD&#x2013;mRMR&#x2013;LSTM model is more robust than the VMD&#x2013;MIM&#x2013;LSTM model when the number of input features is changed.</p>
</sec>
<sec id="s3-6-3">
<title>3.5.3 Effect of Input Data</title>
<p>To investigate the effect of different input data on the accuracy of the model, the training set in the 3.1 section is used as the raw data input to the hybrid model. <xref ref-type="fig" rid="F10">Figure&#x20;10A</xref> shows the results of the training set decomposed by VMD. Comparing the decomposition results in <xref ref-type="fig" rid="F6">Figure&#x20;6A</xref>, <xref ref-type="fig" rid="F10">10A</xref>, it can be seen that the trend of the training set is similar to the original data. Further comparing the spectrograms in <xref ref-type="fig" rid="F6">Figure&#x20;6B</xref>, <xref ref-type="fig" rid="F10">10B</xref>, although the peak frequency of the training data set and the original data set is different, they appear at the same locations. The decomposition results of the training set are input into the hybrid model, and the prediction result is shown in <xref ref-type="fig" rid="F11">Figure&#x20;11</xref>. The predicted values of MAE, RMSE, and <italic>R</italic>
<sup>2</sup> are 30.96&#xa0;kWh, 38.96&#xa0;kWh, and 0.95, respectively. Compared with the results of decomposing the original data (VMD&#x2013;mRMR&#x2013;LSTM), the <italic>R</italic>
<sup>2</sup> of the decomposed training dataset model (Training_Set-VMD&#x2013;mRMR&#x2013;LSTM) decreased by 2%, and the RMSE and MAE increased by 27 and 44.9%, respectively, indicating that the selection of the input data can have an impact on the accuracy of the model. Training_Set- VMD&#x2013;mRMR&#x2013;LSTM still has higher accuracy than the single LSTM model, and the <italic>R</italic>
<sup>2</sup> improved by 7%, RMSE and MAE reduced by 42.4 and 25.9%, respectively. In conclusion, the use of training data as model input reduces the accuracy of the model, but the impact is small in general. Compared with a single model, the proposed hybrid model still has a greater superiority.</p>
<fig id="F10" position="float">
<label>FIGURE 10</label>
<caption>
<p>
<bold>(A)</bold> The decomposition results of training set. <bold>(B)</bold> The decomposition spectrogram of training&#x20;set.</p>
</caption>
<graphic xlink:href="fenrg-09-772508-g010.tif"/>
</fig>
<fig id="F11" position="float">
<label>FIGURE 11</label>
<caption>
<p>Prediction results of model with training&#x20;set.</p>
</caption>
<graphic xlink:href="fenrg-09-772508-g011.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec id="s4">
<title>4 Conclusion</title>
<p>A hybrid short-term load forecasting model, namely VMD&#x2013;mRMR&#x2013;LSTM, was proposed in this paper. To solve the modal mixing problem presented by the EMD algorithm, the VMD algorithm was used, and the value of its decomposition number K was determined by the average instantaneous frequency. For feature selection, the mRMR algorithm was used to select the related feature by analysing the correlation between each component and feature as well as the redundancy between features. Finally, the LSTM model was used for the prediction model. The case study in this paper demonstrated the following:<list list-type="simple">
<list-item>
<p>1) Compared to single prediction models, hybrid models have higher accuracy and are more robust in the field of energy consumption prediction and have a broad application prospect for the short-term prediction of building energy consumption.</p>
</list-item>
<list-item>
<p>2) Using VMD to decompose the original sequence can have a better decomposition effect than when EEMD is used. Decomposition by VMD solves the problem of modal confusion so that the decomposed sequence is stable. The prediction results of the hybrid model using VMD are higher than those of the hybrid model using&#x20;EEMD.</p>
</list-item>
<list-item>
<p>3) The mRMR algorithm can eliminate the redundancy between features and show the influencing factors of the modal components. The experimental results prove that the features selected by the mRMR algorithm have a higher prediction accuracy and better interpretability than those selected by MIM, which is supported by the value of <italic>R</italic>
<sup>2</sup> increasing by 3%, the value of MAE decreasing by 35.6% and the value of RMSE decreasing by&#x20;34.1%.</p>
</list-item>
<list-item>
<p>4) The hybrid model proposed in this paper can achieve an <italic>R</italic>
<sup>2</sup> value of 0.97, and its prediction results are higher than those of the single model (LSTM) and the general integrated model (VMD&#x2013;MIM&#x2013;LSTM). Therefore, the proposed VMD&#x2013;mRMR&#x2013;LSTM approach has a high potential for practical applications in energy systems, such as forecasting building energy consumption.</p>
</list-item>
<list-item>
<p>5) By varying the number of input feature parameters and the number of neurons in the hidden layer, the model is proven to have good robustness.</p>
</list-item>
</list>
</p>
<p>In this paper, all decomposed components were predicted using the LSTM model. However, since the frequency of each component varied, the LSTM may not have produced ideal results for each component. For example, in <xref ref-type="table" rid="T3">Table&#x20;3</xref>, the difference between the <italic>R</italic>
<sup>2</sup> values of the <italic>u</italic>
<sub>
<italic>1</italic>
</sub> and <italic>u</italic>
<sub>
<italic>4</italic>
</sub> components for the VMD&#x2013;mRMR&#x2013;LSTM model was not negligible. Choosing appropriate prediction models for the different frequency components may lead to better results. We will conduct more research in this direction in the future.</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>The data analyzed in this study is subject to the following licenses/restrictions: This study is applicable to the air conditioning cooling and heating load data of various building users. Requests to access these datasets should be directed to FQ, <email>qianfanyue91@163.com</email>.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>YR gave guidance on the framework and ideas of the paper. GW built the load forecasting model and compared it with conventional method. HM gave guidance on the framework and ideas of the paper. FQ gave guidance on the framework and ideas of the paper.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This research has been supported by the national key R&#x26;D project (No.2020YFD1100504-05).</p>
</sec>
<sec sec-type="COI-statement" id="s7">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Amasyali</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>El-Gohary</surname>
<given-names>N. M.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>A Review of Data-Driven Building Energy Consumption Prediction Studies[J]</article-title>. <source>Renew. Sust. Energ. Rev.</source> <volume>81</volume>. <pub-id pub-id-type="doi">10.1016/j.rser.2017.04.095</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>de Paiva</surname>
<given-names>G. M.</given-names>
</name>
<name>
<surname>Pires Pimentel</surname>
<given-names>Sergio.</given-names>
</name>
<name>
<surname>Pinheiro Alvarenga</surname>
<given-names>Bernardo.</given-names>
</name>
<name>
<surname>Marra</surname>
<given-names>E. G.</given-names>
</name>
<name>
<surname>Mussetta</surname>
<given-names>Marco.</given-names>
</name>
<name>
<surname>Leva</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Multiple Site Intraday Solar Irradiance Forecasting by Machine Learning Algorithms: MGGP and MLP Neural Networks[J]</article-title>. <source>Energies</source> <volume>13</volume> (<issue>11</issue>). </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dong</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Rahman</surname>
<given-names>S. M. M.</given-names>
</name>
<name>
<surname>Vega</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>A Hybrid Model Approach for Forecasting Future Residential Electricity Consumption</article-title>. <source>Energy and Buildings</source> <volume>117</volume>, <fpage>341</fpage>&#x2013;<lpage>351</lpage>. <pub-id pub-id-type="doi">10.1016/j.enbuild.2015.09.033</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dragomiretskiy</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zosso</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Variational Mode Decomposition</article-title>. <source>IEEE Trans. Signal. Process.</source> <volume>62</volume> (<issue>3</issue>), <fpage>531</fpage>&#x2013;<lpage>544</lpage>. <pub-id pub-id-type="doi">10.1109/TSP.2013.2288675</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Eom</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Clarke</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>S. H.</given-names>
</name>
<name>
<surname>Kyle</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Patel</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>China&#x27;s Building Energy Demand: Long-Term Implications from a Detailed Assessment</article-title>. <source>Energy</source> <volume>46</volume> (<issue>1</issue>), <fpage>405</fpage>&#x2013;<lpage>419</lpage>. <pub-id pub-id-type="doi">10.1016/j.energy.2012.08.009</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guo</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Multi-step Forecasting for Wind Speed Using a Modified EMD-Based Artificial Neural Network Model[J]</article-title>. <source>Renew. Energ.</source> <volume>37</volume> (<issue>1</issue>). </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Feng</surname>
<given-names>Z-k.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>A Hybrid Short-Term Load Forecasting Model Based on Variational Mode Decomposition and Long Short-Term Memory Networks Considering Relevant Factors with Bayesian Optimization Algorithm[J]</article-title>. <source>Appl. Energ.</source> <volume>237</volume>. <pub-id pub-id-type="doi">10.1016/j.apenergy.2019.01.055</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2018</year>). <source>A Hybrid System for Short-Term Wind Speed forecasting[J]</source>. <publisher-loc>Kidlington, Oxford</publisher-loc>: <publisher-name>Elsevier BV</publisher-name>, <fpage>226</fpage>. </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hochreiter</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Schmidhuber</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>1997</year>). <article-title>Long Short-Term Memory</article-title>. <source>Neural Comput.</source> <volume>9</volume> (<issue>8</issue>), <fpage>1735</fpage>&#x2013;<lpage>1780</lpage>. <pub-id pub-id-type="doi">10.1162/neco.1997.9.8.1735</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>A Hybrid Forecasting Approach Applied to Wind Speed Time Series[J]</article-title>. <source>Renew. Energ.</source> <volume>60</volume>. <pub-id pub-id-type="doi">10.1016/j.renene.2013.05.012</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang Norden</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Long Steven</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>M. C.</given-names>
</name>
<name>
<surname>Shih Hsing</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>Q.</given-names>
</name>
<etal/>
</person-group> (<year>1998</year>). <article-title>The Empirical Mode Decomposition and the Hilbert Spectrum for Nonlinear and Non-stationary Time Series Analysis[J]</article-title>. <source>Proc. R. Soc. A: Math. Phys. Eng. Sci.</source>, <fpage>454</fpage>. </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kong</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Dong</surname>
<given-names>Z. Y.</given-names>
</name>
<name>
<surname>Jia</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>David</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hill</surname>
<given-names>Y. X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Short-Term Residential Load Forecasting Based on LSTM Recurrent Neural Network.[J]</article-title>. <source>IEEE Trans. Smart Grid</source> <volume>10</volume> (<issue>1</issue>). </citation>
</ref>
<ref id="B13">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Z. Y.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>J.&#x20;B.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Power Load Forecasting Based on the Combined Model of LSTM and XGBoost[C]//PRAI &#x2019;19</article-title>,&#x201d; in <source>Proceedings of the 2019 the International Conference on Pattern Recognition and Artificial Intelligence</source> (<publisher-loc>Wenzhou, China</publisher-loc>: <publisher-name>ACM</publisher-name>), <fpage>46</fpage>&#x2013;<lpage>51</lpage>. </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Tao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Ao</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bai</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Improving Forecasting Accuracy of Daily enterprise Electricity Consumption Using a Random forest Based on Ensemble Empirical Mode Decomposition[J]</article-title>. <source>Energy</source> <volume>165</volume>. <pub-id pub-id-type="doi">10.1016/j.energy.2018.10.113</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lin</surname>
<given-names>Y .</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Combined Model Based on EMD-SVM for Short-Term Wind Power Prediction</article-title>. <source>Proc. CSEE</source> <volume>31</volume>, <fpage>102</fpage>&#x2013;<lpage>108</lpage>. </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Tian</surname>
<given-names>H-Q.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y-F.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>A Hybrid Model for Wind Speed Prediction Using Empirical Mode Decomposition and Artificial Neural Networks[J]</article-title>. <source>Renew. Energ.</source> <volume>48</volume>. <pub-id pub-id-type="doi">10.1016/j.renene.2012.06.012</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Mi</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Smart Multi-step Deep Learning Model for Wind Speed Forecasting Based on Variational Mode Decomposition, Singular Spectrum Analysis, LSTM Network and ELM[J]</article-title>. <source>Energ. Convers. Manag.</source> <volume>159</volume>. <pub-id pub-id-type="doi">10.1016/j.enconman.2018.01.010</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Short-term Wind Speed Forecasting Based on the Jaya-SVM Model[J]</article-title>. <source>Int. J.&#x20;Electr. Power Energ. Syst.</source> <volume>121</volume>. <pub-id pub-id-type="doi">10.1016/j.ijepes.2020.106056</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Lun</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Accuracy Analyses and Model Comparison of Machine Learning Adopted in Building Energy Consumption Prediction[J]</article-title>. <source>Energy Exploration &#x26; Exploitation</source> <volume>37</volume> (<issue>4</issue>). <pub-id pub-id-type="doi">10.1177/0144598718822400</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Marcello Anderson</surname>
<given-names>F. B.</given-names>
</name>
<name>
<surname>Lima</surname>
<given-names>P. C. M. C.</given-names>
</name>
<name>
<surname>Carneiro</surname>
<given-names>T. C.</given-names>
</name>
<name>
<surname>Leite</surname>
<given-names>J.&#x20;R.</given-names>
</name>
<name>
<surname>Luiz</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Neto</surname>
<given-names>D. B.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Portfolio Theory Applied to Solar and Wind Resources Forecast[J]</article-title>. <source>IET Renew. Power Generation</source> <volume>11</volume> (<issue>7</issue>). <pub-id pub-id-type="doi">10.1049/iet-rpg.2017.0006</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Muhammad</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Francesco</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Sonia</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Marco</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Comparison of echo State Network and Feed-Forward Neural Networks in Electrical Load Forecasting for Demand Response Programs[J]</article-title>. <source>Mathematics Comput. Simulation</source>, <fpage>184</fpage>. </citation>
</ref>
<ref id="B22">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Niu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2020</year>). <source>A Hybrid Stock price index Forecasting Model Based on Variational Mode Decomposition and LSTM network[J]</source>. <publisher-loc>Dordrecht, Netherlands</publisher-loc>: <publisher-name>Springer US</publisher-name>. </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Novakovic</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Strbac</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Bulatovic</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Toward Optimal Feature Selection Using Ranking Methods and Classification Algorithms</article-title>. <source>Yugoslav J.&#x20;Operation</source> <volume>21</volume> (<issue>1</issue>), <fpage>119</fpage>&#x2013;<lpage>135</lpage>. <pub-id pub-id-type="doi">10.2298/YJOR1101119N</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pan</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Application of XGBoost Algorithm in Hourly PM2.5 Concentration Prediction[J]</article-title>. <source>IOP Conf. Series:Earth Environ. Sci.</source>, <volume>113</volume>, <fpage>1</fpage>&#x2013;<lpage>7</lpage>. </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pei</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Qin</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Yao</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Multi-Step Ahead Short-Term Load Forecasting Using Hybrid Feature Selection and Improved Long Short-Term Memory Network[J]</article-title>. <source>Energies</source> <volume>13</volume> (<issue>16</issue>). <pub-id pub-id-type="doi">10.3390/en13164121</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Peng</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Long</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Ding</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Feature Selection Based on Mutual Information: Criteria of max-dependency, max-relevance, and Min-Redundancy</article-title>. <source>IEEE Trans. Pattern Anal. Mach Intell.</source> <volume>27</volume> (<issue>8</issue>), <fpage>1226</fpage>&#x2013;<lpage>1238</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2005.159</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Said</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>A New Ensemble Empirical Mode Decomposition (EEMD) Denoising Method for Seismic Signals[J]</article-title>. <source>Energ. Proced.</source> <volume>97</volume>. </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Short-Term Wind Power Forecasting on Multiple Scales Using VMD Decomposition, K-Means Clustering and LSTM Principal Computing[J]</article-title>. <source>IEEE Access</source> <volume>7</volume>. <pub-id pub-id-type="doi">10.1109/access.2019.2942040</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Short-Term Wind Power Forecasting on Multiple Scales Using VMD Decomposition, K-Means Clustering and LSTM Principal Computing</article-title>. <source>IEEE Access</source> <volume>7</volume>, <fpage>166917</fpage>&#x2013;<lpage>166929</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2019.2942040</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>HyperAdam: A Learnable Task-Adaptive Adam for Network Training[J]</article-title>. <source>Proc. AAAI Conf. Artif. Intelligence</source> <volume>33</volume>. <pub-id pub-id-type="doi">10.1609/aaai.v33i01.33015297</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>Y-X.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>Q-B.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>J-Q.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Improved EEMD-Based Crude Oil price Forecasting Using LSTM Networks[J]</article-title>. <source>Physica A: Stat. Mech. its Appl.</source>, <fpage>516</fpage>. </citation>
</ref>
<ref id="B32">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>Y. J.</given-names>
</name>
</person-group> (<year>2016</year>). <source>Research on Fault Diagnosis of Wind Turbine Transmission System Based on Variational Mode Decomposition. Dissertation</source>. <publisher-loc>Beijing</publisher-loc>: <publisher-name>North China Electric Power University</publisher-name>. </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>N. E.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Ensemble Empirical Mode Decomposition: A Noise-Assisted Data Analysis Method[J]</article-title>. <source>Adv. Adaptive Data Anal.</source> <volume>1</volume> (<issue>1</issue>). </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Xi</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Energy Consumption Prediction of Chiller Based on Long Short-Term Memory [J]</article-title>. <source>Refrigeration &#x26; Air Conditioning</source> <volume>34</volume> (<issue>06</issue>), <fpage>664</fpage>&#x2013;<lpage>669</lpage>. </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Kun</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Research on Variational Mode Decomposition in Rolling Bearings Fault Diagnosis of the Multistage Centrifugal Pump[J]</article-title>. <source>Mech. Syst. Signal Process.</source> <volume>93</volume>. <pub-id pub-id-type="doi">10.1016/j.ymssp.2017.02.013</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y .</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>A Hybrid EMD-SVM Based Short-Term Wind Power Forecasting Model</article-title>,&#x201d; in <conf-name>Proceedings of the 2015 IEEE PES Asia-Pacific Power and Energy Engineering Conference (APPEEC)</conf-name> (<publisher-loc>BrisbaneAustralia</publisher-loc>: <publisher-name>QLD</publisher-name>), <fpage>151</fpage>&#x2013;<lpage>185</lpage>. <pub-id pub-id-type="doi">10.1109/appeec.2015.7380872</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yan</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Aasma</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A Novel Deep Learning Framework: Prediction and Analysis of Financial Time Series Using CEEMD and LSTM[J]</article-title>. <source>Expert Syst. Appl.</source> <volume>159</volume>. <pub-id pub-id-type="doi">10.1016/j.eswa.2020.113609</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Yan</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Comparison of Building Energy Simulation Software: DeST, EnergyPlus and DOE-2[J]</article-title>. <source>Building Sci.</source> <volume>28</volume> (<issue>S2</issue>), <fpage>213</fpage>&#x2013;<lpage>222</lpage>. </citation>
</ref>
</ref-list>
</back>
</article>