<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Energy Res.</journal-id>
<journal-title>Frontiers in Energy Research</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Energy Res.</abbrev-journal-title>
<issn pub-type="epub">2296-598X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">730640</article-id>
<article-id pub-id-type="doi">10.3389/fenrg.2021.730640</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Energy Research</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Potential Analysis of the Attention-Based LSTM Model in Ultra-Short-Term Forecasting of Building HVAC Energy Consumption</article-title>
<alt-title alt-title-type="left-running-head">Xu et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Attention-Based LSTM HVAC Forecast</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Xu</surname>
<given-names>Yang</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1385180/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Gao</surname>
<given-names>Weijun</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Qian</surname>
<given-names>Fanyue</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Li</surname>
<given-names>Yanxue</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>Innovation Institute for Sustainable Maritime Architecture Research and Technology, Qingdao University of Technology, <addr-line>Qingdao</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>Faculty of Environmental Engineering, The University of Kitakyushu, <addr-line>Kitakyushu</addr-line>, <country>Japan</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/874728/overview">Yongping Sun</ext-link>, Hubei University of Economics, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/97763/overview">Yi Zong</ext-link>, Technical University of Denmark, Denmark</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/880530/overview">Muhammad Mohsin</ext-link>, Jiangsu University, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Weijun Gao, <email>gaoweijun@me.com</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Sustainable Energy Systems and Policies, a section of the journal Frontiers in Energy Research</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>23</day>
<month>08</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>9</volume>
<elocation-id>730640</elocation-id>
<history>
<date date-type="received">
<day>25</day>
<month>06</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>04</day>
<month>08</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Xu, Gao, Qian and Li.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Xu, Gao, Qian and Li</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Predicting system energy consumption accurately and adjusting dynamic operating parameters of the HVAC system in advance is the basis of realizing the model predictive control (MPC). In recent years, the LSTM network had made remarkable achievements in the field of load forecasting. This paper aimed to evaluate the potential of using an attentional-based LSTM network (A-LSTM) to predict HVAC energy consumption in practical applications. To evaluate the application potential of the A-LSTM model in real cases, the training set and test set used in experiments are the real energy consumption data collected by Kitakyushu Science Research Park in Japan. Pearce analysis was first carried out on the source data set and built the target database. Then five baseline models (A-LSTM, LSTM, RNN, DNN, and SVR) were built. Besides, to optimize the super parameters of the model, the Tree-structured of Parzen Estimators (TPE) algorithm was introduced. Finally, the applications are performed on the target database, and the results are analyzed from multiple perspectives, including model comparisons on different sizes of the training set, model comparisons on different system operation modes, graphical examination, etc. The results showed that the performance of the A-LSTM model was better than other baseline models, it could provide accurate and reliable hourly forecasting for HVAC energy consumption.</p>
</abstract>
<kwd-group>
<kwd>energy consumption prediction</kwd>
<kwd>ultra-short-term forecast</kwd>
<kwd>deep learning</kwd>
<kwd>LSTM network</kwd>
<kwd>attention mechanism</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>According to statistics, building energy consumption accounts for about 40% of the total energy consumption (<xref ref-type="bibr" rid="B31">Mohsin et&#x20;al., 2020</xref>), and the proportion of building carbon dioxide emissions is as high as 36% of the total emissions (<xref ref-type="bibr" rid="B12">Jradi et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B30">Mohsin et&#x20;al., 2021</xref>). Heating, ventilation, and air-conditioning (HVAC) systems account for 40% (or even higher) of the commercial building energy consumption (<xref ref-type="bibr" rid="B13">Kim et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B43">Wang et&#x20;al., 2021</xref>). Higher energy consumption usually means greater energy saving potential, so the research of energy-saving technology combining big data and artificial intelligence has become one of the hot spots in recent years (<xref ref-type="bibr" rid="B37">Sun et&#x20;al., 2021</xref>). Simulation modeling is mainly used to predict energy consumption in the design stage of HVAC. The factors considered include building physical parameters, outdoor meteorological parameters, indoor environmental parameters, and room usage (<xref ref-type="bibr" rid="B26">Ma et&#x20;al., 2017</xref>). Since there are different degrees of assumptions and simplifications in the process of establishing the model, there may be a large error between the predicted energy consumption in the design stage and the actual operation time. The data-driven forecasting model is mainly used to predict energy consumption in the operation stage of HVAC. The HVAC energy consumption is affected by many factors, not only by the building itself, but also by meteorological conditions (such as outdoor temperature and illumination), internal personnel activities (such as occupancy and power consumption), time lag, and the actual use of air conditioning (such as control deviation and operation scheme adjustment) (<xref ref-type="bibr" rid="B39">Wang et&#x20;al., 2016</xref>). Under the influence of these factors, the HVAC load curve has the characteristics of strong fluctuation, large randomness, and not obvious periodicity compared with the electric load curve. This presents great challenges in designing data-driven HVAC models. To solve this problem, Wasim Iqbal et&#x20;al. proposed the method of negative Binomial regression (NBR) model analysis (<xref ref-type="bibr" rid="B11">Iqbal et&#x20;al., 2021</xref>), BinbinYu proposed a method based on the Dynamic Spatial Panel Model (DSPM) Model (<xref ref-type="bibr" rid="B48">Yu, 2021</xref>), and WeiqingLi et&#x20;al. proposed the data envelopment analysis (DEA) and entropy method (<xref ref-type="bibr" rid="B15">Li et&#x20;al., 2021a</xref>) to analyze the interaction between various factors. At present, the operating parameters of the HVAC system in buildings are mainly set according to the load prediction in the design stage. It will lead to the HVAC system working in a state of low efficiency, resulting in a large amount of energy consumption.</p>
<p>To solve this problem, Model predictive control (MPC) was proposed and had been widely used (<xref ref-type="bibr" rid="B29">Mayne, 2014</xref>; <xref ref-type="bibr" rid="B36">Sultana et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B50">Zhan and Chong, 2021</xref>). There have been many examples of load forecasting applied to MPC in other energy fields. For example, Xueyuan Zhao proposed a short-term household load forecasting model to guide users&#x2019; power consumption behavior and adjust power grid load (<xref ref-type="bibr" rid="B51">Zhao et&#x20;al., 2021</xref>); Based on the forecast of wind power generation and electricity price, J.J.&#x20;Yang proposed a charge-discharge control strategy for energy storage equipment based on the data drive, which realized the maximization of income of energy storage equipment (<xref ref-type="bibr" rid="B46">Yang et&#x20;al., 2020a</xref>). Fangyuan Chang proposed a charging cost control solution based on reinforcement learning by using long and short-term memory networks (LSTM) to predict real-time electricity prices (<xref ref-type="bibr" rid="B4">Chang et&#x20;al., 2020</xref>). Jin Li proposed an adaptive genetic algorithm based on neighborhood search to optimize the total cost related to energy consumption and running time of electric vehicles. (<xref ref-type="bibr" rid="B14">Li et&#x20;al., 2020</xref>). This shows that predicting system energy consumption accurately and adjusting dynamic operating parameters of the HVAC system in advance is the basis of realizing MPC (<xref ref-type="bibr" rid="B7">Hazyuk et&#x20;al., 2012</xref>).</p>
<p>Load forecasting can be divided into ultra-short-term, short-term, medium-term, and long-term according to different purposes, and the ultra-short-term load forecast refers to the load forecast within 1&#xa0;h in the future (<xref ref-type="bibr" rid="B6">Guo et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B33">Qian et&#x20;al., 2020</xref>). Due to the time delay of building thermal load, the change of influencing factors in the short term cannot immediately change the HVAC energy consumption (<xref ref-type="bibr" rid="B10">Huang and Chow, 2011</xref>). Therefore, the prediction of ultra-short-term load can minimize the impact of the uncertainty of input variables on the prediction results. The behavior of HVAC can be expressed as time-series data with a certain periodically. The prediction model can learn the load mode of the system from the time-series data, and use these modes to make load predictions. After the prediction range is determined, the appropriate algorithm needs to be selected. The algorithms in the field of time series prediction can be divided into two categories: traditional machine learning algorithms and deep learning algorithms. Traditional machine learning algorithms are mostly based on statistical models, Current popular algorithms include Autoregressive Moving Average (ARIMA) (<xref ref-type="bibr" rid="B20">Liu et&#x20;al., 2021</xref>), support vector machine (SVM) (<xref ref-type="bibr" rid="B27">Ma et&#x20;al., 2018</xref>), Regression Tree (<xref ref-type="bibr" rid="B45">Yang et&#x20;al., 2020b</xref>), Random Forest (<xref ref-type="bibr" rid="B16">Li et&#x20;al., 2021b</xref>), and artificial neural networks (ANN) (<xref ref-type="bibr" rid="B44">Wei et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B3">Bui et&#x20;al., 2020</xref>).</p>
<p>In recent years, with the rapid development of deep learning, the deep neural network has been more and more applied in the field of load prediction. Deep learning is a series of new structures and new methods evolved based on multi-layer neural networks (<xref ref-type="bibr" rid="B25">Lv et&#x20;al., 2021a</xref>). Deep learning models have obvious advantages over traditional machine learning models in predicting multivariable time series problems. In the real world, time series prediction presents multiple challenges, such as having multiple input variables, the need to predict multiple time steps, and the need to perform the same type of prediction for multiple actual observation stations (<xref ref-type="bibr" rid="B1">Askari et&#x20;al., 2015</xref>). In particular, a deep learning model can support any number but a fixed number of inputs and outputs. Multivariable time series have multiple time-varying variables, each of which depends not only on its past value but also on other variables. These characteristics are correlated with each other, and in this case, multiple variables need to be considered to give the best-predicted energy consumption.</p>
<p>The most basic Deep learning model is Deep Neural Networks (DNN), also known as multi-layer perceptron (MLP). DNN has more hidden layers than ordinary artificial neural networks, which gives it the ability to learn complex patterns. Deb proposed a DNN - based model for predicting the daily cooling load of buildings (<xref ref-type="bibr" rid="B5">Deb et&#x20;al., 2016</xref>). Massana proposed a DNN model-based method for predicting short-term electrical loads in non-residential buildings (<xref ref-type="bibr" rid="B28">Massana et&#x20;al., 2015</xref>). Zhihan Lv et&#x20;al. proposed a layered DAE support vector machine (SDAE-SVM) model based on a three-layer neural network and achieved good prediction results (<xref ref-type="bibr" rid="B23">Lv et&#x20;al., 2021b</xref>). However, the DNN model cannot retain time-series information. It can only be predicted according to the current input and output values, and cannot learn the time dependence of data, which limits the accuracy of its prediction time series.</p>
<p>Recursive neural network (RNN), as a special deep neural network, can retain and consider the time variation of time series in the training process (<xref ref-type="bibr" rid="B9">Hochreiter and Schmidhuber, 1997</xref>), which makes it very suitable for time series data with periodicity. The time series of HVAC energy consumption is influenced by environmental variables and human living habits and has a strong periodicity. However, due to the problems of gradient explosion and gradient disappearance in the RNN network, the long-term dependence in time series cannot be retained, which limits the prediction accuracy of the RNN network. The long-short term memory (LSTM) network adds a series of multi-threshold gates based on the RNN network, which can deal with a long-term dependency relationship to a certain extent. LSTM networks were first used in natural language processing (<xref ref-type="bibr" rid="B38">Verwimp et&#x20;al., 2020</xref>), machine translation (<xref ref-type="bibr" rid="B35">Su et&#x20;al., 2020</xref>) and video recognition (<xref ref-type="bibr" rid="B24">Lv et&#x20;al., 2021c</xref>), etc. In recent years, LSTM networks have attracted more and more attention in the field of load prediction. The authors of (<xref ref-type="bibr" rid="B40">Wang et&#x20;al., 2019a</xref>) used the LSTM network to build a regional scale building energy consumption prediction model. Sendra-Arranz proposed a variety of multi-step prediction models based on the LSTM network to predict residential HVAC consumption (<xref ref-type="bibr" rid="B34">Sendra-Arranz and Guti&#xe9;rrez, 2020</xref>). Zhe Wang proposed a new method for predicting plug loads using the LSTM network. Through the data collected from a real office building in Berkeley, California, the prediction accuracy of this method was verified to be better than that of the traditional machine learning algorithm (<xref ref-type="bibr" rid="B42">Wang et&#x20;al., 2019b</xref>; <xref ref-type="bibr" rid="B41">Wang et&#x20;al., 2020</xref>).</p>
<p>Since the LSTM network adopts the code-decoding framework, the limitations of the code-decoding framework will lead to information loss when processing long time series. Bahdanau first introduced the attention mechanism into the code-decoding framework in 2014 (<xref ref-type="bibr" rid="B2">Bahdanau et&#x20;al., 2014</xref>). The attention mechanism can quantitatively assign a weight to each specific time step in the time series feature, which improves the attention distraction defect of traditional LSTM (<xref ref-type="bibr" rid="B18">Li et&#x20;al., 2019</xref>). As far as the author knows, the LSTM model based on the ATTENTION mechanism has not been applied in the field of HVAC energy consumption prediction, but some researchers have started experiments in other load fields and achieved some results (<xref ref-type="bibr" rid="B21">Lu et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B49">Yu et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B19">Liang, et&#x20;al, 2018</xref>). Such as Heidari used an attention-based LSTM (A-LSTM) model to predict the load of the solar-assisted hot water system and proved that the prediction accuracy of the A-LSTM model was better than that of the traditional LSTM model (<xref ref-type="bibr" rid="B8">Heidari and Khovalyg, 2020</xref>). Jince Li proposed an improved attention-based LSTM (A-LSTM) model for multivariate time series of predictions of two process industry cases (<xref ref-type="bibr" rid="B17">Li, et&#x20;al, 2021c</xref>). Tongguang Yang proposed an attention-based LSTM model to predict the day-ahead PV power output (<xref ref-type="bibr" rid="B47">Yang et&#x20;al., 2019</xref>). All these cases show that the A-LSTM model has significant advantages over the traditional LSTM model in dealing with time series problems.</p>
<p>This paper aims to develop a new HVAC ultra-short-term energy consumption prediction model. To achieve this goal, this paper first conducted potential rule analysis and feature engineering for 9&#xa0;years&#x2019; operation data of Kitakyushu Science and Research Park&#x2019;s (KSRP) Energy Center and constructed a data set for modeling. Then, we developed the LSTM model based on this data set and developed the A-LSTM model by adding the attention layer to the LSTM model. Besides, the RNN model, DNN model, and SVR model were also developed to compare performance. The hyper-parameters of the above models were optimized by the TPE algorithm to ensure prediction accuracy. Next, we used the data from the Energy Center from 2002 to 2009 as the training set and the data from 2010 as the test set to conduct experiments, and gradually reduced the size of data sets to evaluate the performance of the above five models in different training sets. Finally, we also evaluated the small-scale prediction effects of the above five models under four typical operating&#x20;modes.</p>
<p>Based on the literature review the main novelty of the paper can be summarised as follows:<list list-type="simple">
<list-item>
<p>&#x2022; This paper studies and verifies that the deep learning method with the attention mechanism has advantages in the memory and feature selection of time series information of air conditioning load prediction. As far as the author knows, the LSTM model based on the ATTENTION mechanism has not been applied in the field of HVAC energy consumption prediction.</p>
</list-item>
<list-item>
<p>&#x2022; In this paper, we use the Tree-Structured of Parzen Estimators (TPE) algorithm to optimize the parameters of each baseline model, to ensure that all the baseline models involved in horizontal comparison are in the best prediction performance, and to ensure the authenticity of the prediction performance comparison</p>
</list-item>
<list-item>
<p>&#x2022; In this paper, we used the real building data set instead of the standard data set to verify the load prediction potential of the A-LSTM model, which verified the value of this model in actual engineering. The results show that the A-LSTM model has advantages over the traditional machine learning algorithm in various operating modes, especially in the heating high load period, and has great potential in the field of ultra-short-term load prediction of&#x20;HVAC.</p>
</list-item>
</list>
</p>
</sec>
<sec sec-type="methods" id="s2">
<title>Methodology</title>
<sec id="s2-1">
<title>LSTM Neural Network</title>
<p>RNN network is a kind of neural network used to process time series. Compared with the traditional DNN network and CNN network, RNN adopts a cyclic structure to replace the hidden layer of the feedforward neural network. In the process of information transmission, there will be a part of the information left in the current neuron during each cycle, and the retained information will be used as the input of the next neural unit with the new information. In this way, the RNN network implements &#x201c;memory&#x201d;. However, when the input time series is too long, it is difficult to retain the information in the RNN network, this is also known as the phenomenon of gradient disappearance and gradient explosion (<xref ref-type="bibr" rid="B9">Hochreiter and Schmidhuber, 1997</xref>).</p>
<p>Based on the RNN unit, the input gate <inline-formula id="inf1">
<mml:math id="m1">
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, output gate <inline-formula id="inf2">
<mml:math id="m2">
<mml:mrow>
<mml:msub>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, forgetting gate <inline-formula id="inf3">
<mml:math id="m3">
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and cell state <inline-formula id="inf4">
<mml:math id="m4">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are added to the LSTM unit to control the inheritance and abandonment of information. There are three inputs of the LSTM unit: the input vector <inline-formula id="inf5">
<mml:math id="m5">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> at the current time slot t, the unit state <inline-formula id="inf6">
<mml:math id="m6">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> at the time slot t&#x2212;1, and the state of the hidden layer <inline-formula id="inf7">
<mml:math id="m7">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> at the time slot t&#x2212;1. The final output of the LSTM unit is the cell state <inline-formula id="inf8">
<mml:math id="m8">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> at the current time slot t and the state of the hidden layer at the current time <inline-formula id="inf9">
<mml:math id="m9">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> To figure out the <inline-formula id="inf10">
<mml:math id="m10">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, we first let the <inline-formula id="inf11">
<mml:math id="m11">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mtext>&#x2009;</mml:mtext>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>o</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf12">
<mml:math id="m12">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>f</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> be the weight matrix of the input gate, the output gate, and the forgetting gate, and let the <inline-formula id="inf13">
<mml:math id="m13">
<mml:mrow>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represent the combination of the hidden state at the moment t&#x2212;1 and the input of the unit at the time slot t into a new vector. Besides, let the <inline-formula id="inf14">
<mml:math id="m14">
<mml:mrow>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf15">
<mml:math id="m15">
<mml:mrow>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>o</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, and <inline-formula id="inf16">
<mml:math id="m16">
<mml:mrow>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>f</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> be their bias vectors, and let the <inline-formula id="inf17">
<mml:math id="m17">
<mml:mi>&#x3c3;</mml:mi>
</mml:math>
</inline-formula> represent the Sigmoid activation function. The formulas of the <inline-formula id="inf18">
<mml:math id="m18">
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf19">
<mml:math id="m19">
<mml:mrow>
<mml:msub>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, and <inline-formula id="inf20">
<mml:math id="m20">
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are shown below:<disp-formula id="e1">
<mml:math id="m21">
<mml:mrow>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x22c5;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>
<disp-formula id="e2">
<mml:math id="m22">
<mml:mrow>
<mml:msub>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>o</mml:mi>
</mml:msub>
<mml:mo>&#x22c5;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>o</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>
<disp-formula id="e3">
<mml:math id="m23">
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>f</mml:mi>
</mml:msub>
<mml:mo>&#x22c5;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>f</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>
</p>
<p>Finally, let <inline-formula id="inf21">
<mml:math id="m24">
<mml:mrow>
<mml:mtext>tan</mml:mtext>
<mml:mi>h</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> represent the activation function, <inline-formula id="inf22">
<mml:math id="m25">
<mml:mtext>&#x2a;</mml:mtext>
</mml:math>
</inline-formula> represent the Hadamard product, and let the <inline-formula id="inf23">
<mml:math id="m26">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>C</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> be the state of the intermediate unit input at time slot t, then we can calculate the <inline-formula id="inf24">
<mml:math id="m27">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> as follows:<disp-formula id="e4">
<mml:math id="m28">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>C</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mtext>tan</mml:mtext>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>C</mml:mi>
</mml:msub>
<mml:mo>&#x22c5;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>
<disp-formula id="e5">
<mml:math id="m29">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x2217;</mml:mo>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x2217;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>C</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>
<disp-formula id="e6">
<mml:math id="m30">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x2217;</mml:mo>
<mml:mtext>tan</mml:mtext>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>
</p>
</sec>
<sec id="s2-2">
<title>Attention Mechanism</title>
<p>Although the LSTM model has a memory function, and it can save some time-series information, but because the standard LSTM model uses the traditional encoder-decoder structure, it still has some limitations. When processing the time series <inline-formula id="inf25">
<mml:math id="m31">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, the Encoder will first encode the input sequence into a fixed-length implicit vector <inline-formula id="inf26">
<mml:math id="m32">
<mml:mi>h</mml:mi>
</mml:math>
</inline-formula> and give the same weight to the implicit vector. However, when the length of <inline-formula id="inf27">
<mml:math id="m33">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> increases, the average weight distribution will reduce the discrimination of <inline-formula id="inf28">
<mml:math id="m34">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, and some important time-series information will be ignored in the process of training the model, thus affecting the prediction accuracy of the&#x20;model.</p>
<p>The Attention mechanism is a mechanism used to optimize the Encoder-Decoder structural model, It can be combined with a variety of models depending on the actual situation. An encoder-decoder model with an Attention mechanism first learns the weight of each element from the sequence and then recombines the elements by weight. By assigning different weight parameters to each input element, the Attention mechanism can focus more on the parts that are relevant to the input element, thereby suppressing other useless information. Its biggest advantage is that it can consider global connection and local connection in one step, and can realize the parallel computation, which is particularly important for big data computation. In this paper, the Bahdanau algorithm (<xref ref-type="bibr" rid="B2">Bahdanau et&#x20;al., 2014</xref>) is adopted to realize the Attention mechanism, and the structure of the A-LSTM model adopted is shown in <xref ref-type="fig" rid="F1">Figure&#x20;1</xref>:</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>The structure of the A-LSTM model.</p>
</caption>
<graphic xlink:href="fenrg-09-730640-g001.tif"/>
</fig>
<p>The calculation process of the Attention layer is shown in <xref ref-type="disp-formula" rid="e7">Eqs 7</xref>&#x2013;<xref ref-type="disp-formula" rid="e9">9</xref>. <inline-formula id="inf29">
<mml:math id="m35">
<mml:mtext>i</mml:mtext>
</mml:math>
</inline-formula> denotes the moment; <inline-formula id="inf30">
<mml:math id="m36">
<mml:mi>j</mml:mi>
</mml:math>
</inline-formula> denotes the j element in the sequence; <inline-formula id="inf31">
<mml:math id="m37">
<mml:mi>T</mml:mi>
</mml:math>
</inline-formula> denotes the length of the sequence; <inline-formula id="inf32">
<mml:math id="m38">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> denotes the weight value; <inline-formula id="inf33">
<mml:math id="m39">
<mml:mrow>
<mml:mtext>&#xa0;</mml:mtext>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the matching degree between the element to be encoded and other elements.<disp-formula id="e7">
<mml:math id="m40">
<mml:mrow>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>
<disp-formula id="e8">
<mml:math id="m41">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>exp</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3a3;</mml:mi>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>t</mml:mi>
</mml:msubsup>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(8)</label>
</disp-formula>
<disp-formula id="e9">
<mml:math id="m42">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:munderover>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
<label>(9)</label>
</disp-formula>
</p>
</sec>
<sec id="s2-3">
<title>Tree-Structured of Parzen Estimators</title>
<p>The performance of machine learning models largely depends on the selection of hyperparameters. With the increase of model complexity and the amount of training data, automatic hyperparameter optimization plays an increasingly important role in the developments (<xref ref-type="bibr" rid="B32">Nguyen et&#x20;al., 2020</xref>). Compared with traditional manual parameter adjustment, automatic parameter optimization has the following advantages: 1) It reduces the manpower of development work. 2) Improve the performance of machine learning models. 3) Improve the reproducibility of the results (<xref ref-type="bibr" rid="B22">Luo et&#x20;al., 2021</xref>). In this paper, the Tree-Structured of Parzen Estimators (TPE) algorithm was used to achieve automatic optimization of the model&#x2019;s hyperparameters. TPE algorithm is an improved algorithm of Bayesian optimization algorithm (BO). It solves the limitation of the traditional BO algorithm to deal with classification parameters and conditional parameters, so it has higher efficiency.</p>
<p>The main process of the TPE algorithm is to convert the hyperparametric space into the nonparametric density distribution first, and then model the process <inline-formula id="inf34">
<mml:math id="m43">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. As shown in <xref ref-type="disp-formula" rid="e10">Eq. 10</xref>, TPE uses two density distributions of Equation to define <inline-formula id="inf35">
<mml:math id="m44">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf36">
<mml:math id="m45">
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mo>&#x3c;</mml:mo>
<mml:msup>
<mml:mi>y</mml:mi>
<mml:mtext>&#x2a;</mml:mtext>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> indicates that the value of the objective function is less than the threshold, and <inline-formula id="inf37">
<mml:math id="m46">
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mo>&#x2265;</mml:mo>
<mml:msup>
<mml:mi>y</mml:mi>
<mml:mtext>&#x2a;</mml:mtext>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> denotes that the value of the objective function is greater than or equal to the threshold.<disp-formula id="e10">
<mml:math id="m47">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mtext>if</mml:mtext>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>&#x3c;</mml:mo>
<mml:msup>
<mml:mi>y</mml:mi>
<mml:mo>&#x2217;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mtext>if</mml:mtext>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>&#x2265;</mml:mo>
<mml:msup>
<mml:mi>y</mml:mi>
<mml:mo>&#x2217;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(10)</label>
</disp-formula>
</p>
<p>The calculation of Expected Improvement (EI) is shown in <xref ref-type="disp-formula" rid="e11">Eqs 11</xref>&#x2013;<xref ref-type="disp-formula" rid="e13">13</xref>.<disp-formula id="e11">
<mml:math id="m48">
<mml:mrow>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x222b;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x221e;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mi>y</mml:mi>
<mml:mo>&#x2217;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:munderover>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi>y</mml:mi>
<mml:mo>&#x2217;</mml:mo>
</mml:msup>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
<mml:mtext>d</mml:mtext>
<mml:mi>y</mml:mi>
</mml:mrow>
</mml:math>
<label>(11)</label>
</disp-formula>
<disp-formula id="e12">
<mml:math id="m49">
<mml:mrow>
<mml:mi>&#x3b3;</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mo>&#x3c;</mml:mo>
<mml:msup>
<mml:mi>y</mml:mi>
<mml:mo>&#x2217;</mml:mo>
</mml:msup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(12)</label>
</disp-formula>
<disp-formula id="e13">
<mml:math id="m50">
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mstyle displaystyle="true">
<mml:mo>&#x222b;</mml:mo>
</mml:mstyle>
<mml:mi>R</mml:mi>
</mml:munder>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mtext>d</mml:mtext>
<mml:mi>y</mml:mi>
</mml:mrow>
</mml:math>
<label>(13)</label>
</disp-formula>
</p>
<p>Substitute <xref ref-type="disp-formula" rid="e12">Eqs 12</xref>, <xref ref-type="disp-formula" rid="e13">13</xref> into <xref ref-type="disp-formula" rid="e11">Eq. 11</xref> to get the final <xref ref-type="disp-formula" rid="e14">Eq. 14</xref>.<disp-formula id="e14">
<mml:math id="m51">
<mml:mrow>
<mml:mi>E</mml:mi>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mrow>
<mml:msup>
<mml:mi>y</mml:mi>
<mml:mo>&#x2217;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>&#x3b3;</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3b3;</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
<label>(14)</label>
</disp-formula>
</p>
<p>It can be seen from <xref ref-type="disp-formula" rid="e13">Eq. 13</xref> that point <inline-formula id="inf38">
<mml:math id="m52">
<mml:mrow>
<mml:msup>
<mml:mi>x</mml:mi>
<mml:mtext>&#x2a;</mml:mtext>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> with the largest Ei is the point with the smallest <inline-formula id="inf39">
<mml:math id="m53">
<mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>/</mml:mo>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. The TPE algorithm evaluates the improvement points according to <inline-formula id="inf40">
<mml:math id="m54">
<mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>/</mml:mo>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> in each iteration, and finally returns a point <inline-formula id="inf41">
<mml:math id="m55">
<mml:mrow>
<mml:msup>
<mml:mi>x</mml:mi>
<mml:mtext>&#x2a;</mml:mtext>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> with the largest EI. The corresponding process is shown in <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Flowchart of the TPE algorithm.</p>
</caption>
<graphic xlink:href="fenrg-09-730640-g002.tif"/>
</fig>
</sec>
</sec>
<sec id="s3">
<title>Case Study</title>
<sec id="s3-1">
<title>Case Introduction</title>
<p>In this paper, data collected by the CCHP system of Kitakyushu Science Research Park (KSRP) in Japan were used as the research object. The KSRP system was a distributed energy system consisting of a gas engine (160&#xa0;kW), a fuel cell (200&#xa0;kW), and a photovoltaic system (150&#xa0;kW). The system mainly supplied energy to the main teaching building of The University of Kitakyushu, which can meet the teaching and office needs of more than 3,000 people. The target building was divided into four floors, the first floor included the student center, meeting rooms and classrooms. The second to fourth floors were teachers&#x2019; offices and student research&#x20;rooms.</p>
<p>The system was powered by the gas engine, fuel cells, photovoltaic system and the utility grid. The cooling load, heating load and hot water load were mainly provided by the absorption chiller, while the gas engine and fuel cell also provided part of the cooling and heating load when generating electricity. The basic schematic diagram of CCHP system at KSRP was shown in <xref ref-type="fig" rid="F3">Figure&#x20;3</xref>. The system also included a detailed data acquisition system that recorded not only detailed operational data for each device, but also environmental data related to the target building. Using the temperature and flow data collected by the data acquisition system for hot and cold water supply and recovery, we could calculate the hot and cold load requirements of the target building in real-time. The KSRP cogeneration system was established in 2001. To make the model reflect the most real operating state of the system, we selected the data from 2002 to 2010 as the research object (78,820 pieces of data), because only 3&#xa0;days of system failure occurred during this period, which could reduce the impact of data missing on the modeling.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>The basic schematic diagram of the CCHP system at KSRP.</p>
</caption>
<graphic xlink:href="fenrg-09-730640-g003.tif"/>
</fig>
</sec>
<sec id="s3-2">
<title>Potential Analysis of Input Data Set</title>
<p>We first calculated the cooling and heating output of the equipment from January 1, 2002, to December 31, 2010, based on the gas consumption of the equipment and the annual average COP(cooling 1.00, heating 0.85). To verify the authenticity of the data, we also calculated the cooling and heating output based on the temperature and flow rate of cold and hot water supply and recovery collected by the system. To explore the distribution rule of these data in time series, we calculated the mean value of load in units of the month, week and hour respectively, and the results are shown in <xref ref-type="fig" rid="F4">Figure&#x20;4</xref>. As can be seen from <xref ref-type="fig" rid="F4">Figure&#x20;4A</xref>, the average load varies greatly each month. The annual peak value of total heating load output occurs in January, and that of total cooling load output occurs in August. Therefore, December, January, February, and March were defined as the heating season; July, August, and September as the cooling season; April, May, June, October, and November as the low-load season. The prediction effect of the model will be evaluated respectively according to this division. It can be seen from <xref ref-type="fig" rid="F4">Figure&#x20;4B</xref> that the cooling and heating loads are higher on weekdays than on weekends; It also can be seen from <xref ref-type="fig" rid="F4">Figure&#x20;4C</xref> that the average daily load distribution in the heating season and the cooling season is significantly different. All of the above time information can reflect the impact of human activities on load, so they can be used as characteristic factors for database construction.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>The diagram of average load distributed by the time: <bold>(A)</bold> average load distinguished by month, <bold>(B)</bold> average load distinguished by day of the week, <bold>(C)</bold> average load distinguished by hour.</p>
</caption>
<graphic xlink:href="fenrg-09-730640-g004.tif"/>
</fig>
<p>We also selected other environmental factors that might affect the heat and cold output to build the initial database, including data collected by the Energy Center every hour from January 1, 2002, to December 31, 2010, a total of 78,820 pieces of data. Each group of data includes time information, outdoor temperature (&#xb0;C), relative humidity (%), irradiance (<inline-formula id="inf42">
<mml:math id="m56">
<mml:mrow>
<mml:mi>W</mml:mi>
<mml:mi mathvariant="normal">&#x2215;</mml:mi>
<mml:msup>
<mml:mi>m</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>), wind speed (m/s), wind direction, and load output (The positive load indicates the heating load, and the negative load indicates the refrigeration load). In addition, the index &#x201c;trend&#x201d; indicates the ordinal number of the data in the time series. Since features with small correlation will provide unnecessary information in the training of the model, which will affect the robustness of the model, Pearson correlation analysis was conducted for all features, and the results are shown in <xref ref-type="fig" rid="F5">Figure&#x20;5</xref>. It can be seen that the biggest factor affecting the load is outdoor temperature, followed by the time serial number. The correlation between wind direction and load was too small (&#x2212;0.0087), and we had deleted this feature in later modeling. Examples of these data were shown in <xref ref-type="table" rid="T1">Table&#x20;1</xref>. In the following experiments, we used the data from 2002 to 2009 as the training set, the data from 2010 as the test set, and then randomly extracted 20% data from the training set as the verification set. Since the numerical dimensions of different variables are greatly different, it is necessary to normalize the data so that the data can be uniformly mapped to the interval [0,1]. In the next section, based on this data set, we will build a prediction model that can predict the next hour&#x2019;s load according to the input data of the first N&#xa0;hours.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Correlation analysis between the available features.</p>
</caption>
<graphic xlink:href="fenrg-09-730640-g005.tif"/>
</fig>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Example of the database.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Trend</th>
<th align="center">Month</th>
<th align="center">Weekday</th>
<th align="center">Hour</th>
<th align="center">Temperature&#xa0;(&#xb0;C)</th>
<th align="center">Humidity&#xa0;(%)</th>
<th align="center">Illuminance (<inline-formula id="inf43">
<mml:math id="m57">
<mml:mrow>
<mml:mrow>
<mml:mi>W</mml:mi>
<mml:mo>/</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>m</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>)</th>
<th align="center">Windspeed (m/s)</th>
<th align="center">Load (kW)</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">0</td>
<td align="char" char=".">1</td>
<td align="char" char=".">5</td>
<td align="char" char=".">1</td>
<td align="char" char=".">1.7</td>
<td align="char" char=".">18</td>
<td align="char" char=".">0</td>
<td align="char" char=".">3.6</td>
<td align="char" char=".">0</td>
</tr>
<tr>
<td align="left">1</td>
<td align="char" char=".">1</td>
<td align="char" char=".">5</td>
<td align="char" char=".">2</td>
<td align="char" char=".">1.4</td>
<td align="char" char=".">19</td>
<td align="char" char=".">0</td>
<td align="char" char=".">9.1</td>
<td align="char" char=".">0</td>
</tr>
<tr>
<td align="left">2</td>
<td align="char" char=".">1</td>
<td align="char" char=".">5</td>
<td align="char" char=".">3</td>
<td align="char" char=".">1.4</td>
<td align="char" char=".">19</td>
<td align="char" char=".">0</td>
<td align="char" char=".">5.5</td>
<td align="char" char=".">147.208</td>
</tr>
<tr>
<td align="left">3</td>
<td align="char" char=".">1</td>
<td align="char" char=".">5</td>
<td align="char" char=".">4</td>
<td align="char" char=".">1.2</td>
<td align="char" char=".">20</td>
<td align="char" char=".">0</td>
<td align="char" char=".">6.9</td>
<td align="char" char=".">165.156</td>
</tr>
<tr>
<td align="left">4</td>
<td align="char" char=".">1</td>
<td align="char" char=".">5</td>
<td align="char" char=".">5</td>
<td align="char" char=".">1.1</td>
<td align="char" char=".">20</td>
<td align="char" char=".">0</td>
<td align="char" char=".">2.2</td>
<td align="char" char=".">111.105</td>
</tr>
<tr>
<td align="left">5</td>
<td align="char" char=".">1</td>
<td align="char" char=".">5</td>
<td align="char" char=".">6</td>
<td align="char" char=".">1.3</td>
<td align="char" char=".">19</td>
<td align="char" char=".">0</td>
<td align="char" char=".">5.7</td>
<td align="char" char=".">79.517</td>
</tr>
<tr>
<td align="left">6</td>
<td align="char" char=".">1</td>
<td align="char" char=".">5</td>
<td align="char" char=".">7</td>
<td align="char" char=".">1.6</td>
<td align="char" char=".">19</td>
<td align="char" char=".">0</td>
<td align="char" char=".">6.4</td>
<td align="char" char=".">127.376</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s4">
<title>Result and Discussion</title>
<sec id="s4-1">
<title>Model Parameter Setting</title>
<p>In this study, we used the Hyperopt framework to implement the TPE algorithm and automatically optimized the hyperparameters of all baseline models. The programming language is Python and the deep learning framework is TensorFlow2.0. Hyperopt is a Python library for hyperparametric optimization based on Bayesian optimization. It supports the optimization of continuous, discrete, and condition variables. Using the Hyperopt framework requires setting four parameters: specifying the objective function to be optimized, defining the search space with super parameters, Trails Database, and the search algorithm. This section will take the LSTM model as an example to outline the method of constructing the model. After the parameter optimization of the LSTM baseline model was completed, we added an attention layer after the hidden layer of the LSTM model to build the A-LSTM&#x20;model.</p>
<p>The LSTM baseline model needs to optimize four parameters, which are the time step L of each layer in LSTM (using the length of the previous data), the size of the hidden unit m of each layer, the size of the batch processing b in the training process (we used the two-layer LSTM structure, and set the same hidden unit for each layer by default), and the drop rate of the Dropout layer. To determine the range of L, we first performed autocorrelation analysis on load data to identify data cycle patterns, and the results are shown in <xref ref-type="fig" rid="F6">Figure&#x20;6</xref>. In <xref ref-type="fig" rid="F6">Figure&#x20;6</xref>, the X-axis represents &#x201c;hours&#x201d; and the Y-axis represents the autocorrelation coefficient. We found that the overall autocorrelation of the load is in the form of cycle decline, and the autocorrelation of the load is a cycle of 24&#xa0;h, which means the autocorrelation peak occurs every 24&#xa0;h. Therefore, we define the conditional parameters of L as (12, 24, 36, 48). To avoid overfitting, we added a dropout layer after each LSTM layer, and the conditional parameters of drop rate are (0.2, 0.3, 0.4, 0.5). Due to the limited computational force, based on ensuring the prediction accuracy, we set the conditional parameter sets of m and b based on the empirical method: m &#x2208;{32,64,128,256} and b &#x2208;{32,64,128,256} (<xref ref-type="bibr" rid="B42">Wang et&#x20;al., 2019b</xref>). We input the above conditional parameters into the Hyperopt framework and use the TPE algorithm to optimize the model&#x2019;s super parameters. <xref ref-type="fig" rid="F7">Figure&#x20;7</xref> shows the optimized RNN, LSTM, and A-LSTM model structure, the hyperparameters of these models are determined by the TPE algorithm.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Load autocorrelation analysis results.</p>
</caption>
<graphic xlink:href="fenrg-09-730640-g006.tif"/>
</fig>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Structure of RNN, LSTM, and A-LSTM&#x20;model.</p>
</caption>
<graphic xlink:href="fenrg-09-730640-g007.tif"/>
</fig>
<p>Since the above three models are all recurrent neural networks, we also set up the DNN model and the SVR model for horizontal comparison. These two models also input all the data 24&#xa0;h before the time slot t and the time data at the time slot t to predict and finally output the load data at the time slot t. The optimal hyperparameters of the DNN model optimized by the TPE algorithm are shown in <xref ref-type="table" rid="T2">Table&#x20;2</xref>. The optimal hyperparameters of the SVR model optimized by the TPE algorithm are shown in <xref ref-type="table" rid="T3">Table&#x20;3</xref>. The topology of the above five models depends on the characteristics of the KSRP dataset, so for the other datasets, the structure and hyperparameters of the model should be adjusted according to the data. Since the focus of this study was to explore the potential of the LSTM model with the attention mechanism in the field of load prediction. Therefore, on the premise of ensuring the prediction accuracy, the topological structure and input characteristics of the model were simplified as far as possible, to improve the generalization ability of the model and reduce the required computational&#x20;force.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Hyperparameters for DNN&#x20;model.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Model</th>
<th align="center">Layer1 Units</th>
<th align="center">Layer2 Units</th>
<th align="center">Layer3 Units</th>
<th align="center">Batch size</th>
<th align="center">Drop rate</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">DNN</td>
<td align="char" char=".">128</td>
<td align="char" char=".">64</td>
<td align="char" char=".">32</td>
<td align="char" char=".">64</td>
<td align="char" char=".">0.2</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Hyperparamers for SVR&#x20;model.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Model</th>
<th align="center">Kernel</th>
<th align="center">C</th>
<th align="center">Gamma</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">SVR</td>
<td align="center">Rbf</td>
<td align="char" char=".">97.227588</td>
<td align="char" char=".">0.001032</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4-2">
<title>Annual Prediction Performance Comparison</title>
<p>To evaluate the time series prediction effect of the A-LSTM model on this data set, we compared it with the same type of RNN, LSTM model, and DNN model without memory function in this experiment. All models have been trained and tested 5 times, and the final data used for comparison is the average of the five test results to reduce the errors caused by random numbers. Root Mean Square Error (RMSE), Mean Absolute Error (MAE), and R-Square Value (R2_SCORE) were used as indicators of the evaluation model, which were calculated according to <xref ref-type="disp-formula" rid="e15">Eqs 15</xref>&#x2013;<xref ref-type="disp-formula" rid="e17">17</xref>. The <inline-formula id="inf44">
<mml:math id="m58">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> denotes the real observations, <inline-formula id="inf45">
<mml:math id="m59">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> denotes the average of the observed value, <inline-formula id="inf46">
<mml:math id="m60">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> denotes the predicted value, N denotes the number of test samples.<disp-formula id="e15">
<mml:math id="m61">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>M</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>E</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msqrt>
<mml:mrow>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>n</mml:mi>
</mml:mfrac>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:munderover>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:math>
<label>(15)</label>
</disp-formula>
<disp-formula id="e16">
<mml:math id="m62">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>E</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>N</mml:mi>
</mml:mfrac>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(16)</label>
</disp-formula>
<disp-formula id="e17">
<mml:math id="m63">
<mml:mrow>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3a3;</mml:mi>
<mml:mrow>
<mml:mtext>i</mml:mtext>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x2dc;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3a3;</mml:mi>
<mml:mrow>
<mml:mtext>i</mml:mtext>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(17)</label>
</disp-formula>
</p>
<p>We first used the data of 8&#xa0;years (from 2002 to 2009) as the training set to train the models and evaluated the effect of the load forecast in 2010. The results of the five models respectively predicting the annual data of 2010 are shown in <xref ref-type="table" rid="T4">Table&#x20;4</xref>. It could be seen that although the prediction results of each model were close to, the prediction accuracy of the A-LSTM model was the highest. Compared with the second-best predicted LSTM, A-LSTM&#x2019;s RMSE decreased by 3.06%, MSE decreased by 6.54%, and R<sup>2</sup> value increased by 0.43%. The reason why the evaluation results are close is that the system operates under low or zero load for a large amount of time in a year, and the prediction error during these periods is very small, which may reduce the overall average prediction error. We will explore this phenomenon in the next section.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Comparison of prediction errors between different models.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left"/>
<th align="center">SVR</th>
<th align="center">DNN</th>
<th align="center">RNN</th>
<th align="center">LSTM</th>
<th align="center">A-LSTM</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">RMSE (kW)</td>
<td align="char" char=".">106.490</td>
<td align="char" char=".">78.788</td>
<td align="char" char=".">80.638</td>
<td align="char" char=".">77.340</td>
<td align="char" char=".">74.977</td>
</tr>
<tr>
<td align="left">MAE (kW)</td>
<td align="char" char=".">85.632</td>
<td align="char" char=".">48.752</td>
<td align="char" char=".">48.979</td>
<td align="char" char=".">47.929</td>
<td align="char" char=".">44.793</td>
</tr>
<tr>
<td align="left">R<sup>2</sup>
</td>
<td align="char" char=".">0.854</td>
<td align="char" char=".">0.922</td>
<td align="char" char=".">0.918</td>
<td align="char" char=".">0.925</td>
<td align="char" char=".">0.929</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To explore the influence of the size of the training set on the prediction accuracy of the model, we also conducted the following experiments: keeping the topological structure of the above five models unchanged, gradually reducing the training set in a unit of 2&#xa0;years, and the 2010 data were used as the test set to evaluate each model separately. The experimental results are shown in <xref ref-type="fig" rid="F8">Figure&#x20;8</xref>. We found that the prediction accuracy of each model decreases with the reduction of the training set. The experiment shows that the prediction accuracy of the A-LSTM model was the best when the data of 8, 6, and 4&#xa0;years were used as the training set. Compared with the suboptimal LSTM model, its RMSE decreased by 3.06, 10.86, and 11.29%, respectively. R<sup>2</sup> value increased by 0.43, 2.21, and 2.57%, respectively. However, when 2&#xa0;years&#x2019; data were used as the training set, the prediction accuracy of the A-LSTM model decreased significantly, and its prediction accuracy was only better than that of the SVR model. This indicates that the prediction accuracy of the A-LSTM model will increase with the length of training set, and the prediction accuracy of 4-years or 6-years data sets of the A-LSTM model has obvious advantages compared with other models. This means that when the length of training set is greater than a certain threshold (6&#x2013;8&#xa0;years), the advantage of its prediction accuracy will gradually decrease compared with other cyclic neural network models. Besides, when the length of the training set is less than A certain threshold value (4&#x2013;2&#xa0;years), the prediction accuracy of the A-LSTM model will decrease significantly.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>The prediction accuracy of each model under different lengths of training&#x20;set: <bold>(A)</bold> the comparison of RMSE, <bold>(B)</bold> the comparison of R<sup>2</sup> value.</p>
</caption>
<graphic xlink:href="fenrg-09-730640-g008.tif"/>
</fig>
</sec>
<sec id="s4-3">
<title>Prediction Performance Comparison at High and Low Loads</title>
<p>In the previous section, we discussed how a large number of zero-load and low-load forecasts over a year may affect the average error. To more intuitively evaluate the prediction effect of the A-LSTM model, we selected data of 2&#xa0;weeks in each of four periods in the 2010&#x20;years for comparison. Among the data selected for the experiment, two groups were high-load period data (2010.1.1 to 2010.1.14 and 2010.8.1 to 2010.1.14), and the other two groups were low-load period data (2010.3.1 to 2010.3.14 and 2010.5.16 to 2010.5.30). the results are shown in <xref ref-type="table" rid="T5">Table&#x20;5</xref>. As can be seen from <xref ref-type="table" rid="T4">Table&#x20;4</xref>, the prediction accuracy of the A-LSTM model was significantly higher than that of other models. In the high heating load stage, compared with the second-best predicted LSTM, A-LSTM&#x2019;s RMSE decreased by 10.02%, MSE decreased by 5.93%, and R<sup>2</sup> value increased by 2.59%; In the high cooling load stage, RMSE of A-LSTM decreased by 9.21%, MSE decreased by 8.80%, and R<sup>2</sup> value increased by 1.88%, compared with that of the second-best predicted LSTM. In the low heating load stage, compared with the second-best predicted LSTM, A-LSTM&#x2019;s RMSE decreased by 6.25%, MSE decreased by 3.94%, and R<sup>2</sup> value increased by 5.14%; In the low cooling load stage, RMSE of A-LSTM decreased by 5.24%, MSE decreased by 4.47%, and R<sup>2</sup> value increased by 2.36%, compared with that of the second-best predicted RNN. This indicates that compared with the low-load stage, A-LSTM in the high-load stage has an obvious improvement compared with other models, which also indicates that A-LSTM has more potential in peak prediction. It can be seen from <xref ref-type="table" rid="T5">Table&#x20;5</xref> that the prediction accuracy of the five models for the cooling load is higher than that for the heating load. Taking the A-LSTM model as an example, RMSE decreased by 18.515 (kW), MAE decreased by 15.733 (kW), and R<sup>2</sup> value increased by 0.051 in the peak cooling period compared with the peak heating period. This change was also evident during periods of low&#x20;load.</p>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>Performance of A-LSTM models compared to the baseline&#x20;model.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left"/>
<th colspan="5" align="center">2010.1.1&#x223c;2010.1.14</th>
<th colspan="5" align="center">2010.8.1&#x223c;2010.8.14</th>
</tr>
<tr>
<td align="left"/>
<td align="center">SVR</td>
<td align="center">DNN</td>
<td align="center">RNN</td>
<td align="center">LSTM</td>
<td align="center">A-LSTM</td>
<td align="center">SVR</td>
<td align="center">DNN</td>
<td align="center">RNN</td>
<td align="center">LSTM</td>
<td align="center">A-LSTM</td>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">RMSE (kW)</td>
<td align="char" char=".">123.508</td>
<td align="char" char=".">103.035</td>
<td align="char" char=".">93.796</td>
<td align="char" char=".">96.558</td>
<td align="char" char=".">86.876</td>
<td align="char" char=".">108.849</td>
<td align="char" char=".">78.129</td>
<td align="char" char=".">79.959</td>
<td align="char" char=".">75.294</td>
<td align="char" char=".">68.361</td>
</tr>
<tr>
<td align="left">MAE (kW)</td>
<td align="char" char=".">102.083</td>
<td align="char" char=".">72.996</td>
<td align="char" char=".">66.898</td>
<td align="char" char=".">66.192</td>
<td align="char" char=".">62.262</td>
<td align="char" char=".">88.199</td>
<td align="char" char=".">57.575</td>
<td align="char" char=".">54.966</td>
<td align="char" char=".">50.127</td>
<td align="char" char=".">46.529</td>
</tr>
<tr>
<td align="left">R<sup>2</sup>
</td>
<td align="char" char=".">0.737</td>
<td align="char" char=".">0.817</td>
<td align="char" char=".">0.839</td>
<td align="char" char=".">0.848</td>
<td align="char" char=".">0.870</td>
<td align="char" char=".">0.800</td>
<td align="char" char=".">0.896</td>
<td align="char" char=".">0.892</td>
<td align="char" char=".">0.904</td>
<td align="char" char=".">0.921</td>
</tr>
<tr>
<td align="left"/>
<td colspan="5" align="center">
<bold>2010.3.1&#x223c;2010.3.14</bold>
</td>
<td colspan="5" align="center">
<bold>2010.5.16&#x223c;2010.5.30</bold>
</td>
</tr>
<tr>
<td align="left"/>
<td align="center">
<bold>SVR</bold>
</td>
<td align="center">
<bold>DNN</bold>
</td>
<td align="center">
<bold>RNN</bold>
</td>
<td align="center">
<bold>LSTM</bold>
</td>
<td align="center">
<bold>A-LSTM</bold>
</td>
<td align="center">
<bold>SVR</bold>
</td>
<td align="center">
<bold>DNN</bold>
</td>
<td align="center">
<bold>RNN</bold>
</td>
<td align="center">
<bold>LSTM</bold>
</td>
<td align="center">
<bold>A-LSTM</bold>
</td>
</tr>
<tr>
<td align="left">RMSE (kW)</td>
<td align="char" char=".">99.075</td>
<td align="char" char=".">79.973</td>
<td align="char" char=".">81.230</td>
<td align="char" char=".">78.325</td>
<td align="char" char=".">73.427</td>
<td align="char" char=".">68.277</td>
<td align="char" char=".">54.021</td>
<td align="char" char=".">50.452</td>
<td align="char" char=".">53.515</td>
<td align="char" char=".">47.805</td>
</tr>
<tr>
<td align="left">MAE (kW)</td>
<td align="char" char=".">78.165</td>
<td align="char" char=".">48.027</td>
<td align="char" char=".">53.816</td>
<td align="char" char=".">49.141</td>
<td align="char" char=".">47.206</td>
<td align="char" char=".">56.184</td>
<td align="char" char=".">35.185</td>
<td align="char" char=".">31.627</td>
<td align="char" char=".">38.351</td>
<td align="char" char=".">29.855</td>
</tr>
<tr>
<td align="left">R<sup>2</sup>
</td>
<td align="char" char=".">0.521</td>
<td align="char" char=".">0.687</td>
<td align="char" char=".">0.678</td>
<td align="char" char=".">0.701</td>
<td align="char" char=".">0.737</td>
<td align="char" char=".">0.642</td>
<td align="char" char=".">0.776</td>
<td align="char" char=".">0.805</td>
<td align="char" char=".">0.781</td>
<td align="char" char=".">0.824</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To explain this phenomenon, the characteristic correlation coefficients of the cooling season and heating season were statistically analyzed. <xref ref-type="fig" rid="F9">Figure&#x20;9</xref> can explain the reasons for the above phenomena from one perspective. It can be seen from <xref ref-type="fig" rid="F9">Figure&#x20;9</xref> that the correlation coefficients between the load and other characteristics in the refrigeration season are higher than those in the heating season, especially the temperature, illumination, and humidity. This indicates that the output of cooling load is more affected by environmental factors, while the output of heat load is more affected by the laws of human production and life. The existing data cannot fully reflect the laws of human production and life, but it reflects the environmental factors more comprehensively, so this phenomenon occurs.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>Absolute value of correlation coefficients of load and other features in the database between heating season and cooling season.</p>
</caption>
<graphic xlink:href="fenrg-09-730640-g009.tif"/>
</fig>
<p>The actual prediction curves corresponding to <xref ref-type="table" rid="T5">Table&#x20;5</xref> are shown in <xref ref-type="fig" rid="F10">Figure&#x20;10</xref>. As can be seen from the fitting curve results, the load output of the KSRP system in the high load stage was distributed discretely, and it fluctuated greatly in the short term, as the peaks and troughs often appear alternately in the time series. Except for the SVR model, other models had a good fitting effect. By comparing R<sup>2</sup> values in <xref ref-type="table" rid="T5">Table&#x20;5</xref>, we also found that the prediction curve fitting rate of all models, including the A-LSTM model, was higher in the high load period than in the low load period. It could also be seen from <xref ref-type="fig" rid="F10">Figure&#x20;10</xref> that the curve fitting effect in the period of the high load was better than that in the period of low load. This indicates that the load in the low load stage is more affected by random factors and is more difficult to predict.</p>
<fig id="F10" position="float">
<label>FIGURE 10</label>
<caption>
<p>Predicted load use <italic>versus</italic> measured load use by different models for 2&#xa0;weeks as a test period: <bold>(A)</bold> Forecasting effect of high load period in the heating season, <bold>(B)</bold> Forecasting effect of high load period in the cooling season, <bold>(C)</bold> Forecasting effect of low load period in the heating season, <bold>(D)</bold> Forecasting effect of low load period in the cooling season.</p>
</caption>
<graphic xlink:href="fenrg-09-730640-g010.tif"/>
</fig>
<p>There are two limitations in the current study. First, due to the limited computational force, the search method adopted in the hyperparameter optimization in this paper is based on the conditional parameters, rather than the search based on the assignment interval. Although the conditional parameters based on the empirical method can ensure the accuracy of the prediction, it is undeniable that there is room for further optimization of the super parameters of the models. Secondly, the Bahdanau algorithm adopted is the classical gradient-based method to obtain the optimal solution. The gradient-based method has the advantage of easy implementation, but at the same time, it will bring premature convergence and the problem of falling into a locally optimal solution. Therefore, there is room for further optimization at the algorithm level of this&#x20;study.</p>
</sec>
</sec>
<sec sec-type="conclusion" id="s5">
<title>Conclusion</title>
<p>Predictive control had attracted more and more attention in building energy efficiency. Previous studies had shown that the HVAC system of large buildings was complex in structure, and its operation was affected by random environmental factors and human activities, so it was very challenging to predict its short-term HVAC load. In this paper, we first analyzed the underlying patterns in the data based on the actual operation data of KSRP Energy Center in 9&#xa0;years and then determined the factors used to establish the model according to the results of the Pearce correlation calculation. The results showed that the cooling and heating load of HVAC was most affected by the outdoor temperature, and the time of daily peak load was concentrated in a specific period.</p>
<p>Therefore, this paper proposed a new model combining the attention mechanism with the LSTM neural network, which was implemented by the following steps: First, according to the autocorrelation analysis results of HVAC load, we determined the data of the previous 24&#xa0;h as the time step to predict the load of the next hour. In the second step, we used the TPE optimization method to optimize the hyperparameters of the baseline LSTM model. The test results showed that the LSTM model with two layers of 64 neurons had the best prediction effect. Thirdly, we added the attention layer to the baseline LSTM model to build the A-LSTM model. Finally, we also set up RNN, DNN, and SVR models as horizontal comparison objects.</p>
<p>Finally, we took the data from KSRP Energy Center from 2002 to 2009 as the training set and the data from 2010 as the test set to test the above five models respectively. The results showed that the prediction accuracy of the A-LSTM model was the best. Compared with the LSTM model, the overall RMSE decreased by 3.06%, MSE decreased by 6.54%, and R<sup>2</sup> value increased by 0.43%. By progressively reducing the size of the training set, we found that the performance advantage of the A-LSTM model was most significant when the length of training set was between 4 and 6&#xa0;years. Besides, when the size of the training set dropped to 2&#xa0;years, the prediction accuracy of the A-LSTM model declined sharply, which indicates that it has limitations in predicting small sample data. To verify the impact of low-load and zero-load samples on the experimental results, we respectively selected four typical operating mode samples in 2010 for evaluation and drew the effect chart of the predicted results. The results showed that the prediction effect of the A-LSTM model for refrigeration load was better than that for heating load, and the prediction effect for the high load period is better than that for the low load period.</p>
<p>In conclusion, for the cooling and heating load prediction of large buildings, the introduction of the attention mechanism can not only effectively improve the prediction accuracy of the traditional LSTM model, but also improve the accuracy of peak prediction. However, in practical application, the prediction effect of this model for different operating modes is different, which in-depth influence mechanism and solutions need to be further analyzed. Besides, there is still room for optimization in the algorithm of attention mechanism. In future work, we will try to apply the A-LSTM model to real-time HVAC energy-saving control. Since the traditional MPC system is a model-based control system, it needs to model the controlled objects accurately, which may affect the generality of the model. Therefore, we are more inclined to adopt a model-free deep reinforcement learning (RL) algorithm to solve this problem, such as taking the predicted value as the observed state of the agent to improve the control accuracy of the RL&#x20;model.</p>
</sec>
</body>
<back>
<sec id="s6">
<title>Data Availability Statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>WG, FQ, and YL contributed to the conception and design of the study. YX organized the database. YX performed the statistical analysis. YX wrote the first draft of the manuscript. WG, FQ, and YL wrote sections of the manuscript. All authors contributed to manuscript revision, read, and approved the submitted version.</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s9" sec-type="disclaimer">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Askari</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Montazerin</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Zarandi</surname>
<given-names>M. H. F.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>A Clustering Based Forecasting Algorithm for Multivariable Fuzzy Time Series Using Linear Combinations of Independent Variables[J]</article-title>. <source>Appl. Soft Comput.</source> <volume>35</volume>, <fpage>151</fpage>&#x2013;<lpage>160</lpage>. </citation>
</ref>
<ref id="B2">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Bahdanau</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Cho</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Bengio</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Neural Machine Translation by Jointly Learning to Align and Translate[J]</article-title>. <source>arXiv</source>.</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bui</surname>
<given-names>D-K.</given-names>
</name>
<name>
<surname>Nguyen</surname>
<given-names>T. N.</given-names>
</name>
<name>
<surname>Ngo</surname>
<given-names>T. D.</given-names>
</name>
<name>
<surname>Nguyen-Xuan</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>An Artificial Neural Network (ANN) Expert System Enhanced with the Electromagnetism-Based Firefly Algorithm (EFA) for Predicting the Energy Consumption in Buildings[J]</article-title>. <source>Energy</source> <volume>190</volume>, <fpage>116370</fpage>. </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Alsafasfeh</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Control of Battery Charging Based on Reinforcement Learning and Long Short-Term Memory Networks[J]</article-title>. <source>Comput. Electr. Eng.</source> <volume>85</volume>, <fpage>106670</fpage>. </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Deb</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Eang</surname>
<given-names>L. S.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Santamouris</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Forecasting Diurnal Cooling Energy Load for Institutional Buildings Using Artificial Neural Networks[J]</article-title>. <source>Energy and Buildings</source> <volume>121</volume>, <fpage>284</fpage>&#x2013;<lpage>297</lpage>. </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guo</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>A Deep Learning Model for Short-Term Power Load and Probability Density Forecasting[J]</article-title>. <source>Energy</source> <volume>160</volume>, <fpage>1186</fpage>&#x2013;<lpage>1200</lpage>. </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hazyuk</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Ghiaus</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Penhouet</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Optimal Temperature Control of Intermittently Heated Buildings Using Model Predictive Control: Part II &#x2013; Control Algorithm[J]</article-title>. <source>Building Environ.</source> <volume>51</volume>, <fpage>388</fpage>&#x2013;<lpage>394</lpage>. </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Heidari</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Khovalyg</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Short-term Energy Use Prediction of Solar-Assisted Water Heating System: Application Case of Combined Attention-Based LSTM and Time-Series Decomposition[J]</article-title>. <source>Solar Energy</source> <volume>207</volume>, <fpage>626</fpage>&#x2013;<lpage>639</lpage>. </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hochreiter</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Schmidhuber</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>1997</year>). <article-title>Long Short-Term Memory[J]</article-title>. <source>Neural Comput.</source> <volume>9</volume> (<issue>8</issue>), <fpage>1735</fpage>&#x2013;<lpage>1780</lpage>. </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Chow</surname>
<given-names>T-T.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Uncertainty Shift in Robust Predictive Control Design for Application in CAV Air-Conditioning Systems[J]</article-title>. <source>Building Serv. Eng. Res. Tech.</source> <volume>32</volume> (<issue>4</issue>), <fpage>329</fpage>&#x2013;<lpage>343</lpage>. </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Iqbal</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>Y. M.</given-names>
</name>
<name>
<surname>Chau</surname>
<given-names>K. Y.</given-names>
</name>
<name>
<surname>Irfan</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Mohsin</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Nexus between Air Pollution and NCOV-2019 in China: Application of Negative Binomial Regression Analysis[J]</article-title>. <source>Process Saf. Environ. Prot.</source> <volume>150</volume>, <fpage>557</fpage>&#x2013;<lpage>565</lpage>. </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jradi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Veje</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>J&#xf8;rgensen</surname>
<given-names>B. N.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Deep Energy Renovation of the M&#xe6;rsk Office Building in Denmark Using a Holistic Design Approach[J]</article-title>. <source>Energy and Buildings</source> <volume>151</volume>, <fpage>306</fpage>&#x2013;<lpage>319</lpage>. </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kim</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Yamaguchi</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Kimura</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ko</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Ikeda</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Shimoda</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Urban Building Energy Modeling Considering the Heterogeneity of HVAC System Stock: A Case Study on Japanese Office Building Stock[J]</article-title>. <source>Energy and Buildings</source> <volume>199</volume>, <fpage>547</fpage>&#x2013;<lpage>561</lpage>. </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Electric Vehicle Routing Problem with Battery Swapping Considering Energy Consumption and Carbon Emissions[J]</article-title>. <source>Sustainability</source> <volume>12</volume> (<issue>24</issue>). </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Qi</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2021A</year>), <article-title>DTDR&#x2013;ALSTM: Extracting Dynamic Time-Delays to Reconstruct Multivariate Data for Improving Attention-Based LSTM Industrial Time Series Prediction Models[J]</article-title>, <source>Knowledge-Based Systems</source>, <volume>211</volume>, <fpage>106508</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2020.106508</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Cai</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Shen</surname>
<given-names>F.</given-names>
</name>
<etal/>
</person-group> (<year>2021B</year>). <article-title>Phenology-based Classification of Crop Species and Rotation Types Using Fused MODIS and Landsat Data: The Comparison of a random-forest-based Model and a Decision-Rule-Based Model[J]</article-title>. <source>Soil Tillage Res.</source> <volume>206</volume>, <fpage>104838</fpage>. </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Chien</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Hsu</surname>
<given-names>C-C.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Nawaz</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Iqbal</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2021C</year>). <article-title>Nexus between Energy Poverty and Energy Efficiency: Estimating the Long-Run Dynamics[J]</article-title>. <source>Resour. Pol.</source> <volume>72</volume>, <fpage>102063</fpage>. </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Kong</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>EA-LSTM: Evolutionary Attention-Based LSTM for Time Series Prediction[J]</article-title>. <source>Knowledge-Based Syst.</source> <volume>181</volume> (<issue>Oct.1</issue>), <fpage>1047851</fpage>&#x2013;<lpage>1047858</lpage>. </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Ke</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Yi</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2018</year>), <article-title>GeoMAN: Multi-Level Attention Networks for Geo-Sensory Time Series Prediction[C]</article-title>, <conf-name>Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence IJCAI-18</conf-name>.<conf-loc>Stockholm, Sweden</conf-loc>, <conf-date>July 13-19, 2018</conf-date>. </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>M-D.</given-names>
</name>
<name>
<surname>Ding</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Bai</surname>
<given-names>Y-L.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Application of Hybrid Model Based on Empirical Mode Decomposition, Novel Recurrent Neural Networks and the ARIMA to Wind Speed Prediction[J]</article-title>. <source>Energ. Convers. Manag.</source> <volume>233</volume>, <fpage>113917</fpage>. </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Xiong</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Parikh</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Socher</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Knowing when to Look: Adaptive Attention <italic>via</italic> A Visual Sentinel for Image Captioning[C]</article-title>, <conf-name>Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <conf-loc>Honolulu, HI, USA</conf-loc>, <conf-date>21-26 July 2017</conf-date>, <pub-id pub-id-type="doi">10.1109/CVPR.2017.345</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Luo</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Cai</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A Multi-Task Deep Learning Model for Short-Term Taxi Demand Forecasting Considering Spatiotemporal Dependences[J]</article-title>. <source>J.&#x20;Traffic Transportation Eng. (English Edition)</source> <volume>8</volume> (<issue>1</issue>), <fpage>83</fpage>&#x2013;<lpage>94</lpage>. </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lv</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Qiao</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Deep-Learning-Enabled Security Issues in the Internet of Things[J]</article-title>. <source>IEEE Internet Things J.</source> <volume>8</volume> (<issue>12</issue>), <fpage>9531</fpage>&#x2013;<lpage>9538</lpage>. </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lv</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Qiao</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Singh</surname>
<given-names>A. K.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Fine-grained Visual Computing Based on Deep Learning[J]</article-title>. <source>ACM Transactions on Multimidia Computing Communications and Applications</source>, <volume>17</volume>. <pub-id pub-id-type="doi">10.1145/3418215</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lv</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Singh</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Deep Learning for Security Problems in 5G Heterogeneous Networks[J]</article-title>. <source>IEEE Netw.</source> <volume>35</volume>, <fpage>67</fpage>&#x2013;<lpage>73</lpage>. </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ma</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Fang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Modeling of District Load Forecasting for Distributed Energy System[J]</article-title>. <source>Appl. Energ.</source> <volume>204</volume>, <fpage>181</fpage>&#x2013;<lpage>205</lpage>. </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ma</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Ye</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Applying Support Vector Machines to Predict Building Energy Consumption in China[J]</article-title>. <source>Clean. Energ. Clean. Cities</source> <volume>152</volume>, <fpage>780</fpage>&#x2013;<lpage>786</lpage>. </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Massana</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Pous</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Burgas</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Melendez</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Colomer</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Short-term Load Forecasting in a Non-residential Building Contrasting Models and Attributes[J]</article-title>. <source>Energy and Buildings</source> <volume>92</volume>, <fpage>322</fpage>&#x2013;<lpage>330</lpage>. </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mayne</surname>
<given-names>D. Q.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Model Predictive Control: Recent Developments and Future Promise[J]</article-title>. <source>Automatica</source> <volume>50</volume> (<issue>12</issue>), <fpage>2967</fpage>&#x2013;<lpage>2986</lpage>. </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mohsin</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hanif</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Taghizadeh-Hesary</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Abbas</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Iqbal</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Nexus between Energy Efficiency and Electricity Reforms: A DEA-Based Way Forward for Clean Power Development[J]</article-title>. <source>Energy Policy</source> <volume>149</volume>, <fpage>112052</fpage>. </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mohsin</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Taghizadeh-Hesary</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Panthamit</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Anwar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Abbas</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Vo</surname>
<given-names>X. V.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Developing Low Carbon Finance Index: Evidence from Developed and Developing Economies[J]</article-title>. <source>Finance Res. Lett.</source>, <fpage>101520</fpage>. </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nguyen</surname>
<given-names>H-P.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zio</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A Long-Term Prediction Approach Based on Long Short-Term Memory Neural Networks with Automatic Parameter Optimization by Tree-Structured Parzen Estimator and Applied to Time-Series Data of NPP Steam Generators[J]</article-title>. <source>Appl. Soft Comput.</source> <volume>89</volume>, <fpage>106116</fpage>. </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Qian</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Potential Analysis of the Transfer Learning Model in Short and Medium-Term Forecasting of Building HVAC Energy Consumption[J]</article-title>. <source>Energy</source> <volume>193</volume>, <fpage>116724</fpage>. </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sendra-Arranz</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Guti&#xe9;rrez</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A Long Short-Term Memory Artificial Neural Network to Predict Daily HVAC Consumption in Buildings[J]</article-title>. <source>Energy and Buildings</source> <volume>216</volume>, <fpage>109952</fpage>. </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Su</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Shi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Jian</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Shi</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Neural Machine Translation with Gumbel Tree-LSTM Based Encoder[J]</article-title>. <source>J.&#x20;Vis. Commun. Image Representation</source> <volume>71</volume>, <fpage>102811</fpage>. </citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sultana</surname>
<given-names>W. R.</given-names>
</name>
<name>
<surname>Sahoo</surname>
<given-names>S. K.</given-names>
</name>
<name>
<surname>Sukchai</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Yamuna</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Venkatesh</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>A Review on State of Art Development of Model Predictive Control for Renewable Energy Applications[J]</article-title>. <source>Renew. Sust. Energ. Rev.</source> <volume>76</volume>, <fpage>391</fpage>&#x2013;<lpage>406</lpage>. </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Awan</surname>
<given-names>R. U.</given-names>
</name>
<name>
<surname>Nawaz</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Mohsin</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Rasheed</surname>
<given-names>A. K.</given-names>
</name>
<name>
<surname>Iqbal</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Assessing the Socio-Economic Viability of Solar Commercialization and Electrification in South Asian Countries</article-title>, <source>Environment, Dev. Sustainability</source>, <volume>23</volume>, <fpage>9875</fpage>&#x2013;<lpage>9897</lpage>. </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Verwimp</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Van Hamme</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wambacq</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>State Gradients for Analyzing Memory in LSTM Language Models[J]</article-title>. <source>Comp. Speech Lang.</source> <volume>61</volume>, <fpage>101034</fpage>. </citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Event-driven Optimization of Complex HVAC Systems[J]</article-title>. <source>Energy and Buildings</source> <volume>133</volume>, <fpage>79</fpage>&#x2013;<lpage>87</lpage>. </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Hong</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Forecasting District-Scale Energy Dynamics through Integrating Building Network and Long Short-Term Memory Learning Algorithm[J]</article-title>. <source>Appl. Energ.</source> <volume>248</volume>, <fpage>217</fpage>&#x2013;<lpage>230</lpage>. </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Hong</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Piette</surname>
<given-names>M. A.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Building thermal Load Prediction through Shallow Machine Learning and Deep Learning[J]</article-title>. <source>Appl. Energ.</source> <volume>263</volume>, <fpage>114683</fpage>. </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Hong</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Piette</surname>
<given-names>M. A.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Predicting Plug Loads with Occupant Count Data through a Deep Learning Approach[J]</article-title>. <source>Energy</source> <volume>181</volume>, <fpage>29</fpage>&#x2013;<lpage>42</lpage>. </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yuan</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Srinivasan</surname>
<given-names>R. S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Practical Issues in Implementing Machine-Learning Models for Building Energy Efficiency: Moving beyond Obstacles[J]</article-title>. <source>Renew. Sust. Energ. Rev.</source> <volume>143</volume>, <fpage>110929</fpage>. </citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Pan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Prediction of Occupancy Level and Energy Consumption in Office Building Using Blind System Identification and Neural Networks[J]</article-title>. <source>Appl. Energ.</source> <volume>240</volume>, <fpage>276</fpage>&#x2013;<lpage>294</lpage>. </citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Tsui</surname>
<given-names>K-L.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Lifespan Prediction of Lithium-Ion Batteries Based on Various Extracted Features and Gradient Boosting Regression Tree Model[J]</article-title>. <source>J.&#x20;Power Sourc.</source> <volume>476</volume>, <fpage>228654</fpage>. </citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>J.&#x20;J.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>M. X.</given-names>
</name>
<name>
<surname>Du</surname>
<given-names>P. J.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Y. X.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A Deep Reinforcement Learning Method for Managing Wind Farm Uncertainties through Energy Storage System Control and External reserve Purchasing[J]</article-title>. <source>Int. J.&#x20;Electr. Power Energ. Syst.</source> <volume>119</volume>, <fpage>105928</fpage>. </citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Xun</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2019</year>), <article-title>LSTM-Attention-Embedding Model-Based Day-Ahead Prediction of Photovoltaic Power Output Us-Ing Bayesian Optimization[J]</article-title>. <source>IEEE Access</source>, <volume>7</volume>, <fpage>171471</fpage>-<lpage>171484</lpage>, 1-1.PP(99, <pub-id pub-id-type="doi">10.1109/ACCESS.2019.2954290</pub-id> </citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Urban Spatial Structure and Total-Factor Energy Efficiency in Chinese Provinces[J]</article-title>. <source>Ecol. Indicators</source> <volume>126</volume>, <fpage>107662</fpage>. </citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Fan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Tao</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2017</year>), <article-title>Multi-modal Factorized Bilinear Pooling with Co-attention Learning for Visual Question Answering[C]</article-title>, <conf-name>2017 IEEE International Conference on Computer Vision (ICCV)</conf-name>, <conf-loc>Venice, Italy</conf-loc>, <conf-date>22-29 Oct. 2017</conf-date>, <pub-id pub-id-type="doi">10.1109/ICCV.2017.202</pub-id> </citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Chong</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Data Requirements and Performance Evaluation of Model Predictive Control in Buildings: A Modeling Perspective[J]</article-title>. <source>Renew. Sust. Energ. Rev.</source> <volume>142</volume>, <fpage>110835</fpage>. </citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Qian</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Ge</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Electricity Cost Comparison of Dynamic Pricing Model Based on Load Forecasting in home Energy Management System[J]</article-title>. <source>Energy</source> <volume>229</volume>, <fpage>120538</fpage>. </citation>
</ref>
</ref-list>
</back>
</article>