<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article article-type="methods-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Energy Res.</journal-id>
<journal-title>Frontiers in Energy Research</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Energy Res.</abbrev-journal-title>
<issn pub-type="epub">2296-598X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">858895</article-id>
<article-id pub-id-type="doi">10.3389/fenrg.2022.858895</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Energy Research</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Microgrid Energy Management Strategy Base on UCB-A3C Learning</article-title>
<alt-title alt-title-type="left-running-head">Yang et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">UCB-A3C Based Microgrid Energy Management</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Yang</surname>
<given-names>Yanhong</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1449274/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Li</surname>
<given-names>Haitao</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1619078/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Shen</surname>
<given-names>Baochen</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Pei</surname>
<given-names>Wei</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1451104/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Peng</surname>
<given-names>Dajian</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Institute of Electrical Engineering</institution>, <institution>Chinese Academy of Sciences</institution>, <addr-line>Beijing</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Faculty of Information Technology</institution>, <institution>Beijing University of Technology</institution>, <addr-line>Beijing</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1394665/overview">Junhui Li</ext-link>, Northeast Electric Power University, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1278893/overview">Yassine Amirat</ext-link>, ISEN Yncr&#xe9;a Ouest, France</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1442315/overview">Qifang Chen</ext-link>, Beijing Jiaotong University, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Haitao Li, <email>lihaitao@bjut.edu.cn</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Smart Grids, a section of the journal Frontiers in Energy Research</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>23</day>
<month>03</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>10</volume>
<elocation-id>858895</elocation-id>
<history>
<date date-type="received">
<day>20</day>
<month>01</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>03</day>
<month>03</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Yang, Li, Shen, Pei and Peng.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Yang, Li, Shen, Pei and Peng</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>The uncertainty of renewable energy and demand response brings many challenges to the microgrid energy management. Driven by the recent advances and applications of deep reinforcement learning a microgrid energy management strategy, i.e.,&#x20;upper confidence bound based advantage actor-critic (A3C), is proposed to utilize a novel action exploration mechanism to learn the power output of wind power generation, the price of electricity trading and power load. The simulation results indicate that the UCB-A3C learning based energy management strategy is better than conventional PPO, actor critical and A3C algorithm.</p>
</abstract>
<kwd-group>
<kwd>microgrid</kwd>
<kwd>energy management</kwd>
<kwd>A3C</kwd>
<kwd>UCB</kwd>
<kwd>edge computing</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>In the context of transition towards sustainable and cleaner energy production, microgrid (MG) has become an effective way for tackling energy crisis and environmental pollution issues. The microgrid is a small-scale energy system consisting of distributed energy sources and loads, which can operate independently from, or in parallel with, the main power grid (<xref ref-type="bibr" rid="B17">Yang et&#x20;al., 2018</xref>). A typical microgrid system is illustrated in <xref ref-type="fig" rid="F1">Figure&#x20;1A</xref>, which includes distributed generation resources (DERs), energy storage systems (ESS) and electric loads. Establishment of microgrid by integrating local renewable energy sources and loads, provides reliability guarantee for local service and strengthen grid resilience, and is a significant step towards Smart Grids (<xref ref-type="bibr" rid="B1">Chen et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B9">Lee, 2022</xref>; <xref ref-type="bibr" rid="B10">Li et&#x20;al., 2020</xref>).</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>
<bold>(A)</bold> is the microgrid model, <bold>(B)</bold> is the microgrid edge computing architecture, <bold>(C)</bold> is the UCB-A3C algorithm, <bold>(D)</bold> is the neural network structure, <bold>(E)</bold> is the microgrid energy management strategy.</p>
</caption>
<graphic xlink:href="fenrg-10-858895-g001.tif"/>
</fig>
<p>With the fluctuation of renewable energy supply and the uncertainty of power load change, how to manage microgrid more efficiently is a major challenge. When dealing with microgrid energy storage management, model-based algorithms such as particle swarm algorithm and ant colony algorithm have been proposed to solve this problem (<xref ref-type="bibr" rid="B19">Zhang et&#x20;al., 2019</xref>). However, the dynamic characteristics of microgrid and the interaction between its components are described by building a model, which is not portable and scalable in practical application.</p>
<p>Recently, since the requirement of an explicit system model can be relaxed by learning-based scheme, this scheme has been introduced as an alternative to model-based approaches, and is used to improve the scalability of microgrid management (<xref ref-type="bibr" rid="B7">Kim et&#x20;al., 2022</xref>; <xref ref-type="bibr" rid="B3">Fan et al., 2021</xref>; <xref ref-type="bibr" rid="B14">Nakabi and Toivanen, 2021</xref>; <xref ref-type="bibr" rid="B15">Pourmousavi et al., 2010</xref>; <xref ref-type="bibr" rid="B18">Yu et al., 2019</xref>). The deep reinforcement learning paradigm, which treats the microgrid as a black box, is the most promising learning-based method to find an optimal microgrid energy management strategy from interactions with it. Recently, microgrid energy management adopting a variety of DRL methods, has been investigated, such as DQN (S.A et&#x20;al., 2010), SARSA (<xref ref-type="bibr" rid="B12">Ming et&#x20;al., 2017</xref>), and Double DQN (<xref ref-type="bibr" rid="B13">Mnih et&#x20;al., 2016</xref>).</p>
<p>Furthermore, literature (Finland, 2018) has explored A3C algorithm based on the policy gradient, and demonstrated that it has better performance than the value function based DRL algorithms in different microgrid operation scenarios. However, note that the conventional A3C approach adopts the heuristic <italic>&#x3b5;</italic>-greedy method in the process of exploration. It always chooses the current best action with probability 1-<italic>&#x3b5;</italic> or choose action randomly with probability <italic>&#x3b5;</italic>. This greedy exploration approach leads to computation complexity proportional to time and the learning performance may be deteriorated (<xref ref-type="bibr" rid="B2">Dong et&#x20;al., 2021</xref>). Based on these observations, an improved A3C learning algorithm with the novel exploration mechanism is proposed to deal with this problem in the learning process, which is benefit to real electricity price and renewable energy production in microgrid energy management.</p>
<p>On the other side, with the continuous advancement of new generation information and communication technologies in recent years, the newly emerging technologies, such as Internet-of-Things (iot), cloud computing and big data analytics, are deeply integrated to form the energy internet and enable new microgrid operational opportunities. Since the iot devices generate tremendous data during the microgrid operation, and they have the limited computation and storage capacity, it is necessary to introduce the cloud computing facilities to cope with these data. However, centralizing data processing in the cloud side would result in significant communication overhead and&#x20;delay.</p>
<p>To address this issue, the edge computing architecture, which performs computational tasks at the edge of the communication network, is able to bring the cloud computing in close to the internet of thing devices (Finland, 2018). It provides the opportunity to integrate the training and inference process of microgrid energy management based on DRL at the edge, which is different from the conventional centralized cloud computing platform for mission-critical and delay-sensitive applications. The edge computing architecture and corresponding cloud-edge coordination mechanism enable the edge gateway to execute the decision-making tasks for timely energy management. Therefore, the edge computing is considered a promising solution to significantly reduce communication delay, improve microgrid energy management performance and bring distributed intelligence for the microgrid system. And then, with the help of the UCB-A3C learning algorithm, we present an intelligent energy management policy in this industrial edge computing environments.</p>
<p>The main contributions of this paper can be summarized as follows.<list list-type="simple">
<list-item>
<p>1) We integrate energy iot communication and cloud-edge coordination and present an edge computing architecture for microgrid energy management and optimization problem. Further, we designed a Markov decision process (MDP) with an objective of minimizing the daily operating cost to model this energy management&#x20;issue.</p>
</list-item>
<list-item>
<p>2) To handle the formulated MDP optimization problem, we propose UCB based A3C learning algorithm with gross margin reward function, which can utilize a novel action exploration mechanism to learn the power output of wind power generation, the price of electricity trading and power&#x20;load.</p>
</list-item>
</list>
</p>
<p>The rest of this paper is organized as follows. <italic>System Model</italic> describes the microgrid energy management architecture with edge computing and the MDP model of energy optimization problem. Then, the UCB based A3C learning algorithm with better learning efficiency is proposed in <italic>UCB-A3C Based Energy Management</italic>. Further on, an energy management approach based on the proposed UCB-A3C learning algorithm is proposed in this section. The performance evaluation of the UCB-A3C based energy management strategy is analyzed with simulations in <italic>Performance Evaluation</italic>. Finally, the conclusions are drawn in <italic>Conclusion</italic>.</p>
</sec>
<sec id="s2">
<title>System Model</title>
<p>The proposed microgrid energy management architecture with edge computing is illustrated in <xref ref-type="fig" rid="F1">Figure&#x20;1B</xref>. It integrates energy iot platforms with edge computing to implement ubiquitous sensing, computing and communication, and can effectively deploy learning based microgrid operation functionalities. This architecture consists of microgrid equipment layer, edge layer and cloud layer. The microgrid equipment layer is composed of various power components and is responsible for supplying the electricity to meet the local demand. We assume that the microgrid includes a group of TCLs, a wind-based DER, a communal ESS and a group of residential price-responsive loads, and these components are managed by edge gateway. Moreover, the microgrid uses these components to trade electricity with the main network to achieve a balance between supply and demand. In this process, if the power required by the power load component is greater than the power generation capacity of wind power generation, it adjusts the energy storage component to dynamically balance the power purchased from the main grid. If the power required by the power load component is less than the power generation capacity of wind power generation, it adjusts the power sold by the energy storage component to the main grid for dynamic balance.</p>
<p>The edge layer, which includes edge platforms, edge gateways and edge services, is located between the cloud platform and the underlying physical equipment layer. It is the key part of the entire architecture and provides functions such as storage, computing and application on the edge side. The hardware platform of edge layer is edge gateway, which is composed of communication modules, storage units and computing units, and is leveraging to perform data acquisition, transmission and microgrid equipment control. Edge gateway can support communication protocols such as RS485, WiFi and 5G, as well as network transmission protocols such as HTTPS and MQTT. Under the co-scheduling of the edge platform, edge gateway can obtain microgrid operational states and control the microgrid equipment according to the instructions from the cloud server.</p>
<p>Edge computing platform is a software environment that is used to write and run software applications and is operated by the distributed edge gateways, and it is usually a standardized interoperability framework deployed on edge gateway to provide plug and play functions for iot sensors. All kinds of micro services for the microgrid based on some popular edge computing platform, such as EdgeX Foundry and KubeEdge, can run on edge platform. The edge platform aggregates microgrid equipment data which is collected from edge gateway. Meanwhile, it may store temporary sensor data, upload long-term data to the cloud center for data monitoring, analysis, storage and visualization, and receive the control instructions from the cloud side at the same time. In addition, edge service as the external interface of the whole microservice system, it collects iot equipment resources in the edge layer and provides services to users, accesses RESTful requests and forwards them to internal microservices. Moreover, in order to improve the capacity of computing and storage, edge service takes advantage of cloud-edge collaboration to cooperate edge resources with cloud resources, and supports powerful expansion capability to implement service mapping, request parsing, encryption and decryption, and authentication.</p>
<p>The cloud layer uses the cloud platform to provide various cloud services, we can deploy different cloud infrastructure environments, such as public cloud, private cloud or hybrid cloud, for different scale microgrid using the elastic expansion capability of the cloud platform, and run deep reinforce learning based intelligent microgrid energy management strategy in the cloud side. Due to the cloud platform can provide sufficient computation and storage resources, the exhaustive analysis of the massive historical data through a model training process is supported by cloud service. The well-trained model can be further transfer to the edge computing layer to implement the local microgrid energy management functionalities.</p>
<p>Specifically, based on the physical architecture of microgrid energy management given above, next we describe the theoretical model of energy management. Note that the agent selects an action under the state and the environment gets the next state in the microgrid scenario, and that the next state of the agent only depends on the current state and the action, and it is not related to previous states and actions. Consequently, microgrid energy management can be formulated as MDP problem. In general, the state space <italic>S</italic>, action space <italic>A</italic>, and reward function R are used for the agent-environment interaction modeling of the MDP. We define that the state space includes exogenous state component and controllable state component, and the action space includes the energy deficiency action, price action, and TCL action. After the agent transfers from state S<sub>
<italic>t</italic>
</sub> to state s<sub>
<italic>t</italic>&#x2b;1</sub>, it received the immediate reward <inline-formula id="inf1">
<mml:math id="m1">
<mml:mrow>
<mml:msub>
<mml:mi>R</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> when an action &#x3b1; is given. With an objective of minimizing the daily operating cost, the reward function, as gross margin from operations, is given by:<disp-formula id="e1">
<mml:math id="m2">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">R</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">Re</mml:mi>
<mml:msub>
<mml:mi mathvariant="bold-italic">v</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="bold-italic">Co</mml:mi>
<mml:msub>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>where <inline-formula id="inf2">
<mml:math id="m3">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>e</mml:mi>
<mml:msub>
<mml:mi>v</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msub>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:msubsup>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>C</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:msubsup>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>w</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mi>t</mml:mi>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mi>t</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:msubsup>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mi>t</mml:mi>
</mml:msubsup>
<mml:mtext>&#xa0;</mml:mtext>
</mml:mrow>
</mml:math>
</inline-formula> is the microgrid revenues from selling electricity to the external grid, <inline-formula id="inf3">
<mml:math id="m4">
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mi>o</mml:mi>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>G</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mi>t</mml:mi>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mi>t</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:msubsup>
<mml:mi>E</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>P</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> is the costs related to purchases from the power generation and external grid. <inline-formula id="inf4">
<mml:math id="m5">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the price of the price-responsive loads, <inline-formula id="inf5">
<mml:math id="m6">
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the number of the price-responsive load. <inline-formula id="inf6">
<mml:math id="m7">
<mml:mrow>
<mml:msubsup>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mi>b</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>S</mml:mi>
<mml:msubsup>
<mml:mi>L</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>P</mml:mi>
<mml:msubsup>
<mml:mi>B</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> represents the power consumption of the price-responsive loads at time <italic>t</italic>. <inline-formula id="inf7">
<mml:math id="m8">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the price of the direct controllable loads, <inline-formula id="inf8">
<mml:math id="m9">
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>C</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the number of the direct controllable loads. <inline-formula id="inf9">
<mml:math id="m10">
<mml:mrow>
<mml:msubsup>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> represents the power consumption of the direct controllable loads at time <italic>t</italic>, it can be calculated by <inline-formula id="inf10">
<mml:math id="m11">
<mml:mrow>
<mml:msubsup>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mi>u</mml:mi>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>. <inline-formula id="inf11">
<mml:math id="m12">
<mml:mrow>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mi>t</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf12">
<mml:math id="m13">
<mml:mrow>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mi>t</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> are the power transmission costs respectively for exporting to and importing from the external grid. <inline-formula id="inf13">
<mml:math id="m14">
<mml:mrow>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mi>t</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> is the up-regulation price and <inline-formula id="inf14">
<mml:math id="m15">
<mml:mrow>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>w</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mi>t</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> is down-regulation price. The power generation cost is <inline-formula id="inf15">
<mml:math id="m16">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. The energies purchased, sold to, and generated from the external grid are <inline-formula id="inf16">
<mml:math id="m17">
<mml:mrow>
<mml:msubsup>
<mml:mi>E</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>P</mml:mi>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mi>t</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf17">
<mml:math id="m18">
<mml:mrow>
<mml:msub>
<mml:mi>G</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> respectively.</p>
</sec>
<sec id="s3">
<title>UCB-A3C Based Energy Management</title>
<p>In view of the fact that the dimension of state space and action space in microgrid are large. To solve this MDP problem, the A3C learning algorithm, which is a state-of-the-art actor-critic method that exploits multi-threading to create several learning agents, is effective to handle the large scale decisions-making problem and is considered in this&#x20;paper.</p>
<sec id="s3-1">
<title>A3C Algorithm</title>
<p>Different from the classical actor-critic algorithm with only one learning agent, the A3C method adopts asynchronously parallel learning of multiple actor on different threads. The key advantage of this parallel learn scheme in different threads is that it breaks the interdependence of gradient updates and decorrelates past experiences gained by each learning agent, and it is an online learning algorithm and converges rapidly (<xref ref-type="bibr" rid="B6">Jia et al., 2015</xref>; <xref ref-type="bibr" rid="B11">Liu et al., 2019</xref>; <xref ref-type="bibr" rid="B8">Lee et al., 2020</xref>).</p>
<p>In A3C algorithm with a multi-threaded training framework, it has one global network consisting of actor network and critical network. These two neural networks have different function. To be specific, the policy gradient schemes is utilized by actor network to choose the action, and the parameterized policy with a set of actor parameters <italic>&#x3b8;</italic>
<sub>
<italic>a</italic>
</sub> is defined by <italic>&#x3c0;</italic>(<italic>a&#x7c;s</italic>; <italic>&#x3b8;</italic>
<sub>
<italic>a</italic>
</sub>) &#x3d; <italic>P</italic>(<italic>a&#x7c;s,&#x3b8;</italic>
<sub>
<italic>a</italic>
</sub>), and the gradient-descent method is applied to update the parameters. The critic network evaluates each action from the actor network and learns the value function while multiple actors are trained in parallel. And in order to qualify the expected reward, the critic network estimates the state-value function <italic>V</italic>(<italic>s</italic>
<sub>
<italic>t</italic>
</sub>;<italic>&#x3b8;</italic>
<sub>
<italic>c</italic>
</sub>) on account of state <italic>s</italic> with critic parameters <italic>&#x3b8;</italic>
<sub>c</sub>.</p>
<p>During algorithm execution, each agent makes use of the value function to evaluate its policy to achieve the long-term cumulative reward. For the given policy <italic>&#x3c0;</italic>, the state-action value function, called <italic>Q</italic>-function, of state action pair (<italic>s</italic>, <italic>a</italic>) can be achieved by action <italic>a</italic>, it is defined as the expected reward by an action <italic>a</italic> in the state <italic>s</italic>,<disp-formula id="e2">
<mml:math id="m19">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">Q</mml:mi>
<mml:mi mathvariant="bold-italic">&#x3c0;</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="bold-italic">a</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">R</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mtext>&#x7c;</mml:mtext>
<mml:msub>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">a</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">a</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>and the state value function<disp-formula id="e3">
<mml:math id="m20">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">V</mml:mi>
<mml:mi mathvariant="bold-italic">&#x3c0;</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="double-struck">E</mml:mi>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">R</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mtext>&#x7c;</mml:mtext>
<mml:msub>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">s</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>where the expectation <inline-formula id="inf18">
<mml:math id="m21">
<mml:mi mathvariant="italic">&#x395;</mml:mi>
</mml:math>
</inline-formula> (&#xb7;) is taken over all possible the state-action transitions following the policy <italic>&#x3c0;</italic>. For the policy and the value function A3C learning algorithm, the parameters are updated by the <italic>n</italic>-step reward and the reward is defined as:<disp-formula id="e4">
<mml:math id="m22">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">r</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b3;</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
<mml:mi mathvariant="bold-italic">r</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">a</mml:mi>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b3;</mml:mi>
<mml:mi mathvariant="bold-italic">n</mml:mi>
</mml:msub>
<mml:mi mathvariant="bold-italic">V</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:mi mathvariant="bold-italic">c</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>where <italic>&#x3b3;</italic> is the discount factor.</p>
<p>For update rule of A3C, it is desire to the agent not only learns how good the action is, but also learn how much better than expected. And the policy gradient scheme is adopted to perform parameters update for the A3C algorithm. However, high variance may be introduced by the policy gradient in the critic network. To deal with the problem, the <italic>Q</italic>(<italic>s</italic>, <italic>a</italic>) function in the policy gradient process has been replaced by the advantage function <inline-formula id="inf19">
<mml:math id="m23">
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, and <inline-formula id="inf20">
<mml:math id="m24">
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is given by:<disp-formula id="e5">
<mml:math id="m25">
<mml:mrow>
<mml:mi mathvariant="bold-italic">A</mml:mi>
<mml:mi mathvariant="bold-italic">(s</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="bold-italic">t)</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">r</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi mathvariant="bold-italic">t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="bold-italic">V</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">R</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3b3;</mml:mi>
<mml:msub>
<mml:mi mathvariant="bold-italic">R</mml:mi>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b3;</mml:mi>
<mml:mrow>
<mml:mi mathvariant="bold-italic">n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi mathvariant="bold-italic">R</mml:mi>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b3;</mml:mi>
<mml:mi mathvariant="bold-italic">n</mml:mi>
</mml:msub>
<mml:mi mathvariant="bold-italic">V</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="bold-italic">V</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>where <inline-formula id="inf21">
<mml:math id="m26">
<mml:mrow>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf22">
<mml:math id="m27">
<mml:mrow>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> are the sate value function in the state <inline-formula id="inf23">
<mml:math id="m28">
<mml:mrow>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf24">
<mml:math id="m29">
<mml:mi>s</mml:mi>
</mml:math>
</inline-formula>, respectively.</p>
<p>In the A3C learning framework, there are two loss functions and they are associated with the outputs of deep neural network, and all the actor-learners update the state value function <inline-formula id="inf25">
<mml:math id="m30">
<mml:mrow>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>s</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> and the policy <italic>&#x3c0;</italic>(<italic>s</italic>, <italic>a</italic>) by the gradient loss. The actor loss function is defined as:<disp-formula id="e6">
<mml:math id="m31">
<mml:mrow>
<mml:mi mathvariant="bold-italic">L</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:mi mathvariant="bold-italic">a</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">log&#x3c0;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">a</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mtext>&#x7c;</mml:mtext>
<mml:msub>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:mi mathvariant="bold-italic">a</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">r</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="bold-italic">V</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:mi mathvariant="bold-italic">c</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3b8;G</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3c0;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:mi mathvariant="bold-italic">a</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>where <inline-formula id="inf26">
<mml:math id="m32">
<mml:mi>&#x3b8;</mml:mi>
</mml:math>
</inline-formula> is the hyperparameter and <inline-formula id="inf27">
<mml:math id="m33">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>&#x3c0;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>a</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the entropy which is used to encourage exploration and discourage premature convergence to a suboptimal policy. The accumulated gradient of the <inline-formula id="inf28">
<mml:math id="m34">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>a</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is expressed as:<disp-formula id="e7">
<mml:math id="m35">
<mml:mrow>
<mml:mi mathvariant="bold-italic">d</mml:mi>
<mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:mi mathvariant="bold-italic">a</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">d</mml:mi>
<mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:mi mathvariant="bold-italic">a</mml:mi>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mo>&#x2207;</mml:mo>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
<mml:msub>
<mml:mi mathvariant="normal">&#x2032;</mml:mi>
<mml:mi>a</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2061;</mml:mo>
<mml:mi mathvariant="italic">log</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>&#x3c0;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">a</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:msub>
<mml:mi mathvariant="normal">&#x2032;</mml:mi>
<mml:mi mathvariant="bold-italic">a</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi mathvariant="bold-italic">A(s</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="bold-italic">t)</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:msub>
<mml:mo>&#x2207;</mml:mo>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
<mml:msub>
<mml:mi mathvariant="normal">&#x2032;</mml:mi>
<mml:mi>a</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="bold-italic">G</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3c0;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:msub>
<mml:mi mathvariant="normal">&#x2032;</mml:mi>
<mml:mi mathvariant="bold-italic">a</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>where <inline-formula id="inf29">
<mml:math id="m36">
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>a</mml:mi>
<mml:mtext>&#x27;</mml:mtext>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> is the thread-specific parameter in actor network.</p>
<p>Similarly, the loss function in critic network is given by:<disp-formula id="e8">
<mml:math id="m37">
<mml:mrow>
<mml:mi mathvariant="bold-italic">L</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:mi mathvariant="bold-italic">c</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">r</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="bold-italic">V</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:mi mathvariant="bold-italic">c</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
<label>(8)</label>
</disp-formula>and the accumulated gradient of <inline-formula id="inf30">
<mml:math id="m38">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is defined as:<disp-formula id="e9">
<mml:math id="m39">
<mml:mrow>
<mml:mi mathvariant="bold-italic">d</mml:mi>
<mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:mi mathvariant="bold-italic">c</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">d</mml:mi>
<mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:mi mathvariant="bold-italic">c</mml:mi>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mo>&#x2207;</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:mi mathvariant="bold-italic">c</mml:mi>
<mml:mi mathvariant="bold-italic">&#x2032;</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:msub>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">r</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="bold-italic">V</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">s</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:mi mathvariant="bold-italic">c</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
<label>(9)</label>
</disp-formula>where <inline-formula id="inf31">
<mml:math id="m40">
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>c</mml:mi>
<mml:mtext>&#x27;</mml:mtext>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> is the thread-specific parameter in critic network. In order to achieve the loss function minimization in our presented A3C framework, the standard noncentered RMSProp algorithm (<xref ref-type="bibr" rid="B16">Tijmen and Geoffrey, 2012</xref>) is utilized to perform training until the accumulated gradients shown in <xref ref-type="disp-formula" rid="e7">Eqs 7</xref>, <xref ref-type="disp-formula" rid="e9">9</xref> is updated.</p>
<p>Consider that the sufficient exploration is needed to avoid a suboptimal policy with worse reward and the exploitation adopts the policy with the best reward, the optimal learning strategy, which can implement the balance between exploration and exploitation, is expected to be achieved.</p>
</sec>
<sec id="s3-2">
<title>Proposed UCB-A3C Algorithm</title>
<p>To further improve the performance of the A3C algorithm, this paper is leveraging the idea of the UCB algorithm. Firstly, the agent selects an action through UCB exploration and executes it. Then, after sampling from the experience pool and calculating the loss function, the priority of action is calculated. Finally, the priority will be assigned to the action to be performed. The priority of action is calculated by the following equation:<disp-formula id="e10">
<mml:math id="m41">
<mml:mrow>
<mml:mi mathvariant="bold-italic">p</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">acts</mml:mi>
<mml:mo>_</mml:mo>
<mml:mi mathvariant="bold-italic">prob</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3c4;</mml:mi>
<mml:msqrt>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mi>ln</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3b5;</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3c3;</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">N</mml:mi>
<mml:mi mathvariant="bold-italic">j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:math>
<label>(10)</label>
</disp-formula>where <italic>N</italic>
<sub>
<italic>j</italic>
</sub> represents the number of the <italic>j</italic>th action is selected, <inline-formula id="inf32">
<mml:math id="m42">
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo>_</mml:mo>
<mml:mi>p</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is the probability value returned by the actor network output, <inline-formula id="inf33">
<mml:math id="m43">
<mml:mi>&#x3c4;</mml:mi>
</mml:math>
</inline-formula> is the parameter that adjusts the influence of priority on its selection action, <italic>&#x3b5;</italic> is the parameter that keeps decreasing, and <italic>&#x3c3;</italic> is the parameter for <italic>&#x3b5;</italic> correction. The second term in <xref ref-type="disp-formula" rid="e10">Eq. 10</xref> is the confidence factor. In the initial stage of the algorithm, the confidence factor is large, so it has a great impact on the priority. With the progress of training, the time step <italic>t</italic> continues to increase, and the influence of confidence factor will gradually decrease, so as to increase the chance of relative attempt and ensure the diversity of samples. At time <italic>t</italic>, if an action has been selected many times, the higher the reward value of the action, the greater the probability of being continued. If an action is selected few times, its confidence factor will be higher and the probability of being continued will be lower. When the algorithm reaches the convergence state, the benefit of optimal action selection can be maximized. As mentioned above, the framework of the UCB-A3C algorithm is shown in <xref ref-type="fig" rid="F1">Figure&#x20;1C</xref>.</p>
<p>Besides, we also carefully design the neural network (NN) structure of the UCB-A3C algorithm. The input layer of NN is composed of 107 neurons, corresponding to the input 107 environmental states. The hidden layer is designed as a combination of convolutional layer, pooled layer and fully connected layer. After the data is input through the input layer, the data is convolved through a convolution layer. The convolution layer adopts a 3&#x20;<inline-formula id="inf34">
<mml:math id="m44">
<mml:mo>&#xd7;</mml:mo>
</mml:math>
</inline-formula> 3 convolution kernel. After output data from the convolution layer, the global average pooling layer is used for data pooling. Then the data is output to actor and critic network through the fully connected layer of two layers with the number of neurons being 200 and 100, respectively. The actor network is designed as a fully connected layer with the number of neurons being 80, while the critic network is designed as a fully connected layer with the number of neurons being 1. The structure of the neural network is shown in <xref ref-type="fig" rid="F1">Figure&#x20;1D</xref>.</p>
<p>Consequently, the details of our proposed UCB-A3C algorithm is described in <xref ref-type="table" rid="T1">Table&#x20;1</xref>, and the flowchart of the proposed algorithm is shown in <xref ref-type="fig" rid="F1">Figure&#x20;1E</xref>.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>The details of UCB-A3C algorithm is described.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Algorithm improvement A3C</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Input: state values of each component of the microgrid</td>
</tr>
<tr>
<td align="left">Output: action of each component of the microgrid</td>
</tr>
<tr>
<td align="left">Initialization: discount factor <italic>&#x3bc;</italic>, parameters of global A3C neural network <inline-formula id="inf35">
<mml:math id="m45">
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>&#x3c9;</mml:mi>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> parameters of current thread neural network <inline-formula id="inf36">
<mml:math id="m46">
<mml:mrow>
<mml:msup>
<mml:mi>&#x3b8;</mml:mi>
<mml:mo>&#x27;</mml:mo>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mi>&#x3c9;</mml:mi>
<mml:mo>&#x27;</mml:mo>
</mml:msup>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> the number of samples selected for training is <italic>d</italic>, the number of iteration rounds globally shared <italic>T</italic>, the maximum number of iteration rounds globally shared is <italic>T</italic>
<sub>max</sub>, initial time <italic>t</italic>
<sub>
<italic>start</italic>
</sub>
</td>
</tr>
<tr>
<td align="left">1: Reset the gradient update amount of public neural network, reset <inline-formula id="inf37">
<mml:math id="m47">
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> &#x3d; 0, <inline-formula id="inf38">
<mml:math id="m48">
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>&#x3c9;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> &#x3d; 0</td>
</tr>
<tr>
<td align="left">2: Update the parameters of the current thread neural network <inline-formula id="inf39">
<mml:math id="m49">
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mi>&#x3b8;</mml:mi>
<mml:mo>&#x27;</mml:mo>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:mi>&#x3c9;</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mi>&#x3c9;</mml:mi>
<mml:mo>&#x27;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td align="left">3: Observe the current system state <italic>s</italic>
<sub>
<italic>t</italic>
</sub>
</td>
</tr>
<tr>
<td align="left">4: Select action <inline-formula id="inf40">
<mml:math id="m50">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> base on strategy <inline-formula id="inf41">
<mml:math id="m51">
<mml:mrow>
<mml:mi>&#x3c0;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x7c;</mml:mo>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>&#x3b8;</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td align="left">5: Calculate the reward value <inline-formula id="inf42">
<mml:math id="m52">
<mml:mrow>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> at the current time <italic>t</italic> and observe the state <italic>s</italic>
<sub>
<italic>t&#x2b;1</italic>
</sub>&#xa0;at the next time</td>
</tr>
<tr>
<td align="left">6: Store the resulting quaterple (<italic>s,a,r,s</italic>&#x27;) in experience pool <italic>D</italic>
</td>
</tr>
<tr>
<td align="left">7: If the experience pool is full, take a batch of samples <italic>d</italic> from the experience pool <italic>D</italic> to train the network</td>
</tr>
<tr>
<td align="left">8: Calculate the priority of the selected action <inline-formula id="inf43">
<mml:math id="m53">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo>_</mml:mo>
<mml:mi>p</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>b</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3c4;</mml:mi>
<mml:msqrt>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mi>ln</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>&#x3b5;</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td align="left">9: <italic>N</italic>
<sub>
<italic>j</italic>
</sub> <italic>&#x3d; N</italic>
<sub>
<italic>j</italic>
</sub>
<italic>&#x2b;1</italic>
</td>
</tr>
<tr>
<td align="left">10: Select the next action moment <inline-formula id="inf44">
<mml:math id="m54">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td align="left">11: <italic>t&#x2190;t&#x2b;1,T&#x2190;T&#x2b;1</italic>
</td>
</tr>
<tr>
<td align="left">12: Determine whether the current state of <italic>s</italic>
<sub>
<italic>t</italic>
</sub> is a terminated state, if not, return to step 5</td>
</tr>
<tr>
<td align="left">13: Calculate <italic>Q</italic>(<italic>s</italic>
<sub>
<italic>t</italic>
</sub>
<italic>,t</italic>) of the last time series position state <italic>s</italic>
<sub>
<italic>t</italic>
</sub>
</td>
</tr>
<tr>
<td align="left">14: <italic>For</italic> <inline-formula id="inf45">
<mml:math id="m55">
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>3</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:msub>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td align="left">15: Calculate <italic>Q</italic>(<italic>s</italic>
<sub>
<italic>i</italic>
</sub>
<italic>,i</italic>) of the state <italic>s</italic>
<sub>
<italic>i</italic>
</sub> corresponding to the current time <italic>t</italic>
</td>
</tr>
<tr>
<td align="left">16: Update the local gradient <inline-formula id="inf46">
<mml:math id="m56">
<mml:mrow>
<mml:msup>
<mml:mi>&#x3b8;</mml:mi>
<mml:mo>&#x27;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> of the current thread</td>
</tr>
<tr>
<td align="left">17: Update the local gradient <inline-formula id="inf47">
<mml:math id="m57">
<mml:mrow>
<mml:msup>
<mml:mi>&#x3c9;</mml:mi>
<mml:mo>&#x27;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> of the current thread</td>
</tr>
<tr>
<td align="left">18: end for</td>
</tr>
<tr>
<td align="left">19: Update the neural network parameters <inline-formula id="inf48">
<mml:math id="m58">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>&#x3c9;</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td align="left">20: Until <italic>T</italic>&#x20;&#x3e; <italic>T</italic>
<sub>max</sub>, otherwise, return to step 3</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Clearly, the proposed algorithm indicates that the probability of the selected action with larger reward is effectively increased with the help of&#x20;UCB.</p>
</sec>
<sec id="s3-3">
<title>Implementation of Microgrid Energy Management</title>
<p>UCB-A3C based microgrid energy management consists of offline and online stages. The offline stage stores the action records of microgrid operation and historical events, improves and perfects the data set through historical data accumulation, performs the offline simulation of microgrid operation control to train the agent, and updates the agent model and parameters for the use of online agents. When the microgrid works in real time in the online stage, the agent calculates the output action and control command according to the state variables and rewards fed back by the microgrid. Moreover, in line with the control command from cloud server, the microgrid operates and feeds back the updated status and reward to the online agent, and stores it in the edge computing platform. The implementation of microgrid energy management strategy based on UCB-A3C is shown in <xref ref-type="fig" rid="F1">Figure&#x20;1F</xref>.</p>
</sec>
</sec>
<sec id="s4">
<title>Performance Evaluation</title>
<p>In order to verify the proposed microgrid energy management strategy based on UCB-A3C, a simulation model, as shown in <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>, is a built-in Python environment, and the simulation parameters are shown in <xref ref-type="table" rid="T2">Table&#x20;2</xref>.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Analysis of the UCB-A3C algorithm, <bold>(A)</bold> is the reward of the UCB-A3C algorithm, <bold>(B)</bold> is the total economic profits, <bold>(C)</bold> is the daily economic profits.</p>
</caption>
<graphic xlink:href="fenrg-10-858895-g002.tif"/>
</fig>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>The simulation parameters.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Parameter</th>
<th align="center">Numerical</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Maximum capacity of ESS</td>
<td align="center">500&#xa0;KWh</td>
</tr>
<tr>
<td align="left">Charging power of ESS</td>
<td align="center">250&#xa0;KW</td>
</tr>
<tr>
<td align="left">Discharge power of ESS</td>
<td align="center">250&#xa0;KW</td>
</tr>
<tr>
<td align="left">Generation cost of DER</td>
<td align="center">32&#x20ac;/MW</td>
</tr>
<tr>
<td align="left">Generation capacity of DER</td>
<td align="center">data source [<xref ref-type="bibr" rid="B5">Oy, 2018</xref>]</td>
</tr>
<tr>
<td align="left">Number of directly controllable loads</td>
<td align="center">100</td>
</tr>
<tr>
<td align="left">The quantity of non-directly controllable loads</td>
<td align="center">150</td>
</tr>
<tr>
<td align="left">Electricity markets cut prices</td>
<td align="center">data source [<xref ref-type="bibr" rid="B4">Fingrid Open Datasets., 2018</xref>]</td>
</tr>
<tr>
<td align="left">Electricity markets raised prices</td>
<td align="center">data source [<xref ref-type="bibr" rid="B4">Fingrid Open Datasets., 2018</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In our simulation, PPO, Actor Critic, A3C and UCB-A3C algorithm are respectively used for training, and the obtained reward value is shown in <xref ref-type="fig" rid="F2">Figure&#x20;2A</xref>. It can be observed that the UCB-A3C algorithm has higher reward in the learning process than other algorithms.</p>
<p>And that, the PPO, Actor-Critic, A3C and improved A3C algorithm were respectively used to carry out the economic profits of microgrid energy control, and the achieved total 10&#xa0;days&#x2019; economic profits is shown in <xref ref-type="fig" rid="F2">Figure&#x20;2B</xref>. It can be found that the economic profits obtained by the UCB-A3C algorithm based microgrid energy management are greater than those obtained by the other three algorithms.</p>
<p>Simultaneously, the daily economic profits of the A3C algorithm and the UCB-A3C algorithm for ten consecutive days are compared, as shown in <xref ref-type="fig" rid="F2">Figure&#x20;2C</xref>. It can be seen from the figure the UCB-A3C algorithm is superior to the A3C algorithm in six of the 10&#xa0;days of revenue, which has better energy management efficiency.</p>
<p>Further, we make use of UCB-A3C algorithm to optimize microgrid energy management, and the predicted data of power generation and consumption of wind power generation components and power load components are shown in <xref ref-type="fig" rid="F3">Figure&#x20;3A</xref>. At this time, the energy storage system is charged from the 0<sup>th</sup> to 1st hours and discharged from the 17th to 21st hours in <xref ref-type="fig" rid="F3">Figure&#x20;3B</xref>. In the energy trading market, electricity is mainly sold, and the trading price changes with the trading electricity is shown in <xref ref-type="fig" rid="F3">Figure&#x20;3C</xref>.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>The UCB-A3C algorithm is utilized to implement&#x20;the&#x20;microgrid energy management, <bold>(A)</bold> is the electric loads and&#x20;distributed&#x20;energy resource, <bold>(B)</bold> is the energy storage system,&#x20;<bold>(C)</bold>&#x20;is&#x20;the&#x20;electricity transaction volume and price change.</p>
</caption>
<graphic xlink:href="fenrg-10-858895-g003.tif"/>
</fig>
<p>To sum up, the improved A3C algorithm can carry out efficient energy coordination management on the microgrid, and then efficiently trade electricity with the power grid, so as to achieve the purpose of reasonable distribution of electricity, improve economic profits, and reduce the power loss in the process of power distribution.</p>
</sec>
<sec sec-type="conclusion" id="s5">
<title>Conclusion</title>
<p>Microgrids are an effective way to deal with flexible access to renewable energy and varying power loads. In order to deal with such volatility and uncertainty, this paper proposes an A3C algorithm based on UCB exploration mechanism. According to the simulation, we have validated that the proposed UCB-A3C algorithm can adapt the constantly changing microgrid environment, learn the efficient energy management strategy, and provide a more economical scheme for the microgrid operation, so as to achieve the purpose of reducing the economic cost. Moreover, UCB-A3C learning algorithm based microgrid energy management has solved the unscalable application and repeated development problems of traditional domain experts. However, in practical application, there are still some areas that need to be further improved due to its long training time and great dependence on training data in the learning process. Therefore, it is the focus of future research to solve the above problems in order to better apply deep reinforcement learning to microgrid energy management.</p>
</sec>
</body>
<back>
<sec id="s6">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>YY is responsible for algorithm research. HL is responsible for algorithm research. BS is responsible for simulation. WP is responsible for algorithm design. DP is responsible for writing paper.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This work is supported by the National Natural Science Foundation of China (No. U2066211) and the Youth Innovation Promotion Association, Chinese Academy of Sciences (No. 2021136).</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Guan</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Multi-objective Optimal Dispatch Model and its Algorithm in Isolated Microgrid with Renewable Energy Generation as Main Power Supply</article-title>. <source>Power Syst. Technol.</source>. </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dong</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Zomaya</surname>
<given-names>A. Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Machine-learning-based Real-Time Economic Dispatch in Islanding Microgrids in a Cloud-Edge Computing Environment</article-title>. <source>IEEE Internet Things J.</source> <volume>8</volume>, <fpage>13703</fpage>&#x2013;<lpage>13711</lpage>. <pub-id pub-id-type="doi">10.1109/JIOT.2021.3067951</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fan</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Optimal Scheduling of Microgrid Based on Deep Deterministic Policy Gradient and Transfer Learning</article-title>. <source>Energies</source> <volume>14</volume>, <fpage>584</fpage>. <pub-id pub-id-type="doi">10.3390/en14030584</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="web">
<collab>Fingrid Open Datasets</collab> (<year>2018</year>) <article-title>Fingrid Open Datasets</article-title>. <comment>Available: <ext-link ext-link-type="uri" xlink:href="https://data.fingrid.fi/open-dataforms">https://data.fingrid.fi/open-dataforms</ext-link>
</comment>. </citation>
</ref>
<ref id="B5">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Oy</surname>
<given-names>Fortum</given-names>
</name>
</person-group> (<year>2018</year>). <source>Wind Farm Data</source>. <publisher-loc>Finland</publisher-loc>. </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jia</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>X.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Research on Some Key Problems Related to Integrated Energy Systems</article-title>. <source>Automation Electric Power Syst.</source>
<pub-id pub-id-type="doi">10.7500/AEPS20141009011</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kim</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Oh</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Choi</surname>
<given-names>J.&#x20;K.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Learning Based Cost Optimal Energy Management Model for Campus Microgrid Systems</article-title>. <source>Appl. Energ.</source> <volume>311</volume>, <fpage>118630</fpage>. <pub-id pub-id-type="doi">10.1016/j.apenergy.2022.118630</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Niyato</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Demand-Side Scheduling Based on Multi-Agent Deep Actor-Critic Learning for Smart Grids</article-title>. <source>Early Access</source>. <pub-id pub-id-type="doi">10.1109/SmartGridComm47815.2020.9302935</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Choi</surname>
<given-names>D.-H.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Federated Reinforcement Learning for Energy Management of Multiple Smart Homes with Distributed Energy Resources</article-title>. <source>IEEE Trans. Ind. Inf.</source> <volume>18</volume>, <fpage>488</fpage>&#x2013;<lpage>497</lpage>. <pub-id pub-id-type="doi">10.1109/TII.2020.3035451</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wan</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Real-time Residential Demand Response</article-title>. <source>IEEE Trans. Smart Grid</source> <volume>11</volume>, <fpage>4144</fpage>&#x2013;<lpage>4154</lpage>. <pub-id pub-id-type="doi">10.1109/TSG.2020.2978061</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Intelligent Edge Computing for IoT-Based Energy Management in Smart Cities</article-title>. <source>IEEE Netw.</source> <volume>33</volume>, <fpage>111</fpage>&#x2013;<lpage>117</lpage>. <pub-id pub-id-type="doi">10.1109/MNET.2019.1800254</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ming</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Rui</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Zha</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tao</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Multi-Objective Optimization of Hybrid Renewable Energy System with Load Forecasting</article-title>,&#x201d; in <conf-name>2017 IEEE International Conference on Energy Internet (ICEI)</conf-name>, <conf-loc>Beijing, China</conf-loc>, <conf-date>17-21 April 2017</conf-date>. <pub-id pub-id-type="doi">10.1109/icei.2017.27</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Mnih</surname>
<given-names>Volodymyr.</given-names>
</name>
<name>
<surname>Badia</surname>
<given-names>A. P.</given-names>
</name>
<name>
<surname>Mirza</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Graves</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Lillicrap</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Harley</surname>
<given-names>T.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). &#x201c;<article-title>Asynchronous Methods for Deep Reinforcement Learning</article-title>,&#x201d; in <conf-name>International Conference on Machine Learning, ICML 2016</conf-name>. </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nakabi</surname>
<given-names>T. A.</given-names>
</name>
<name>
<surname>Toivanen</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Deep Reinforcement Learning for Energy Management in a Microgrid with Flexible Demand</article-title>. <source>Sustainable&#x20;Energ. Grids Networks</source> <volume>25</volume>, <fpage>100413</fpage>. <pub-id pub-id-type="doi">10.1016/j.segan.2020.100413</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pourmousavi</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Nehrir</surname>
<given-names>M. H.</given-names>
</name>
<name>
<surname>Colson</surname>
<given-names>C. M.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Real-Time Energy Management of a Stand-Alone Hybrid Wind-Microturbine Energy System Using Particle Swarm Optimization</article-title>. <source>IEEE&#x20;Trans.&#x20;Sustain.&#x20;Energ.</source> <volume>1</volume>, <fpage>193</fpage>&#x2013;<lpage>201</lpage>. <pub-id pub-id-type="doi">10.13335/j.1000-3673.pst.2019.0298</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tijmen</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Geoffrey</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Lecture 6.5-rmsprop: Divide the Gradient by a Running Average of its Recent Magnitude</article-title>. <source>Coursera: Neural Networks Machine Learn.</source> <volume>4</volume>. </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Pei</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Huo</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Coordinated Planning Method of Multiple Micro-grids and Distribution Network with Flexible Interconnection</article-title>. <source>Appl. Energ.</source> <volume>228</volume>, <fpage>2361</fpage>&#x2013;<lpage>2374</lpage>. <pub-id pub-id-type="doi">10.1016/j.apenergy.2018.07.047</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>Z.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Deep&#x20;Reinforcement Learning for Smart home Energy Management</article-title>.&#x20;<source>IEEE Internet Things J.</source> <volume>7</volume>, <fpage>2751</fpage>&#x2013;<lpage>2762</lpage>. <pub-id pub-id-type="doi">10.1109/JIOT.2019.2957289</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Qiu</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>A Coordinated Control Method for Hybrid Energy Storage System in Microgrid Based on Deep Reinforcement Learning</article-title>. <source>Power Syst. Technol.</source> </citation>
</ref>
</ref-list>
</back>
</article>