<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="brief-report" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Energy Res.</journal-id>
<journal-title>Frontiers in Energy Research</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Energy Res.</abbrev-journal-title>
<issn pub-type="epub">2296-598X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">741101</article-id>
<article-id pub-id-type="doi">10.3389/fenrg.2021.741101</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Energy Research</subject>
<subj-group>
<subject>Brief Research Report</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Distributed Imitation-Orientated Deep Reinforcement Learning Method for Optimal PEMFC Output Voltage Control</article-title>
<alt-title alt-title-type="left-running-head">Li et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">DDRL Method for PEMFC Control</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Li</surname>
<given-names>Jiawen</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1256910/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Li</surname>
<given-names>Yaping</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Yu</surname>
<given-names>Tao</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>College of Electric Power, South China University of Technology, <addr-line>Guangzhou</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>China Electric Power Research Institute (Nanjing), <addr-line>Beijing</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1236518/overview">Yaxing Ren</ext-link>, University of Warwick, United&#x20;Kingdom</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1408940/overview">Lefeng Cheng</ext-link>, Guangzhou University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1309900/overview">Si Chen</ext-link>, University of Glasgow, United&#x20;Kingdom</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Tao Yu, <email>taoyu1@scut.edu.cn</email>, <email>eptaoyu1@163.com</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Smart Grids, a section of the journal Frontiers in Energy Research</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>01</day>
<month>10</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>9</volume>
<elocation-id>741101</elocation-id>
<history>
<date date-type="received">
<day>14</day>
<month>07</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>27</day>
<month>07</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Li, Li and Yu.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Li, Li and Yu</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>In order to improve the stability of proton exchange membrane fuel cell (PEMFC) output voltage, a data-driven output voltage control strategy based on regulation of the duty cycle of the DC-DC converter is proposed in this paper. In detail, an imitation-oriented twin delay deep deterministic (IO-TD3) policy gradient algorithm which offers a more robust voltage control strategy is demonstrated. This proposed output voltage control method is a distributed deep reinforcement learning training framework, the design of which is guided by the pedagogic concept of imitation learning. The effectiveness of the proposed control strategy is experimentally demonstrated.</p>
</abstract>
<kwd-group>
<kwd>distributed deep reinforcement learning</kwd>
<kwd>proton exchange membrane fuel cell</kwd>
<kwd>DC-DC converter</kwd>
<kwd>output voltage control</kwd>
<kwd>robustness</kwd>
</kwd-group>
<contract-num rid="cn001">U2066212</contract-num>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/100014717</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>The voltage of a proton exchange membrane fuel cell (PEMFC) is highly dependent on the temperature, pressure, humidity, and gas flow rate (<xref ref-type="bibr" rid="B19">Yang et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B14">Sun et&#x20;al., 2019</xref>). In addition, the output voltage of PEMFC also fluctuates widely with varying load current (<xref ref-type="bibr" rid="B20">Yang et&#x20;al., 2019a</xref>; <xref ref-type="bibr" rid="B16">Yang et&#x20;al., 2021a</xref>). In order to improve the stability of the PEMFC output voltage, the PEMFC DC-DC converters should output a stable bus voltage in the event of fluctuating input voltage and output load so as to normalize the load (<xref ref-type="bibr" rid="B17">Yang et&#x20;al., 2021b</xref>; <xref ref-type="bibr" rid="B21">Yang et&#x20;al., 2021c</xref>).</p>
<p>There are a number of existing PEMFC output voltage control methods based on control of DC-DC converters, including the PID control algorithm (<xref ref-type="bibr" rid="B15">Swain and Jena, 2015</xref>), fractional order PID algorithm (<xref ref-type="bibr" rid="B20">Yang et&#x20;al., 2019a</xref>; <xref ref-type="bibr" rid="B22">Yang et&#x20;al., 2019b</xref>; <xref ref-type="bibr" rid="B18">Yang et&#x20;al., 2020</xref>), sliding mode control algorithm (<xref ref-type="bibr" rid="B2">Bougrine et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B5">Jiao and Cui, 2013</xref>), model predictive control algorithm (<xref ref-type="bibr" rid="B1">Bemporad et&#x20;al., 2002</xref>; <xref ref-type="bibr" rid="B3">Ferrari-Trecate et&#x20;al., 2002</xref>), robust control method (<xref ref-type="bibr" rid="B12">Olalla et&#x20;al., 2010</xref>), and optimal control algorithm (<xref ref-type="bibr" rid="B4">Jaen et&#x20;al., 2006</xref>; <xref ref-type="bibr" rid="B11">Olalla et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B9">Montagner et&#x20;al., 2011</xref>; <xref ref-type="bibr" rid="B10">Moreira et&#x20;al., 2011</xref>) methods, and so on. Among them, the PID algorithms are traditional control algorithms whose advantages include simple structure and fast calculation speed. However, these are incompatible with non-linear PEMFC systems. The fractional order PID algorithm is an expanded algorithm based on the PID algorithm, which offers better robustness, but which cannot be adapted for non-linear PEMFC systems. Sliding mode control is an excellent candidate for variable structure systems such as DC-DC converters; however, it is not suitable for PEMFC systems in practice as it is affected by the &#x201c;jitter&#x201d; problem. The model prediction algorithm offers higher accuracy and strong robustness; however, the algorithm is heavily reliant on mathematical models, making the control results in reality very different from the theoretical ones. The goal of robust control is to establish feedback control laws accounting for system uncertainty in order to increase the robustness of closed-loop systems. However, the control performance of a controller employing robust control is compromised it as cannot operate at the optimal&#x20;point.</p>
<p>Optimal control is one of the more advanced control algorithm designs. By expressing the performance of a system as an objective function of time, state, error, and other combinations, optimal control selects an appropriate control law which enables the objective function to include extreme values in order to obtain the optimal performance of the system. As described by <xref ref-type="bibr" rid="B4">Jaen et&#x20;al. (2006)</xref>, the average model of the converter is linearised, and the optimal LQR is obtained by solving the algebraic Riccati equation using the pole configuration, frequency domain metric or integral metric as the optimisation objective function; however, this LQR is not robust enough to cope with large disturbances in the system. <xref ref-type="bibr" rid="B9">Montagner et&#x20;al. (2011)</xref> designed a discrete LQR and determined the existence of the Lyapunov function for the closed-loop system using the LMI method, thus ensuring the stability of the system. <xref ref-type="bibr" rid="B11">Olalla et&#x20;al. (2009)</xref> organised the LQR optimisation problem in the form of an LMI, which was then solved using convex optimisation to obtain a robust linear quadratic regulator. In the study by <xref ref-type="bibr" rid="B10">Moreira et&#x20;al. (2011)</xref>, the application of a digital LQR with Kalman state observer for controlling a BUCK converter was tested in a series of simulations. However, the structure of the above optimal algorithm is complex and computationally intensive, leading to a reduction in its control real-time performance in practice (<xref ref-type="bibr" rid="B6">Li and Yu, 2021</xref>).</p>
<p>For these reasons, there remains the need for a simple structured model-free PMEFC optimal control algorithm for guiding DC-DC converters (<xref ref-type="bibr" rid="B7">Li et&#x20;al., 2021</xref>).</p>
<p>The DDPG algorithm (<xref ref-type="bibr" rid="B8">Lillicrap et&#x20;al., 2015</xref>) is a data-driven model-free optimal control algorithm, a kind of deep reinforcement learning, which is characterised by strong self-adaptive capability and decision-making ability, and which can arrive at decisions within a few milliseconds. It is used widely in power system control and robot coordination control, and for addressing UAV control problems (<xref ref-type="bibr" rid="B25">Zhang et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B13">Qi, 2018</xref>; <xref ref-type="bibr" rid="B27">Zhang et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B23">Zhang et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B26">Zhang and Yu, 2019</xref>; <xref ref-type="bibr" rid="B28">Zhang et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B24">Zhang et&#x20;al., 2021</xref>; <xref ref-type="bibr" rid="B24">Zhang et&#x20;al., 2021</xref>). However, the poor training efficiency of the DDPG algorithm explains the low robustness of controllers belonging to this class of algorithms, and their ineligibility for PEMFC systems.</p>
<p>In order to stabilise the output characteristics of the PEMFC and improve the stability of its output voltage, a data-driven output voltage control strategy for controlling the duty cycle of the DC-DC converter is proposed in this paper. To this end, an imitation-oriented twin delay deep deterministic policy gradient (IO-TD3) algorithm is proposed, the design of which reflects the idea of imitation learning. In this paper, we propose a distributed deep reinforcement learning training framework for improving the robustness of the PEMFC control policy. The effectiveness of the proposed control policy is experimentally demonstrated by comparing the proposed method with a number of existing algorithms.</p>
<p>This paper makes the following unique contributions to the research field:<list list-type="simple">
<list-item>
<p>1) A 75&#xa0;kw ninth order output voltage PEMFC dynamic control model that takes into account the DC/DC converter is demonstrated.</p>
</list-item>
<list-item>
<p>2) A PEMFC output voltage control strategy based on an imitation-oriented twin delay deep deterministic policy gradient algorithm for the purpose of increasing robustness is proposed.</p>
</list-item>
</list>
</p>
<p>The remainder of this paper comprises the following sections: the PEMFC model is demonstrated in <italic>The PEMFC Model</italic>, and the proposed algorithm is described in <italic>Proposed Method</italic>; the experimental results are analysed and discussed in <italic>Experiment</italic>, and the findings in this paper are summarised in <italic>Conclusion</italic>.</p>
</sec>
<sec id="s2">
<title>The PEMFC Model</title>
<sec id="s2-1">
<title>PEMFC Modelling and Characterization</title>
<p>A PEMFC is a device that converts chemical energy directly into electrical energy by means of an electrochemical reaction, the individual output voltage of which can be expressed as follows:<disp-formula id="e1">
<mml:math id="m1">
<mml:mrow>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mtext>cell&#xa0;</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>E</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>&#x3b7;</mml:mi>
<mml:mrow>
<mml:mtext>act&#xa0;</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>&#x3b7;</mml:mi>
<mml:mrow>
<mml:mtext>ohm&#xa0;</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>&#x3b7;</mml:mi>
<mml:mrow>
<mml:mtext>con&#xa0;</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>For a fuel cell stack consisting of <italic>N</italic> single cells connected in series, the output voltage <italic>V</italic> can be expressed as follows:<disp-formula id="e2">
<mml:math id="m2">
<mml:mrow>
<mml:mi>V</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>N</mml:mi>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mtext>cell&#xa0;</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>Theoretically, the electric potential of the PEMFC varies with temperature and pressure, as expressed in the following equation:<disp-formula id="e3">
<mml:math id="m3">
<mml:mrow>
<mml:mi>E</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1.229</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>0.85</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mn>10</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>298.15</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>4.3085</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mn>10</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>5</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mi>T</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>ln</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mtext>H</mml:mtext>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>ln</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mn>0</mml:mn>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>
</p>
<sec id="s2-1-1">
<title>Thermodynamic Electric Potential</title>
<p>The thermodynamic electric potential of the single cell (i.e.,&#x20;the Nernst electric potential) can be obtained from the mechanism of the electrochemical reaction of the gas inside the PEMFC. This is represented by the following equation:<disp-formula id="e4">
<mml:math id="m4">
<mml:mrow>
<mml:mi>E</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>&#x394;</mml:mi>
<mml:mi>G</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>F</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>&#x394;</mml:mi>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>F</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mrow>
<mml:mtext>ref</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>F</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>ln</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mtext>H</mml:mtext>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:mi>ln</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mn>0</mml:mn>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>
</p>
</sec>
<sec id="s2-1-2">
<title>Activation Overvoltage</title>
<p>The activation overvoltage of the PEMFC is expressed as follows:<disp-formula id="e5">
<mml:math id="m5">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b7;</mml:mi>
<mml:mrow>
<mml:mtext>act</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>&#x3be;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>&#x3be;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mi>T</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>&#x3be;</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mi>T</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>ln</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>c</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>O</mml:mtext>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>&#x3be;</mml:mi>
<mml:mn>4</mml:mn>
</mml:msub>
<mml:mi>T</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>ln</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>I</mml:mi>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>Whereby <inline-formula id="inf1">
<mml:math id="m6">
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>O</mml:mtext>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the concentration of dissolved oxygen at the cathode catalyst interface, which can be expressed by Henry&#x2019;s law as follows:<disp-formula id="e6">
<mml:math id="m7">
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>O</mml:mtext>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mtext>O</mml:mtext>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:mn>5.08</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mn>10</mml:mn>
</mml:mrow>
<mml:mn>6</mml:mn>
</mml:msup>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>498</mml:mn>
<mml:mo>/</mml:mo>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>
</p>
</sec>
<sec id="s2-1-3">
<title>Ohmic Voltage Drop</title>
<p>The ohmic overvoltage is represented by the following equation:<disp-formula id="e7">
<mml:math id="m8">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b7;</mml:mi>
<mml:mrow>
<mml:mtext>ohm&#xa0;</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>I</mml:mi>
<mml:msub>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mtext>int&#xa0;</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>I</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>R</mml:mi>
<mml:mtext>m</mml:mtext>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>R</mml:mi>
<mml:mtext>c</mml:mtext>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>Empirically, the internal resistance of the PEMFC is expressed as follows:<disp-formula id="e8">
<mml:math id="m9">
<mml:mrow>
<mml:msub>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mtext>int</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0.01605</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>3.5</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mn>10</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>5</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mi>T</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>8</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mn>10</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>5</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:math>
<label>(8)</label>
</disp-formula>
</p>
</sec>
<sec id="s2-1-4">
<title>Dense Differential Polarization Overvoltage</title>
<p>The differential overvoltage can be expressed as follows:<disp-formula id="e9">
<mml:math id="m10">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b7;</mml:mi>
<mml:mrow>
<mml:mtext>con</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3b2;</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>ln</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>J</mml:mi>
<mml:mo>/</mml:mo>
<mml:msub>
<mml:mi>J</mml:mi>
<mml:mrow>
<mml:mi>max</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(9)</label>
</disp-formula>
</p>
</sec>
<sec id="s2-1-5">
<title>Dynamic and Capacitive Characteristics of the Double Layer Charge</title>
<p>The dynamic characteristics of the double layer charge of the PEMFC are similar to those of the capacitor, and the equivalent circuit diagram is shown in <xref ref-type="fig" rid="F1">Figure&#x20;1A</xref>:</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>PEMFC equivalent circuit and DC-DC converter. <bold>(A)</bold>Equivalent Circuit Diagram of PEMFC. <bold>(B)</bold>DC/DC converter structure.</p>
</caption>
<graphic xlink:href="fenrg-09-741101-g001.tif"/>
</fig>
<p>As detailed in the figure, the polarization voltage across <italic>R</italic>
<sub>
<italic>d</italic>
</sub> is <italic>V</italic>
<sub>
<italic>d</italic>
</sub> and the differential equation for the voltage change of a single cell is expressed as follows:<disp-formula id="e10">
<mml:math id="m11">
<mml:mrow>
<mml:mtext>d</mml:mtext>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mtext>d</mml:mtext>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:mtext>d</mml:mtext>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>I</mml:mi>
<mml:mo>/</mml:mo>
<mml:mi>C</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mtext>d</mml:mtext>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:msub>
<mml:mi>R</mml:mi>
<mml:mtext>d</mml:mtext>
</mml:msub>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:math>
<label>(10)</label>
</disp-formula>
</p>
</sec>
</sec>
<sec id="s2-2">
<title>PEMFC Stack Voltage</title>
<p>The stack voltage is defined as the value of the voltage at the front end of the PEMFC as it passes through the DC/DC converter. It is assumed that hydrogen is supplied from a hydrogen tank, and is available in sufficient quantities at all times. The air, on the other hand, is controlled by a proportional valve, which allows the air to be controlled efficiently and in time to meet the PEMFC requirements.</p>
<p>
<xref ref-type="disp-formula" rid="e11">Eq. 11</xref> can be obtained from The Law of Conservation of Mass, and the Ideal Gas Law:<disp-formula id="e11">
<mml:math id="m12">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mi>o</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>8.314</mml:mn>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#xd7;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>m</mml:mi>
<mml:mrow>
<mml:mi>h</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>K</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>E</mml:mi>
<mml:mi>H</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>0.5</mml:mn>
<mml:mi>N</mml:mi>
<mml:mi>I</mml:mi>
</mml:mrow>
<mml:mi>F</mml:mi>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(11)</label>
</disp-formula>
</p>
</sec>
<sec id="s2-3">
<title>DC-DC Boost Converter Model</title>
<p>The output voltage of the PEMFC is the tap voltage of the DC/DC converter. A boost converter is essentially a step-up power converter, i.e.,&#x20;the voltage is raised and then outputted. An DC/DC boost converter circuit is shown in <xref ref-type="fig" rid="F1">Figure&#x20;1B</xref>:</p>
<p>Whereby the input and output voltage relationship are controlled output voltage by the switch duty cycle, as expressed in Equation:<disp-formula id="e12">
<mml:math id="m13">
<mml:mrow>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mtext>ou</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mtext>stack</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
<label>(12)</label>
</disp-formula>The differential equation for <italic>V</italic>
<sub>
<italic>out</italic>
</sub> is as follows:<disp-formula id="e13">
<mml:math id="m14">
<mml:mrow>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mtable columnalign="left">
<mml:mtr columnalign="left">
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mtext>d</mml:mtext>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mi>L</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mtext>d</mml:mtext>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>L</mml:mi>
</mml:mfrac>
<mml:mo>&#x22c5;</mml:mo>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mtext>stack</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr columnalign="left">
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mtext>d</mml:mtext>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mtext>out</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mtext>d</mml:mtext>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>C</mml:mi>
</mml:mfrac>
<mml:mo>&#x22c5;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mtext>ost</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(13)</label>
</disp-formula>
</p>
</sec>
</sec>
<sec id="s3">
<title>Proposed Method</title>
<sec id="s3-1">
<title>Framework of Control Policy</title>
<p>The control model includes a PEMFC stack, a DC/DC converter and its controller. The controller of the DC/DC converter is equated to an intelligent agent which is trained to adapt to the non-linear characteristics of the PEMFC so as to improve the overall output voltage control performance. When applied online, the intelligent agent outputs the optimal duty cycle according to the state of the DC-DC converter and the state of the output voltage. The control interval of the agent is 0.01&#xa0;s.</p>
<sec id="s3-1-1">
<title>Agent</title>
<p>
<list list-type="simple">
<list-item>
<p>1) Action&#x20;space</p>
</list-item>
</list>
</p>
<p>The action space is set to <italic>u</italic>/100<italic>,</italic> as follows:<disp-formula id="e14">
<mml:math id="m15">
<mml:mrow>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mtable columnalign="left">
<mml:mtr>
<mml:mtd>
<mml:mi>a</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>/</mml:mo>
<mml:mn>100</mml:mn>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mtext>0</mml:mtext>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>u</mml:mi>
<mml:mo>&#x2264;</mml:mo>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mrow>
<mml:mi>max</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(14)</label>
</disp-formula>
<list list-type="simple">
<list-item>
<p>2) State&#x20;space</p>
</list-item>
</list>
</p>
<p>The state space is expressed as follows:<disp-formula id="e15">
<mml:math id="m16">
<mml:mrow>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>e</mml:mi>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mstyle displaystyle="true">
<mml:mrow>
<mml:msubsup>
<mml:mo>&#x222b;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mi>t</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>t</mml:mi>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mi>U</mml:mi>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(15)</label>
</disp-formula>
<list list-type="simple">
<list-item>
<p>3) Reward function</p>
</list-item>
</list>
</p>
<p>The reward function is expressed as follows:<disp-formula id="e16">
<mml:math id="m17">
<mml:mrow>
<mml:mtext>r</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3bc;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>&#x3bc;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mi>u</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
</mml:math>
<label>(16)</label>
</disp-formula>
<disp-formula id="e17">
<mml:math id="m18">
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mtable columnalign="left">
<mml:mtr>
<mml:mtd>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>0.3</mml:mn>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3e;</mml:mo>
<mml:mn>0.09</mml:mn>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mn>0</mml:mn>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:msubsup>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2264;</mml:mo>
<mml:mn>0.09</mml:mn>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(17)</label>
</disp-formula>
</p>
</sec>
</sec>
<sec id="s3-2">
<title>DDPG</title>
<p>The Deterministic Policy Gradient (DDPG) policy determines an action via the policy function <italic>&#xb5;(s)</italic>, which is shown in the following equation:<disp-formula id="e18">
<mml:math id="m19">
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3bc;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mi mathvariant="normal">&#x7c;</mml:mi>
<mml:msup>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>&#x3bc;</mml:mi>
</mml:msup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(18)</label>
</disp-formula>This deep reinforcement learning algorithm uses a value network to fit the function Q(s) and the objective function <inline-formula id="inf2">
<mml:math id="m20">
<mml:mrow>
<mml:mi>J</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>&#x3bc;</mml:mi>
</mml:msup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, the latter which is defined as follows:<disp-formula id="e19">
<mml:math id="m21">
<mml:mrow>
<mml:mi>J</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>&#x3bc;</mml:mi>
</mml:msup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b3;</mml:mi>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msup>
<mml:mi>&#x3b3;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:msub>
<mml:mi>r</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mo>&#x22ef;</mml:mo>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(19)</label>
</disp-formula>In this arrangement, the Q function can be expressed as the expected value of the reward for selecting an action under <italic>&#xb5;</italic>(<italic>s</italic>).</p>
<p>In each step, a specific policy is randomly selected for the agent to be executed, and the best policy is selected by maximizing the fusion objective function. The different policy will be executed in different steps, so that an experience replay pool can be obtained for each agent. Finally, the gradient of the fusion objective function <inline-formula id="inf3">
<mml:math id="m22">
<mml:mrow>
<mml:msub>
<mml:mo>&#x2207;</mml:mo>
<mml:mrow>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi>J</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is solved for the policy parameters of each agent, as expressed in the following equation:<disp-formula id="e20">
<mml:math id="m23">
<mml:mrow>
<mml:msub>
<mml:mo>&#x2207;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mi>J</mml:mi>
<mml:mo>&#x2248;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>S</mml:mi>
</mml:mfrac>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>j</mml:mi>
</mml:munder>
<mml:mrow>
<mml:msub>
<mml:mo>&#x2207;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:msub>
<mml:mi>&#x3bc;</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>O</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:msub>
<mml:mo>&#x2207;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mi>Q</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi mathvariant="normal">&#x3f5;</mml:mi>
</mml:msubsup>
<mml:mo>&#x22c5;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi>x</mml:mi>
<mml:mi>j</mml:mi>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>a</mml:mi>
<mml:mn>1</mml:mn>
<mml:mi>j</mml:mi>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mo>&#x22ef;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x22ef;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>a</mml:mi>
<mml:mi>N</mml:mi>
<mml:mi>j</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3bc;</mml:mi>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>O</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
<label>(20)</label>
</disp-formula>
</p>
<p>Nevertheless, the DDPG algorithm suffers from low robustness. The main reasons for this are as follows:<list list-type="simple">
<list-item>
<p>1) The algorithm lacks effective bootstrapping techniques, and so it tends to fall into the local optimum solution, which undermines the robustness of the strategy.</p>
</list-item>
<list-item>
<p>2) Overestimation of the Q-value leads to overfitting of the algorithm&#x2019;s policy, thus making it less robust.</p>
</list-item>
</list>
</p>
</sec>
<sec id="s3-3">
<title>Framework for Offline Training of IO-TD3</title>
<p>In order to address the low robustness of the DDPG algorithm, the IO-TD3 algorithm incorporates the following two innovations:<list list-type="simple">
<list-item>
<p>1) An imitation-oriented distributed training framework for deep reinforcement learning;&#x20;and,</p>
</list-item>
<list-item>
<p>2) An Integrated anti-Q overestimation policy.</p>
</list-item>
</list>
</p>
<p>The large-scale deep reinforcement learning training framework for the IO-TD3 algorithm is illustrated.</p>
<p>The algorithm contains three roles, an explorer, an expert and a leader. A total of 36 parallel systems are included in the algorithm, each containing the same PEMFC system and different load disturbances, so as to enhance sample diversity.</p>
<sec id="s3-3-1">
<title>Explorer</title>
<p>The Explorer contains only one actor network. The explorers in different parallel systems employ their own different exploration principles. The explorers described in this paper use the following exploration principles: greedy strategy, Gaussian noise, and OU&#x20;noise.</p>
<p>The explorer in parallel system 1&#x2013;6 uses an <italic>&#x3b5;</italic>-greedy strategy with the following actions:<disp-formula id="e21">
<mml:math id="m24">
<mml:mrow>
<mml:msubsup>
<mml:mi>a</mml:mi>
<mml:mi>&#x3b5;</mml:mi>
<mml:mi>l</mml:mi>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mtable columnalign="left">
<mml:mtr columnalign="left">
<mml:mtd columnalign="left">
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3c0;</mml:mi>
<mml:mi>&#x3d5;</mml:mi>
<mml:mi>l</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>s</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mtext>With</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mi>&#x3b5;</mml:mi>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>probability</mml:mtext>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr columnalign="left">
<mml:mtd columnalign="left">
<mml:mrow>
<mml:msubsup>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mtext>rand</mml:mtext>
</mml:mrow>
<mml:mi>l</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mtext>With</mml:mtext>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mn>1</mml:mn>
<mml:mo>-</mml:mo>
<mml:mi>&#x3b5;</mml:mi>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mtext>probability</mml:mtext>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(21)</label>
</disp-formula>The explorer in parallel system 7&#x2013;12 uses an OU noise exploration strategy with the following actions:<disp-formula id="e22">
<mml:math id="m25">
<mml:mrow>
<mml:msubsup>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mi>O</mml:mi>
<mml:mi>U</mml:mi>
</mml:mrow>
<mml:mi>j</mml:mi>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mi>&#x3c0;</mml:mi>
<mml:mi>&#x3d5;</mml:mi>
<mml:mi>j</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>s</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>O</mml:mi>
<mml:mi>U</mml:mi>
</mml:mrow>
<mml:mi>j</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
<label>(22)</label>
</disp-formula>The Gaussian noise exploration strategy used in parallel system 13&#x2013;18 has the following actions:<disp-formula id="e23">
<mml:math id="m26">
<mml:mrow>
<mml:msubsup>
<mml:mi>a</mml:mi>
<mml:mrow>
<mml:mtext>Gaussian</mml:mtext>
</mml:mrow>
<mml:mi>m</mml:mi>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mi>&#x3c0;</mml:mi>
<mml:mi>&#x3d5;</mml:mi>
<mml:mi>m</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>s</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mtext>Gaussian</mml:mtext>
</mml:mrow>
<mml:mi>m</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
<label>(23)</label>
</disp-formula>
</p>
</sec>
<sec id="s3-3-2">
<title>Expert</title>
<p>On the basis of imitation learning, the proposed algorithm employs a large number of expert samples, which are used as learning samples, so that the algorithm can be effectively guided to learn correctly during the early stages of training. In this proposed method, the duty cycle of the DC/DC converter is controlled, whereby the parallel systems generate expert samples for the Leader (described below).</p>
<p>The expert itself uses a variety of controllers based on different principles, including PSO-PID and GA-PID algorithms. The objective function for parameter optimization is as follows:<disp-formula id="e24">
<mml:math id="m27">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:mrow>
<mml:msubsup>
<mml:mo>&#x222b;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mi>&#x221e;</mml:mi>
</mml:msubsup>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mstyle>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mtext>d</mml:mtext>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:math>
<label>(24)</label>
</disp-formula>
</p>
</sec>
<sec id="s3-3-3">
<title>Leader</title>
<p>The leader (termed &#x201c;Leader&#x201d;) entails a complete agent structure which includes a two-actor network, two critic networks, and an experience pool. It learns samples from the explorer and the critic in order to obtain the optimal control strategy, and periodically sends the latest parameters to the actor network for all the explorers.</p>
<p>The critic in each leader employs an integrated mitigation Q over-estimation technique.<list list-type="simple">
<list-item>
<p>1) The critic in Leader uses the Clipped Double Q-learning technique to calculate the target value:</p>
</list-item>
</list>
<disp-formula id="e25">
<mml:math id="m28">
<mml:mrow>
<mml:msubsup>
<mml:mi>y</mml:mi>
<mml:mi>t</mml:mi>
<mml:mn>1</mml:mn>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>r</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b3;</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>min</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1,2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>i</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msubsup>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>&#x3c0;</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3d5;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(25)</label>
</disp-formula>
<list list-type="simple">
<list-item>
<p>2) The critic network inside Leader uses a policy delay update policy. <italic>d</italic> updates to the actor network are performed after every <italic>d</italic> update to the critic.</p>
</list-item>
<list-item>
<p>3) The critic inside Leader uses a goal policy smoothing regularization strategy. The critic introduces a regularization method for reducing the variance of the goal values by bootstrapping the estimates of similar state action&#x20;pairs.</p>
</list-item>
</list>
<disp-formula id="e26">
<mml:math id="m29">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>r</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mtext>E</mml:mtext>
<mml:mi>&#x3b5;</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:msup>
<mml:mi>&#x3b8;</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>&#x3c0;</mml:mi>
<mml:mrow>
<mml:msup>
<mml:mi>&#x3d5;</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b5;</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(26)</label>
</disp-formula>
</p>
<p>Smooth regularization is also achieved by adding a random noise to the target strategy and averaging over the mini-batch:<disp-formula id="e27">
<mml:math id="m30">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>r</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b3;</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>min</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1,2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b8;</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>&#x3c0;</mml:mi>
<mml:mrow>
<mml:msup>
<mml:mi>&#x3d5;</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b5;</mml:mi>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(27)</label>
</disp-formula>
<disp-formula id="e28">
<mml:math id="m31">
<mml:mrow>
<mml:mi>&#x3b5;</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:mi>clip</mml:mi>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>c</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(28)</label>
</disp-formula>
</p>
</sec>
</sec>
</sec>
<sec id="s4">
<title>Experiment</title>
<p>In order to verify the superior effectiveness of the proposed method, the IO-TD3 algorithm control strategy was tested against the following methods in case: Ape-x-MADDPG control algorithm (40), MATD3 control algorithm (41), MADDPG coordinated control algorithm (37), BP neural network control algorithm, RBF neural network control algorithm, PSO optimized PID control algorithm (PSO-PID), GA optimized PID control algorithm (GA-PID), PID control algorithm (PID), Fuzzy-FOPID control algorithm (Fuzzy-FOPID), and the PSO-optimized FOPID control algorithm (PSO-FOPID). The first six (including the IO-TD3 algorithm) are referred to as advanced algorithms, and the last five are conventional algorithms.</p>
<p>At 1&#xa0;s, the load current magnitude appears as a load disturbance which begins at 72.6&#xa0;A and rises to 250.0&#xa0;A. The results are shown in <xref ref-type="fig" rid="F2">Figure&#x20;2A,B</xref>.<list list-type="simple">
<list-item>
<p>1) Comparison between proposed algorithm and advanced algorithms. As shown in <xref ref-type="fig" rid="F2">Figure&#x20;2A</xref>, the IO-TD3 algorithm has a better response time, smoother output voltage profile and no overshoot. The proposed algorithm&#x2019;s minimum output voltage value is smaller than that of the other advanced algorithms. Conversely, each of the output voltages of the other advanced algorithms is characterized by large overshoot, and these results are affected by varying degrees of overshoot and oscillation, which can lead to unstable output voltages. The IO-TD3 algorithm therefore has the best control performance.</p>
</list-item>
<list-item>
<p>2) Possible reasons for these promising patterns are as follows: firstly, other DRL algorithms tend to fall into local optima; they amount to sub-optimal control strategies as they are not effectively guided in pre-learning, resulting in large output voltage overshoot and output voltage fluctuations, which undermine PEMFC output performance.</p>
</list-item>
</list>
</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Results of Case 1. <bold>(A)</bold> Output voltage of advanced algorithms <bold>(B)</bold> Output voltage of conventional algorithms</p>
</caption>
<graphic xlink:href="fenrg-09-741101-g002.tif"/>
</fig>
<p>The BP and RBF algorithms are too dependent on the trained samples, resulting in limited control performance. A neural network control algorithm which lacks self-exploration will have lower adaptive ability, leading to poorer control performance.</p>
<p>The PSO-PID and GA-PID algorithms within the conventional control algorithm group lack the adaptive capability for adjusting the PID parameters, and therefore struggle to adapt to the non-linearity of the PEMFC environment. The PSO-FOPID algorithm enables greater robustness in the environment, but is impaired by poor adaptive capability due to its fixed coefficients, which ultimately leads to severe output voltage overshoot and oscillation. The Fuzzy-FOPID algorithm, despite its better adaptive capability, is underpinned by overly simple rules, resulting in poor control accuracy and therefore a large overshoot despite the fast response of the algorithm.</p>
<p>In summary, the IO-TD3 controller is a more suitable candidate for practical output voltage control systems, with its short response times, and good dynamic and static performance indicators.</p>
</sec>
<sec sec-type="conclusion" id="s5">
<title>Conclusion</title>
<p>In this paper, an imitation-oriented deep reinforcement learning output voltage control strategy for controlling the duty cycle of a DC-DC converter has been proposed. The proposed method is an imitation-oriented twin delay deep deterministic (IO-TD3) policy gradient algorithm, the design of which is structured on the concept of imitation learning. It embodies a distributed deep reinforcement learning training framework designed to improve the robustness of the control policy. The effectiveness of the proposed control policy has been experimentally demonstrated. The simulation results show that the IO-TD3 algorithm has superior control performance compared to other deep reinforcement learning algorithms (e.g., Ape-x-MADDPG, MATD3, MADDPG). Compared to other control algorithms (BP, RBF, PSO-PID, GA-PID, PID, Fuzzy-FOPID, PSO-FOPID), the IO-TD3 algorithm is more adaptable, and, in relation to the output voltage of the PEMFC, has better response speed and stability, and can more effectively track and control the output voltage in a timely and effective manner.</p>
</sec>
</body>
<back>
<sec id="s6">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/Supplementary Material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>JL: conceptualization, methodology, software, data curation, writing-original draft preparation, visualization, investigation, software, validation. YL: Writing-Reviewing and editing. TY: Supervision.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This work was jointly supported by National Natural Science Foundation of China (U2066212).</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s10" sec-type="disclaimer">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bemporad</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Borrelli</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Morari</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Model Predictive Control Based on Linear Programming - the Explicit Solution</article-title>. <source>Ieee Trans. Automat. Contr.</source> <volume>47</volume>, <fpage>1974</fpage>&#x2013;<lpage>1985</lpage>. <pub-id pub-id-type="doi">10.1109/tac.2002.805688</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Bougrine</surname>
<given-names>M. D.</given-names>
</name>
<name>
<surname>Benalia</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Benbouzid</surname>
<given-names>M. H.</given-names>
</name>
</person-group> (<year>2013</year>). &#x201c;<article-title>Nonlinear Adaptive Sliding Mode Control of a Powertrain Supplying Fuel Cell Hybrid Vehicle</article-title>,&#x201d; in <conf-name>3rd International Conference on Systems and Control</conf-name> (<publisher-loc>Algiers, Algeria</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>714</fpage>&#x2013;<lpage>719</lpage>. </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ferrari-Trecate</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Cuzzola</surname>
<given-names>F. A.</given-names>
</name>
<name>
<surname>Mignone</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Morari</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Analysis of Discrete-Time Piecewise Affine and Hybrid Systems</article-title>. <source>Automatica</source> <volume>38</volume>, <fpage>2139</fpage>&#x2013;<lpage>2146</lpage>. <pub-id pub-id-type="doi">10.1016/s0005-1098(02)00142-5</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Jaen</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Pou</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Pindado</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Sala</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Zaragoza</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2006</year>). &#x201c;<article-title>A Linear-Quadratic Regulator with Integral Action Applied to PWM DC-DC Converters</article-title>,&#x201d; in <conf-name>IECON 2006 - 32nd Annual Conference on IEEE Industrial Electronics</conf-name> (<publisher-loc>Paris, France</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2280</fpage>&#x2013;<lpage>2285</lpage>. </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jiao</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Cui</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Cui</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Robustness Analysis of Sliding Mode on DC/DC for Fuel Cell Vehicle</article-title>. <source>Jestr</source> <volume>6</volume>, <fpage>1</fpage>&#x2013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.25103/jestr.065.01</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A New Adaptive Controller Based on Distributed Deep Reinforcement Learning for PEMFC Air Supply System</article-title>. <source>Energ. Rep.</source> <volume>7</volume>, <fpage>1267</fpage>&#x2013;<lpage>1279</lpage>. <pub-id pub-id-type="doi">10.1016/j.egyr.2021.02.043</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Efficient Experience Replay Based Deep Deterministic Policy Gradient for AGC Dispatch in Integrated Energy System</article-title>. <source>Appl. Energ.</source> <volume>285</volume>, <fpage>116386</fpage>. <pub-id pub-id-type="doi">10.1016/j.apenergy.2020.116386</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lillicrap</surname>
<given-names>T. P.</given-names>
</name>
<name>
<surname>Hunt</surname>
<given-names>J.&#x20;J.</given-names>
</name>
<name>
<surname>Pritzel</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Heess</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Erez</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Tassa</surname>
<given-names>Y.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Continuous Control with Deep Reinforcement Learning</article-title>. <comment>arXiv preprint arXiv:1509.02971</comment>. </citation>
</ref>
<ref id="B9">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Montagner</surname>
<given-names>V. F.</given-names>
</name>
<name>
<surname>Maccari</surname>
<given-names>L. A.</given-names>
</name>
<name>
<surname>Dupont</surname>
<given-names>F. H.</given-names>
</name>
<name>
<surname>Pinheiro</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Oliveira</surname>
<given-names>R. C. L. F.</given-names>
</name>
</person-group> (<year>2011</year>). &#x201c;<article-title>A DLQR Applied to Boost Converters with Switched Loads: Design and Analysis</article-title>,&#x201d; in <conf-name>XI Brazilian Power Electronics Conference</conf-name> (<publisher-loc>Natal, Brazil</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>68</fpage>&#x2013;<lpage>73</lpage>. </citation>
</ref>
<ref id="B10">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Moreira</surname>
<given-names>C. O.</given-names>
</name>
<name>
<surname>Silva</surname>
<given-names>F. A.</given-names>
</name>
<name>
<surname>Pinto</surname>
<given-names>S. F.</given-names>
</name>
<name>
<surname>Santos</surname>
<given-names>M. B.</given-names>
</name>
</person-group> (<year>2011</year>). &#x201c;<article-title>Digital LQR Control with Kalman Estimator for DC-DC Buck Converter</article-title>,&#x201d; in <conf-name>2011 IEEE EUROCON - International Conference on Computer as a Tool</conf-name> (<publisher-loc>Lisbon, Portugal</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>4</lpage>. </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Olalla</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Leyva</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>El Aroudi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Queinnec</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Robust LQR Control for PWM Converters: An LMI Approach</article-title>. <source>Ieee Trans. Ind. Electron.</source> <volume>56</volume>, <fpage>2548</fpage>&#x2013;<lpage>2558</lpage>. <pub-id pub-id-type="doi">10.1109/tie.2009.2017556</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Olalla</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Queinnec</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Leyva</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Aroudi</surname>
<given-names>A. E.</given-names>
</name>
</person-group> (<year>2010</year>). &#x201c;<article-title>Robust Control Design of Bilinear DC-DC Converters with Guaranteed Region of Stability</article-title>,&#x201d; in <conf-name>2010 IEEE International Symposium on Industrial Electronics</conf-name> (<publisher-loc>Bari, Italy</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>3005</fpage>&#x2013;<lpage>3010</lpage>. </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Qi</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Rotor Resistance and Excitation Inductance Estimation of an Induction Motor Using Deep-Q-Learning Algorithm</article-title>. <source>Eng. Appl. Artif. Intelligence</source> <volume>72</volume>, <fpage>67</fpage>&#x2013;<lpage>79</lpage>. <pub-id pub-id-type="doi">10.1016/j.engappai.2018.03.018</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Jin</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Pan</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Shen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>K. Y.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Efficiency Analysis and Control of a Grid-Connected PEM Fuel Cell in Distributed Generation</article-title>. <source>Energ. Convers. Manage.</source> <volume>195</volume>, <fpage>587</fpage>&#x2013;<lpage>596</lpage>. <pub-id pub-id-type="doi">10.1016/j.enconman.2019.04.041</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Swain</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Jena</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>PID Control Design for the Pressure Regulation of PEM Fuel Cell</article-title>,&#x201d; in <conf-name>2015 International Conference on Recent Developments in Control, Automation and Power Engineering (RDCAPE)</conf-name> (<publisher-loc>Noida, India</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>286</fpage>&#x2013;<lpage>291</lpage>. </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2021a</year>). <article-title>Parameter Extraction of PEMFC via Bayesian Regularization Neural Network Based Meta-Heuristic Algorithms</article-title>. <source>Energy</source> <volume>228</volume>, <fpage>120592</fpage>. <pub-id pub-id-type="doi">10.1016/j.energy.2021.120592</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Swe</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Shu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<etal/>
</person-group> (<year>2021b</year>). <article-title>Energy Cooperation between Myanmar and China under One Belt One Road: Current State, Challenges and Perspectives</article-title>. <source>Energy</source> <volume>215</volume>, <fpage>119130</fpage>. <pub-id pub-id-type="doi">10.1016/j.energy.2020.119130</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Yao</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Shu</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Comprehensive Overview of Meta-Heuristic Algorithm Applications on PV Cell Parameter Identification</article-title>. <source>Energ. Convers. Manage.</source> <volume>208</volume>, <fpage>112595</fpage>. <pub-id pub-id-type="doi">10.1016/j.enconman.2020.112595</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Shu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Dong</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Robust Sliding-Mode Control of Wind Energy Conversion Systems for Optimal Power Extraction via Nonlinear Perturbation Observers</article-title>. <source>Appl. Energ.</source> <volume>210</volume>, <fpage>711</fpage>&#x2013;<lpage>723</lpage>. <pub-id pub-id-type="doi">10.1016/j.apenergy.2017.08.027</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Shu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Sang</surname>
<given-names>Y.</given-names>
</name>
<etal/>
</person-group> (<year>2019a</year>). <article-title>Dynamic Leader Based Collective Intelligence for Maximum Power point Tracking of PV Systems Affected by Partial Shading Condition</article-title>. <source>Energ. Convers. Manage.</source> <volume>179</volume>, <fpage>286</fpage>&#x2013;<lpage>303</lpage>. <pub-id pub-id-type="doi">10.1016/j.enconman.2018.10.074</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>Z.</given-names>
</name>
<etal/>
</person-group> (<year>2021c</year>). <article-title>Parameter Identification of Proton Exchange Membrane Fuel Cell via Levenberg-Marquardt Backpropagation Algorithm</article-title>. <source>Int. J.&#x20;Hydrogen Energ.</source> <volume>46</volume>, <fpage>22998</fpage>&#x2013;<lpage>23012</lpage>. <pub-id pub-id-type="doi">10.1016/j.ijhydene.2021.04.130</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Zhong</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Shu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2019b</year>). <article-title>Novel Bio-Inspired Memetic Salp Swarm Algorithm and Application to MPPT for PV Systems Considering Partial Shading Condition</article-title>. <source>J.&#x20;Clean. Prod.</source> <volume>215</volume>, <fpage>1203</fpage>&#x2013;<lpage>1222</lpage>. <pub-id pub-id-type="doi">10.1016/j.jclepro.2019.01.150</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Memetic Reinforcement Learning Based Maximum Power point Tracking Design for PV Systems under Partial Shading Condition</article-title>. <source>Energy</source> <volume>174</volume>, <fpage>1079</fpage>&#x2013;<lpage>1090</lpage>. <pub-id pub-id-type="doi">10.1016/j.energy.2019.03.053</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Tan</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Adaptive Distributed Auction-Based Algorithm for Optimal Mileage Based AGC Dispatch with High Participation of Renewable Energy</article-title>. <source>Int. J.&#x20;Electr. Power Energ. Syst.</source> <volume>124</volume>, <fpage>106371</fpage>. <pub-id pub-id-type="doi">10.1016/j.ijepes.2020.106371</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Robust Collaborative Consensus Algorithm for Decentralized Economic Dispatch with a Practical Communication Network</article-title>. <source>Electric Power Syst. Res.</source> <volume>140</volume>, <fpage>597</fpage>&#x2013;<lpage>610</lpage>. <pub-id pub-id-type="doi">10.1016/j.epsr.2016.05.014</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Fast Stackelberg Equilibrium Learning for Real-Time Coordinated Energy Control of a Multi-Area Integrated Energy System</article-title>. <source>Appl. Therm. Eng.</source> <volume>153</volume>, <fpage>225</fpage>&#x2013;<lpage>241</lpage>. <pub-id pub-id-type="doi">10.1016/j.applthermaleng.2019.02.053</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Fan</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>A Cyber-Physical-Social System with Parallel Learning for Distributed Energy Management of a Microgrid</article-title>. <source>Energy</source> <volume>165</volume>, <fpage>205</fpage>&#x2013;<lpage>221</lpage>. <pub-id pub-id-type="doi">10.1016/j.energy.2018.09.069</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Mou</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Ding</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>UAV-enabled Secure Communications by Multi-Agent Deep Reinforcement Learning</article-title>. <source>IEEE Trans. Veh. Technol.</source> <volume>69</volume>, <fpage>11599</fpage>&#x2013;<lpage>11611</lpage>. <pub-id pub-id-type="doi">10.1109/tvt.2020.3014788</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>