<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">749878</article-id>
<article-id pub-id-type="doi">10.3389/frai.2021.749878</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>QF-TraderNet: Intraday Trading <italic>via</italic> Deep Reinforcement With Quantum Price Levels Based Profit-And-Loss Control</article-title>
<alt-title alt-title-type="left-running-head">Qiu et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">QF-TraderNet</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Qiu</surname>
<given-names>Yifu</given-names>
</name>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Qiu</surname>
<given-names>Yitao</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/1424944/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Yuan</surname>
<given-names>Yicong</given-names>
</name>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Chen</surname>
<given-names>Zheng</given-names>
</name>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Lee</surname>
<given-names>Raymond</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1295357/overview"/>
</contrib>
</contrib-group>
<aff>Department of Computer Science and Technology, Division of Science and Technology, BNU-HKBU United International College, <addr-line>Zhuhai</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/592652/overview">Ronald Hochreiter</ext-link>, Vienna University of Economics and Business, Austria</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/590352/overview">Paolo Pagnottoni</ext-link>, University of Pavia, Italy</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/899035/overview">Taiyong Li</ext-link>, Southwestern University of Finance and Economics, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Raymond Lee, <email>raymondshtlee@uic.edu.cn</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Artificial Intelligence in Finance, a section of the journal Frontiers in Artificial Intelligence</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>29</day>
<month>10</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>4</volume>
<elocation-id>749878</elocation-id>
<history>
<date date-type="received">
<day>30</day>
<month>07</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>21</day>
<month>09</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Qiu, Qiu, Yuan, Chen and Lee.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Qiu, Qiu, Yuan, Chen and Lee</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Reinforcement Learning (RL) based machine trading attracts a rich profusion of interest. However, in the existing research, RL in the day-trade task suffers from the noisy financial movement in the short time scale, difficulty in order settlement, and expensive action search in a continuous-value space. This paper introduced an end-to-end RL intraday trading agent, namely QF-TraderNet, based on the quantum finance theory (QFT) and deep reinforcement learning. We proposed a novel design for the intraday RL trader&#x2019;s action space, inspired by the Quantum Price Levels (QPLs). Our action space design also brings the model a learnable profit-and-loss control strategy. QF-TraderNet composes two neural networks: 1) A long short term memory networks for the feature learning of financial time series; 2) a policy generator network (PGN) for generating the distribution of actions. The profitability and robustness of QF-TraderNet have been verified in multi-type financial datasets, including FOREX, metals, crude oil, and financial indices. The experimental results demonstrate that QF-TraderNet outperforms other baselines in terms of cumulative price returns and Sharpe Ratio, and the robustness in the acceidential market&#x20;shift.</p>
</abstract>
<kwd-group>
<kwd>quantum finance</kwd>
<kwd>quantum price level</kwd>
<kwd>reinforcement learning</kwd>
<kwd>automatic trading</kwd>
<kwd>intelligent trading system</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Financial trading is an online decision-making process (<xref ref-type="bibr" rid="B5">Deng et&#x20;al., 2016</xref>). Previous works (<xref ref-type="bibr" rid="B16">Moody and Saffell, 1998</xref>; <xref ref-type="bibr" rid="B17">Moody and Saffell, 2001</xref>; <xref ref-type="bibr" rid="B4">Dempster and Leemans, 2006</xref>) demonstrated the Reinforcement Learning (RL) agent&#x2019;s promising profitability in trading activities. However, traditional RL algorithms face challenges for the intraday trading problem in three aspects: 1) Short-term financial movement is often accompanied by more noisy oscillations. 2) The computational complexity for making decision in daily continuous-value price range. In the <italic>T</italic>&#x20;&#x2b; <italic>n</italic> strategy, RL agents are assigned a long, neutral, or short position in each trading day, including the Fuzzy Deep Recurrent Neural Networks (FDRNN) (<xref ref-type="bibr" rid="B5">Deng et&#x20;al., 2016</xref>) and Direct Reinforcement Learning (DRL) (<xref ref-type="bibr" rid="B17">Moody and Saffell, 2001</xref>). However, in day trade, i.e.,&#x20;<italic>T</italic>&#x20;&#x2b; 0 strategy, the trading task is converted to identify the optimal price to open and close the order. 3) The early stop of orders when applying the intraday strategy. Conventionally, the settlement of orders involved two hyperparameters: Target Profit (TP) and Stop Loss (SL). TP refers to the price to close the activating order and take out the profit if the price moved as expected. SL denotes the price to terminate the transaction and avoid a further loss if the price moved towards a loss direction (e.g., the price dropped down following a long position decision). These two hyperparameters are defined as a fixed shift relative to price to enter the market, as known as, points. If the price touched these two-preset levels, the order will be closed deterministically. An instance of the early-stop order is shown in <xref ref-type="fig" rid="F1">Figure&#x20;1</xref>.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>An early-stop loss problem: a short order is early settled (red dash line: SL) before the price drops to the profitable range. Thus, the strategy loses the potential profit (blue double arrow).</p>
</caption>
<graphic xlink:href="frai-04-749878-g001.tif"/>
</fig>
<p>Focusing on the mentioned challenges, we proposed a deep reinforcement learning-based end-to-end learning model, named QF-TraderNet. Our model directly generates the trading policy to control profit and loss instead of using fixed TP and SL. QF-TraderNet comprises two neural networks with different functions: 1) a Long-short Term Memory (LSTM) networks for extracting the temporal feature in financial time series; 2) a policy generator network (PGN) for generating the distribution of actions (policy) in each state. We especially reference the Quantum Price Levels (QPLs) as illustrated in <xref ref-type="fig" rid="F2">Figure 2</xref> to design the action space for the RL agent, thus discretizing the price-value space. Our method is inspired by the Quantum Finance Theory that QPLs captures the equilibrium states of price movement on a daily basis (<xref ref-type="bibr" rid="B11">Lee, 2020</xref>). We utilize the deep reinforcement learning algorithm to update the trainable parameters of QF-TraderNet iteratively to maximize the cumulative price return.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Illustration of AUDUSD&#x2019;s QPLs in 3 consecutive trading days (23/04/2020&#x2013;27/04/2020) in 30-min K-line graph. The blue lines represent negative QPLs based on the ground state (black dash line); the red lines are positive QPLs. Line color deepens with the rise of the QPL level <italic>n</italic>.</p>
</caption>
<graphic xlink:href="frai-04-749878-g002.tif"/>
</fig>
<p>Experiments on various financial datasets, including the financial indices, metals, crude oil, and FOREX, and comparisons with previous RL and DL-based single-product trading systems have been conducted. Our QF-TraderNet outperforms some state-of-the-art baselines in the profitability evaluated by the cumulative return and the risk-adjusted return (Sharpe ratio), and the robustness facing market turbulence. Our model shows adaptability in the unseen market environment. The generated policy of QF-TraderNet also provides an explainable profit-and-loss order control strategy.</p>
<p>Our main contributions could be summarized as:<list list-type="simple">
<list-item>
<p>&#x2022; We propose a novel end-to-end daytrade model that directly learns the optimal price level to settle, thus solving the early stop in an implicit stop-loss and target-profit setting.</p>
</list-item>
<list-item>
<p>&#x2022; We are the first to present RL agent&#x2019;s action space <italic>via</italic> the daily quantum price level, making the machine day trade tractable.</p>
</list-item>
<list-item>
<p>&#x2022; Under the same market information perception, we achieve better profitability and robustness than previous state-of-the-art RL based models.</p>
</list-item>
</list>
</p>
</sec>
<sec id="s2">
<title>2 Related Work</title>
<p>Our work is in line with two sub-tasks: financial feature extraction and transactions based on deep reinforcement learning. We shortly review past studies.</p>
<sec id="s2-1">
<title>2.1 Financial Feature Extraction and Representation</title>
<p>Computational approaches for the applications in financial modeling have attracted much attention in the past. (<xref ref-type="bibr" rid="B20">Peralta and Zareei, 2016</xref>). utilized the network model to perform the portfolio planning and selection. <xref ref-type="bibr" rid="B8">Giudici et&#x20;al. (2021)</xref> used volatility spillover decomposition methods to model the relations between two currencies. <xref ref-type="bibr" rid="B22">Resta et&#x20;al. (2020)</xref> conducted a technical analysis-based approach to identify the trading opportunities with specific on cryptocurrency. Among these, the neural networks shows promising ability in learning both the structured and unstructured data. Most of the related works in neural financial modeling were made to the relationship embedding (<xref ref-type="bibr" rid="B12">Li et&#x20;al., 2019</xref>) and forecasting (<xref ref-type="bibr" rid="B29">Wei et&#x20;al., 2017</xref>), option pricing (<xref ref-type="bibr" rid="B19">Pagnottoni, 2019</xref>), and forecasting (<xref ref-type="bibr" rid="B18">Neely et&#x20;al., 2014</xref>). The long short-term memory networks (LSTM) (<xref ref-type="bibr" rid="B29">Wei et&#x20;al., 2017</xref>), Elman recurrent neural networks (<xref ref-type="bibr" rid="B27">Wang et&#x20;al., 2016</xref>) were employed in financial time series analysis tasks successfully. <xref ref-type="bibr" rid="B25">Tran et&#x20;al. (2018)</xref> utilized the attention mechanism to refine RNN. (<xref ref-type="bibr" rid="B15">Mohan et&#x20;al., 2019</xref>). leveraged both market and textual information to boost the performance of stock prediction. Some studies also adopted stock embedding to mine the affinity indicators (<xref ref-type="bibr" rid="B1">Chen et&#x20;al., 2019</xref>).</p>
</sec>
<sec id="s2-2">
<title>2.2 Reinforcement Learning in Trading</title>
<p>Algorithmic trading has been widely studied in its different subareas, including risk control (<xref ref-type="bibr" rid="B21">Pichler et&#x20;al., 2021</xref>), portfolio optimization (<xref ref-type="bibr" rid="B7">Giudici et&#x20;al., 2020</xref>), and trading strategy (<xref ref-type="bibr" rid="B13">Marques and Gomes, 2010</xref>; <xref ref-type="bibr" rid="B26">Vella and Ng, 2015</xref>; <xref ref-type="bibr" rid="B2">Chen et&#x20;al., 2021</xref>). Nowadays, the AI-based trading, especially, the reinforcement learning-approach, attracts the interest in both academia and industry. <xref ref-type="bibr" rid="B17">Moody and Saffell (2001)</xref> proposed a direct reinforcement algorithm to trade and performed a comprehensive comparison between the Q-learning with the policy gradient. <xref ref-type="bibr" rid="B9">Huang et&#x20;al. (2016)</xref> further propose a robust trading agent based on the deep-Q networks (DQN). <xref ref-type="bibr" rid="B5">Deng et&#x20;al. (2016)</xref> utilized the fuzzy logic with a deep learning model to extract the financial feature from noisy time series, which achieved state-of-the-art performance in the single-product trading. <xref ref-type="bibr" rid="B31">Xiong et&#x20;al. (2018)</xref> employed the Deep Deterministic Policy Gradient (DDPG) baesd on the standard actor-critic framework to perform the stock trading. The experiments demonstrated their profitability over the baselines including the min-variance portfolio allocation method and the technical approach based on the Dow Jones Industrial Average (DJIA) index. <xref ref-type="bibr" rid="B28">Wang et&#x20;al. (2019)</xref> employed the RL algorithm to construct the winner and loser portfolio and traded in the buy-winner-sell-loser strategy. However, the intraday trading task for reinforced trading agent are still less addressed, which is mainly because the complexity in designing trading space for frequent trading strategy. We dominantly aim at the efficient intraday trading in our research.</p>
</sec>
</sec>
<sec id="s3">
<title>3&#x20;QF-TraderNet</title>
<p>Daytrade refers to the strategy of taking a position and leaving the market within one trading day. We let our model sends an order when the market is opened every trading day. Based on the observed environment, we train QF-TraderNet to learn the optimal QPL to settle. We will introduce the QPL based action space search and model architecture separately.</p>
<sec id="s3-1">
<title>3.1 Quantum Finance Theory Based Action Space Search</title>
<p>Quantum finance theory elaborated on the relationship between the secondary financial market and the classical-quantum mechanics model (<xref ref-type="bibr" rid="B11">Lee, 2020</xref>) (<xref ref-type="bibr" rid="B14">Meng et&#x20;al., 2015</xref>) (<xref ref-type="bibr" rid="B32">Ye and Huang, 2008</xref>). QFT proposes an anharmonic oscillator model to embed the interrelationships among financial products. It considers the dynamics of the financial products are affected by the energy field generated by itself and other financial product (<xref ref-type="bibr" rid="B11">Lee, 2020</xref>). The energy levels generated from the field of particle regulate the equilibrium states of price movement on a daily basis, which is noted as the daily quantum price level (QPL). QPLs could be viewed as the support or resistance in classical financial analysis indeed. Past studies (<xref ref-type="bibr" rid="B10">Lee, 2019</xref>) have shown that QPLs can be used as feature extraction for the financial time series. The procedure of the QPL calculation is given with the following&#x20;steps.</p>
<sec id="s3-1-1">
<title>Step 1: Modeling the Potential Energy of Market Movement <italic>via</italic> Four Major Market Participants</title>
<p>Same with the classical quantum mechanics, the <italic>Hamiltonian</italic> in QFT contains the potential term and the volatility term. Founded on the conventional financial analysis, primary market participants include 1) Investor, 2) Speculator, 3) Arbitrageurs, 4) Hedger, and 5) Market maker; however, there is no available chance for Arbitrager to perform effective trading according to the efficient market hypothesis (<xref ref-type="bibr" rid="B11">Lee, 2020</xref>). Thus we ignore the arbitrageurs&#x2019; effect, and then count the impact of other participants towards the calculation of market potential term:</p>
<p>Market makers provide the facilitator services for other participants, and to absorb the outstanding demand noted as <italic>z</italic>
<sub>
<italic>&#x3c3;</italic>
</sub>, with absorbability factors <italic>&#x3b1;</italic>
<sub>
<italic>&#x3c3;</italic>
</sub>. Thus, the excess demand at any instance is given by &#x394;<italic>z</italic>&#x20;&#x3d; <italic>z</italic>
<sub>&#x2b;</sub> &#x2212; <italic>z</italic>
<sub>&#x2212;</sub>. The relationship between instantaneous returns <inline-formula id="inf1">
<mml:math id="m1">
<mml:mi>r</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>r</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="normal">&#x394;</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="normal">&#x394;</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="normal">&#x394;</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:math>
</inline-formula>, and the excess demand could be approximately noted as <inline-formula id="inf2">
<mml:math id="m2">
<mml:mi>r</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="normal">&#x394;</mml:mi>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b3;</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:math>
</inline-formula>, in which <italic>&#x3b3;</italic> represents the market depth. For an efficient market with the smooth market environment, we assume the absorbability of existing orders with different trading directions will be the same, and the contribution of the market makers is derived as (<xref ref-type="bibr" rid="B11">Lee, 2020</xref>),<disp-formula id="e1">
<mml:math id="m3">
<mml:mfrac>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi mathvariant="normal">&#x394;</mml:mi>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:msub>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>M</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:msub>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>M</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:msub>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>M</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
<label>(1)</label>
</disp-formula>
<disp-formula id="e2">
<mml:math id="m4">
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3b3;</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>M</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
<label>(2)</label>
</disp-formula>where <italic>&#x3c3;</italic> denotes the trading position including &#x2b;: long position, and -: short position. <italic>r</italic>
<sub>
<italic>t</italic>
</sub> denotes the simultaneous price return respect to time&#x20;t.</p>
<p>Speculators are trend-following participants with few senses about risk control. Their behavior mainly contributes to the market movement by its dynamic oscillator term. A damping variable <italic>&#x3b4;</italic> is defined to represent the resistance of trend followers behaviors towards the market. Considering that speculators have less consider risk, there is no high-order anharmonic term regarding the market volatility,<disp-formula id="e3">
<mml:math id="m5">
<mml:mfrac>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi mathvariant="normal">&#x394;</mml:mi>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:msub>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi>&#x3b4;</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
<label>(3)</label>
</disp-formula>
</p>
<p>Investors have a sense of stopping loss. They are 1) earning profit following the trend, 2) minimizing the risk; thus, we define their potential energy by,<disp-formula id="e4">
<mml:math id="m6">
<mml:mfrac>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi mathvariant="normal">&#x394;</mml:mi>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:msub>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>I</mml:mi>
<mml:mi>V</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>I</mml:mi>
<mml:mi>V</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>v</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>I</mml:mi>
<mml:mi>V</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(4)</label>
</disp-formula>where <italic>&#x3b4;</italic>, <italic>v</italic> stand for the harmonic dynamic term (trend following contribution); and anharmonic term (market volatility), respectively.</p>
<p>Hedger also controls the risk but using sophisticated hedging techniques. Commonly, the reverse trading direction has been performed by Hedgers compared with common Investors, especially for the one-product hedging strategy. Hence, the market dynamic caused by Hedger could be summarized as,<disp-formula id="e5">
<mml:math id="m7">
<mml:mfrac>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi mathvariant="normal">&#x394;</mml:mi>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:msub>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>v</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
<label>(5)</label>
</disp-formula>
</p>
<p>To conclude the equations (3.1) from to (3.4), the simultaneous price return dr/dt could be rewritten as,<disp-formula id="e6">
<mml:math id="m8">
<mml:mfrac>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3b3;</mml:mi>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:mfrac>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi mathvariant="normal">&#x394;</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3b3;</mml:mi>
<mml:mi>&#x3b4;</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b3;</mml:mi>
<mml:mi>v</mml:mi>
<mml:msubsup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
<label>(6)</label>
</disp-formula>where <italic>P</italic> denotes the number of types of participants inside markets. <italic>&#x3b4;</italic>, and <italic>v</italic> in <xref ref-type="disp-formula" rid="e5">Eq. 5</xref> are the summary of each term across all participants models, i.e.,&#x20;<italic>&#x3b4;</italic> &#x3d; <italic>&#x3b3;&#x3b1;</italic>
<sub>
<italic>MM</italic>
</sub> &#x2b; <italic>&#x3b4;</italic>
<sub>
<italic>SP</italic>
</sub> &#x2b; <italic>&#x3b4;</italic>
<sub>
<italic>HG</italic>
</sub> &#x2212; <italic>&#x3b4;</italic>
<sub>
<italic>IV</italic>
</sub>, and <italic>v</italic>&#x20;&#x3d; <italic>v</italic>
<sub>
<italic>HG</italic>
</sub> &#x2212; <italic>v</italic>
<sub>
<italic>IV</italic>
</sub>. Combining <italic>dr</italic>/<italic>dt</italic> with the Brownian price returns described by the Langevin equation, the instantaneous potential energy is modeled with the following equation,<disp-formula id="e7">
<mml:math id="m9">
<mml:mi>V</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x222b;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>&#x3b3;</mml:mi>
<mml:mi>&#x3b7;</mml:mi>
<mml:mi>&#x3b4;</mml:mi>
<mml:mi>r</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3b3;</mml:mi>
<mml:mi>&#x3b7;</mml:mi>
<mml:mi>v</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mi>d</mml:mi>
<mml:mi>r</mml:mi>
<mml:mo>&#x2248;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>&#x3b3;</mml:mi>
<mml:mi>&#x3b7;</mml:mi>
<mml:mi>&#x3b4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:msup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>&#x3b3;</mml:mi>
<mml:mi>&#x3b7;</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:msup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:msup>
</mml:math>
<label>(7)</label>
</disp-formula>where <italic>&#x3b7;</italic> is the damping force factor of the market.</p>
</sec>
<sec id="s3-1-2">
<title>Step 2: Modeling the Kinetic Term of Market Movement <italic>via</italic> Price Return</title>
<p>One challenge to model the kinetic term is to replace the displacements in classical particles with an appropriate measurement in finance. Specifically, we replace displacement with price returns <italic>r</italic>(<italic>t</italic>), as <italic>r</italic>(<italic>t</italic>) connects the price change with time unit, which simplifies the Schr&#xf6;dinger equation into the Non-time-dependent one. Hence, the <italic>Hamiltonian</italic> for financial particle could be formulated by,<disp-formula id="e8">
<mml:math id="m10">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>H</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>&#x210f;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mfrac>
<mml:mrow>
<mml:mi>&#x2202;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x2202;</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>V</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(8)</label>
</disp-formula>where <italic>&#x210f;</italic>, <italic>m</italic> denote the plank constant and intrinsic properties of the financial market, such as market capitalization in a stock market. Combining the <italic>Hamiltonian</italic> with the classical Schr&#xf6;dinger equation, the Schr&#xf6;dinger Equation for Quantum Finance Theory (QFSE) comes out with (<xref ref-type="bibr" rid="B11">Lee, 2020</xref>),<disp-formula id="e9">
<mml:math id="m11">
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mi>&#x210f;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2b;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mi>&#x3b3;</mml:mi>
<mml:mi>&#x3b7;</mml:mi>
<mml:mi>&#x3b4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:msup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>&#x3b3;</mml:mi>
<mml:mi>&#x3b7;</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:msup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mi>&#x3d5;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>E</mml:mi>
<mml:mi>&#x3d5;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(9)</label>
</disp-formula>
</p>
<p>
<italic>E</italic> denotes the particle&#x2019;s energy levels, which refers to the Quantum Price Levels for the financial particles. The first term <inline-formula id="inf3">
<mml:math id="m12">
<mml:mfrac>
<mml:mrow>
<mml:mi>&#x210f;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:math>
</inline-formula> is the kinetic energy term. The second term <italic>V</italic>(<italic>r</italic>) represents the potential energy term, i.e. (3.6), of the quantum finance market. <italic>&#x3d5;</italic>(<italic>r</italic>) is the wave-function of QFSE, which is approximated by the probability density function of historical price return.</p>
</sec>
<sec id="s3-1-3">
<title>Step 3: Perform the Action Space Search by Solving the QFSE</title>
<p>According to QFT, if there were no extrinsic incentives such as financial events or the release of critical financial figures, QFPs would remain at their energy levels (i.e.,&#x20;equilibrium states) and perform regular oscillations. If there is an external stimulus, QFPs would absorb or release the quantized energy and jump to other QPLs. Thus, daily QPLs could be viewed as the potential states of the price movements in one trading day. Hence, we employ QPLs as the action candidates in the action space <inline-formula id="inf4">
<mml:math id="m13">
<mml:mi>A</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula> of QF-TraderNet. The detailed numerical method for solving QFSE and the algorithm for the QPL based action space search is given in the supplementary&#x20;file.</p>
</sec>
</sec>
<sec id="s3-2">
<title>3.2 Deep Feature Learning and Representation by LSTM Networks</title>
<p>LSTM networks show promising performance in the sequential feature learning, as its structural adaptability (<xref ref-type="bibr" rid="B6">Gers et&#x20;al., 2000</xref>). We introduce the LSTM networks to extract the temporal features of the financial series, thus improving the perception in the market status of the policy generation network (PGN).</p>
<p>We use the same <italic>look-back window</italic> in (<xref ref-type="bibr" rid="B28">Wang et&#x20;al., 2019</xref>) with size <italic>W</italic> to split the input sequence <bold>
<italic>x</italic>
</bold> from the completed series <inline-formula id="inf5">
<mml:math id="m14">
<mml:mi mathvariant="bold-italic">S</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">1</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">2</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">T</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula>, i.e.,&#x20;agent evaluates the market status by the time period with size <italic>W</italic>. Hence, the input matrix of LSTM could be noted as <inline-formula id="inf6">
<mml:math id="m15">
<mml:mi mathvariant="bold-italic">X</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">1</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">2</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">T</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="bold-italic">W</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn mathvariant="bold-italic">1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula>, where <inline-formula id="inf7">
<mml:math id="m16">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="bold-italic">W</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>w</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula>. We design our input vectors <bold>
<italic>s</italic>
</bold>
<sub>
<bold>
<italic>t</italic>
</bold>
</sub> is constituted by: 1) <italic>Opening, highest, lowest and closing prices</italic> for each trading day. Note: the close price in <italic>t</italic>&#x20;&#x2212; 1 day might be different with the open price in <italic>t</italic> because of the adjustment of the market outside the trading hours; hence, we consider the entire price variables with four types. 2) <italic>Transaction Volume.</italic> 3) <italic>Moving Average Convergence-Divergence</italic> is a technical indicator to identify the market status. 4) <italic>Relative strength index</italic> is a technical indicator measuring the price momentum. 5) <italic>Bollinger Band (main, upper, and lower)</italic> can be applied to identify the potential price range, consequently observing the market trend (<xref ref-type="bibr" rid="B3">Colby and Meyers, 1988</xref>). 6) <italic>KDJ (stochastic oscillator)</italic> is used in short-term oriented trading by the price velocity techniques (<xref ref-type="bibr" rid="B3">Colby and Meyers, 1988</xref>).</p>
<p>The principal components analysis (PCA) (<xref ref-type="bibr" rid="B30">Wold et&#x20;al., 1987</xref>) is utilized to compress the series data <bold>
<italic>S</italic>
</bold> into <inline-formula id="inf8">
<mml:math id="m17">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> dimension and denoise (<xref ref-type="bibr" rid="B30">Wold et&#x20;al., 1987</xref>). Subsequently, the L2 normalization is applied to scale the input features to be in the same magnitude. The preprocessing is calculated as,<disp-formula id="e10">
<mml:math id="m18">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold-italic">X</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:munder>
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mi>C</mml:mi>
<mml:mi>A</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x2192;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
</mml:munder>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi mathvariant="bold-italic">X</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mi>C</mml:mi>
<mml:mi>A</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x2192;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
</mml:munder>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi mathvariant="bold-italic">X</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
</mml:math>
<label>(10)</label>
</disp-formula>, where <inline-formula id="inf9">
<mml:math id="m19">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>F</mml:mi>
</mml:math>
</inline-formula>, and the deep feature learning model could be described as,<disp-formula id="e11">
<mml:math id="m20">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>T</mml:mi>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3be;</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>T</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>W</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(11)</label>
</disp-formula>where <bold>
<italic>&#x3be;</italic>
</bold> is the trainable parameters for&#x20;LSTM.</p>
</sec>
<sec id="s3-3">
<title>3.3 Policy Generator Networks (PGN)</title>
<p>Given the learned feature vector <bold>
<italic>h</italic>
</bold>
<sub>
<bold>
<italic>t</italic>
</bold>
</sub>, PGN directly produces the output policy, i.e.,&#x20;the probability of settling order in each &#x2b; QPL and -QPL, according to the action score <inline-formula id="inf10">
<mml:math id="m21">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> produced by a fully-connected networks (FFBPN).<disp-formula id="e12">
<mml:math id="m22">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>F</mml:mi>
<mml:mi>B</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(12)</label>
</disp-formula>where <bold>
<italic>&#x3b8;</italic>
</bold> deontes the parameters of FFBPN, with the weighted matrix <bold>
<italic>W</italic>
</bold>
<sub>
<bold>
<italic>&#x3b8;</italic>
</bold>
</sub> and bias <bold>
<italic>b</italic>
</bold>
<sub>
<bold>
<italic>&#x3b8;</italic>
</bold>
</sub>. Let <inline-formula id="inf11">
<mml:math id="m23">
<mml:msubsup>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> denotes <italic>i</italic>&#x20;&#x2212; <italic>th</italic> action at time <italic>t</italic>. The output policy <bold>
<italic>a</italic>
</bold>
<sub>
<bold>
<italic>t</italic>
</bold>
</sub> is calculated as,<disp-formula id="e13">
<mml:math id="m24">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>e</mml:mi>
<mml:mi>x</mml:mi>
<mml:mi>p</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2208;</mml:mo>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msub>
<mml:mi>e</mml:mi>
<mml:mi>x</mml:mi>
<mml:mi>p</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfrac>
</mml:math>
<label>(13)</label>
</disp-formula>in timestep <italic>t</italic>, model takes action <italic>a</italic>
<sub>
<italic>t</italic>
</sub> by sampling from the policy <inline-formula id="inf12">
<mml:math id="m25">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> comprised of long (&#x2b;) and short (-) trading direction. <inline-formula id="inf13">
<mml:math id="m26">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> contains <italic>A</italic> dimensions, indicating the number of candidate actions, with the reward of price return <inline-formula id="inf14">
<mml:math id="m27">
<mml:msubsup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> for each,<disp-formula id="e14">
<mml:math id="m28">
<mml:msubsup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="matrix">
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mi>&#x3b4;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>Q</mml:mi>
<mml:mi>P</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mo>,</mml:mo>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mo>&#x2200;</mml:mo>
<mml:mi>Q</mml:mi>
<mml:mi>P</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2208;</mml:mo>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mi>&#x3b4;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mo>,</mml:mo>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mo>&#x2200;</mml:mo>
<mml:mi>Q</mml:mi>
<mml:mi>P</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2209;</mml:mo>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(14)</label>
</disp-formula>where <italic>&#x3b4;</italic> denotes the trading direction: for actions with &#x2b;QPL as the target price level to settle, the trading will be determined as long buy (<italic>&#x3b4;</italic> &#x3d; &#x2b; 1); for the actions in -QPL, short sell (<italic>&#x3b4;</italic> &#x3d; &#x2212; 1) trading will be performed; and <italic>&#x3b4;</italic> is 0 when the decision is made to be neutral, as no trading will be made in <italic>t</italic> trading&#x20;day.</p>
<p>We train our QF-TraderNet with reinforcement learning. The key idea is to maintain a loop with the successive steps: 1) agent <italic>&#x3c0;</italic> aware the environment, 2) <italic>&#x3c0;</italic> make the action, and 3) adjust its behavior to receive more reward until the agent has received its learning goal (<xref ref-type="bibr" rid="B23">Sutton and Barto, 2018</xref>). Therefore, for each training episode, a trajectory <inline-formula id="inf15">
<mml:math id="m29">
<mml:mi mathvariant="bold-italic">&#x3c4;</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">1</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">1</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">2</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold">1</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn mathvariant="bold-italic">1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">T</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula> could be defined as the sequence of state-action tuple, with the corresponding return sequence<xref ref-type="fn" rid="fn1">
<sup>1</sup>
</xref>
<inline-formula id="inf16">
<mml:math id="m30">
<mml:mi mathvariant="bold-italic">r</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula>. The probability of action <italic>Pr</italic> (<italic>action</italic>
<sub>
<italic>t</italic>
</sub> &#x3d; <italic>i</italic>) for each QPL is determined by QF-TraderNet as:<disp-formula id="e15">
<mml:math id="m31">
<mml:msubsup>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>P</mml:mi>
<mml:mi>r</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>Q</mml:mi>
<mml:mi>P</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold-italic">X</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3be;</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(15)</label>
</disp-formula>
<disp-formula id="e16">
<mml:math id="m32">
<mml:msub>
<mml:mrow>
<mml:mfenced open="" close="|">
<mml:mrow>
<mml:mtable class="matrix">
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mo>&#x3d;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c0;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mi>G</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:munder>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>T</mml:mi>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3be;</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
<label>(16)</label>
</disp-formula>let <italic>R</italic>
<sub>
<bold>
<italic>&#x3c4;</italic>
</bold>
</sub> denotes the cumulative price return for trajectory <bold>
<italic>&#x3c4;</italic>
</bold>, with <inline-formula id="inf17">
<mml:math id="m33">
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>W</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3c4;</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>. Then, for all possible explored trajectories, the expectation reward obtained by the RL agent could be evaluated as (<xref ref-type="bibr" rid="B24">Sutton et&#x20;al., 2000</xref>),<disp-formula id="e17">
<mml:math id="m34">
<mml:msub>
<mml:mrow>
<mml:mi>J</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3c0;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3be;</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mo>&#x222b;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3c4;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3c4;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:munder>
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mi>r</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3c4;</mml:mi>
<mml:mo>;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3be;</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3c0;</mml:mi>
</mml:mrow>
</mml:munder>
<mml:mi>d</mml:mi>
<mml:mi mathvariant="bold-italic">&#x3c4;</mml:mi>
</mml:math>
<label>(17)</label>
</disp-formula>where <inline-formula id="inf18">
<mml:math id="m35">
<mml:munder>
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mi>r</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3c4;</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3b8;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3be;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3c0;</mml:mi>
</mml:mrow>
</mml:munder>
</mml:math>
</inline-formula> is the probability for QF-TraderNet agent <italic>&#x3c0;</italic> with parameters <bold>
<italic>&#x3b8;</italic>
</bold> and <bold>
<italic>&#x3be;</italic>
</bold> to generate trajectory <bold>
<italic>&#x3c4;</italic>
</bold> with Monte-Carlo Simulation. Then, the objective is to maximize the expectation of reward, <bold>
<italic>&#x3b8;</italic>
</bold>&#x2a;, <bold>
<italic>&#x3be;</italic>
</bold>&#x2a; &#x3d; <italic>argmax</italic>
<sub>
<bold>
<italic>&#x3b8;,&#x3be;</italic>
</bold>
</sub>
<italic>J</italic> (<bold>
<italic>&#x3b8;,&#x3be;</italic>
</bold>). We substitute objective with its inverse to and use gradient descent to optimize. To avoid the local minimum probelm caused by the multiple postive-reward actions, we use the state-dependent threshold method (<xref ref-type="bibr" rid="B23">Sutton and Barto, 2018</xref>) to allow the RL agent perform a more efficient optimization. The detailed gradient calculation is given in the supplementary.</p>
</sec>
<sec id="s3-4">
<title>3.4 Trading Policy With Learnable Soft Profit and Loss Control</title>
<p>In QF-TraderNet, the LSTM networks learn the hidden representation and feed it into PGN; then PGN generates the learned policy to decide the target QPL to settle. As the action is sampled from the generated policy, QF-TraderNet adopts a soft profit-and-loss control strategy rather than the deterministic TP and SL. The overall summary of QF-TraderNet architecture has been shown in <xref ref-type="fig" rid="F3">Figure&#x20;3</xref>.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>The RL framework for the QF-TraderNet.</p>
</caption>
<graphic xlink:href="frai-04-749878-g003.tif"/>
</fig>
<p>An equivalent way to interpret our strategy is that our model trades with long buy if the decision is made in positive QPL. In reverse, short sell transactions will be delivered. Once the trading direction is decided, the target QPL with the maximum probability will be considered as the soft target price (S-TP), and the soft stop loss line will be the QPL with the highest probability in the opposite trading direction. One exemplification is presented in <xref ref-type="fig" rid="F4">Figure&#x20;4</xref>.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>A case study illustrates our profit-and-loss control strategy. The trading policy is uniformly distributed initially. Ideally, our model assigns the &#x2b;3 QPL action which earns the maximum profit with the largest probability as S-TP. On the short side, &#x2212;1 QPL can take the most considerable reward, leading to being accredited the maximum probability as S-SL.</p>
</caption>
<graphic xlink:href="frai-04-749878-g004.tif"/>
</fig>
<p>Since the S-TP and S-SL control is probability-based, when the price touches the stop loss line prematurely, QF-TraderNet will not be forced to do the settlement. It will think whether there is a better target price for settlement in the entire action space. Therefore, the model is more flexible for the SL and TP control in different states, compared with using a couple of preset &#x201c;hard&#x201d; hyperparameters.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Experiments</title>
<p>We conduct the empirical evaluation for our QF-TraderNet in various types of financial datasets. In our experiment, eight datasets from 4 categories are used, including 1) foreign exchange product: Great Britain Pounds vs. United&#x20;States Dollar (GBPUSD), Australian Dollar vs. United&#x20;States Dollar (AUDUSD), Euro vs. United&#x20;States Dollar (EURUSD), United&#x20;States Dollar vs. Swiss Franc (USDCHF); 2) financial indices: S&#x26;P 500 Index (S&#x26;P500), Hang Seng Index (HSI); 3) Metal: Silver vs. United&#x20;States Dollar (XAGUSD), and 4) Crude oil: Oil vs. United&#x20;States Dollar (OILUSe). The evaluation is conducted from the perspective of earning profits; and the robustness when agents face the unexpected change of market states. We also investigate the impact of different settings of our proposed QPL based action space search for RL trader, and the ablation study of our&#x20;model.</p>
<sec id="s4-1">
<title>4.1 Experiment Settings</title>
<p>All datasets utilized in experiments are fetched from the free and opened historical data center in <italic>MetaTrader 4</italic>, which is a professional trading platform for the FOREX, financial indices, and other securities. We download the raw time series data, around 2048 trading days, and we split the 90<italic>%</italic> front of data for training and validation. The rest will be utilized as out-of-sample verification, i.e.,&#x20;the continuous series from November 2012 to July 2019, has been spliced to construct the sequential training sample; the rest part is applied as testing and validation. To be noticed, the evaluation period has covered the recent fluctuations in the global financial market caused by the COVID-19 pandemic, which could be utilized as the robustness test when the trading agent is handling the unforeseen market fluctuations. The size of <italic>look-back window</italic> is set at 3, and the metrics regarding price return and Sharpe ratio is daily calculated. In the backtest, initial capital is set to the corresponding currency or asset with a value of 10,000, at a transaction cost with 0.3<italic>%</italic> (<xref ref-type="bibr" rid="B5">Deng et&#x20;al., 2016</xref>). All the experiments are conducted in the single NVIDIA GTX Titan X&#x20;GPU.</p>
</sec>
<sec id="s4-2">
<title>4.2 Models Settings</title>
<p>To compare our model with the traditional methods, we select the forecasting based trading model and other state-of-the-art reinforcement learning-based trading agents as the baseline.<list list-type="simple">
<list-item>
<p>&#x2022; <italic>Market baseline</italic> (<xref ref-type="bibr" rid="B9">Huang et&#x20;al., 2016</xref>). This strategy is used to measure the overall performance of the market during this period <italic>T</italic>, by holding the product consistently.</p>
</list-item>
<list-item>
<p>&#x2022; <italic>DDR-RNN.</italic> Following the idea of Deep Direct Reinforcement, but we apply the principal component analysis (PCA) to denoise and composes data. We also employ RNN to learn the features, and a two-layer FFBPN as the policy generator rather than the logistic regression in original design. This model can be regarded as the ablation study of QF-TraderNet without the QPL action space search.</p>
</list-item>
<list-item>
<p>&#x2022; <italic>FCM</italic>, a forecasting model based on RNN trend predictor, consisting of a 7-layer LSTM with 512 hidden dimensions. It trades with a Buy-Winner-Sell-Loser strategy.</p>
</list-item>
<list-item>
<p>&#x2022; <italic>RF.</italic> Same design with FCM but predict the trend <italic>via</italic> Random Forest.</p>
</list-item>
<list-item>
<p>&#x2022; <italic>QF-PGN.</italic> QF-PGN is the policy gradient based RL agent with QPL based order control. Single FFBPN is utilized as the policy generator with 3 ReLU layers, and 128 neurons per layer. This model could be admitted as our model without the deep feature representation&#x20;block.</p>
</list-item>
<list-item>
<p>&#x2022; <italic>FDRNN</italic> (<xref ref-type="bibr" rid="B5">Deng et&#x20;al., 2016</xref>). A state-of-the-art direct reinforcement RL trader following the one-product trading, by using the fuzzy representation and deep autoencoder to extract the features.</p>
</list-item>
</list>
</p>
<p>We implement two versions of QF-TraderNet: 1) <italic>QF-TraderNet Lite (QFTN-L)</italic>: 2 layers LSTM with 128-dimensional hidden vector as the feature representation, and 3 layers of policy generator network with 128, 64, 32 neurons per each. The size of action space is 3.2) <italic>QF-TraderNet Ultra (QFTN-U)</italic>: Same architecture with the Lite, but the number of candidate actions is enlarged to&#x20;7.</p>
<p>Regarding the training settings, the Adaptive Moment Estimation (ADAM) optimizer with 1,500 training epochs is used for all iterative optimization models at a 0.001 learning rate. For the algorithms requiring PCA, the target dimensions <inline-formula id="inf19">
<mml:math id="m36">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> is set at 4, satisfying the composes matrix has embedded 99.5<italic>%</italic> of the interrelationship of features. In the practical implementation, we directly utilize the four prices as the input for USDCHF, S&#x26;P500, XAGUSD, and OILUSe; the normalization step is not performed for the HSI and OILUSe. The reason is that our experimental results show our model can perceive the market state good enough in these settings. For the sake of computational complexity, we remove the extra input features.</p>
</sec>
<sec id="s4-3">
<title>4.3 Performance in 8 Financial Datasets</title>
<p>As displayed in <xref ref-type="fig" rid="F5">Figure&#x20;5</xref> and <xref ref-type="table" rid="T1">Table&#x20;1</xref>, we present the evaluation of each trading system&#x2019;s profitability in 8 datasets, with the metrics of cumulative price return (CPR) and the Sharpe ratio (SR). The CPR is formulated with,<disp-formula id="e18">
<mml:math id="m37">
<mml:mi>C</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>R</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>h</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(18)</label>
</disp-formula>and the Sharpe ratio is calculated by:<disp-formula id="e19">
<mml:math id="m38">
<mml:mi>S</mml:mi>
<mml:mi>R</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>v</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>e</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi mathvariant="bold-italic">C</mml:mi>
<mml:mi mathvariant="bold-italic">P</mml:mi>
<mml:mi mathvariant="bold-italic">R</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>v</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi mathvariant="bold-italic">C</mml:mi>
<mml:mi mathvariant="bold-italic">P</mml:mi>
<mml:mi mathvariant="bold-italic">R</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfrac>
</mml:math>
<label>(19)</label>
</disp-formula>
</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>1st panel: Continuous partition for the training and verification data; 2nd panel: Affected by the global economic situation, most datasets showed a downward trend at the testing interval, accompanied by highly irregular oscillations; the 3rd panel: cumulative reward curve for different methods in testing evaluation.</p>
</caption>
<graphic xlink:href="frai-04-749878-g005.tif"/>
</fig>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Summary of the main comparison results among all models.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Models</th>
<th colspan="2" align="center">HSI</th>
<th colspan="2" align="center">S&#x26;P500</th>
<th colspan="2" align="center">Silver</th>
<th colspan="2" align="center">Crude oil</th>
<th colspan="2" align="center">USDCHF</th>
<th colspan="2" align="center">GBPUSD</th>
<th colspan="2" align="center">EURUSD</th>
<th colspan="2" align="center">AUDUSD</th>
</tr>
<tr>
<th align="center">CPR</th>
<th align="center">SR</th>
<th align="center">CPR</th>
<th align="center">SR</th>
<th align="center">CPR</th>
<th align="center">SR</th>
<th align="center">CPR</th>
<th align="center">SR</th>
<th align="center">CPR</th>
<th align="center">SR</th>
<th align="center">CPR</th>
<th align="center">SR</th>
<th align="center">CPR</th>
<th align="center">SR</th>
<th align="center">CPR</th>
<th align="center">SR</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Market</td>
<td align="char" char=".">555.00</td>
<td align="char" char=".">0.01</td>
<td align="char" char=".">2,122.27</td>
<td align="char" char=".">0.05</td>
<td align="char" char=".">&#x2212;12.66</td>
<td align="char" char=".">&#x2212;0.03</td>
<td align="char" char=".">&#x2212;90.79</td>
<td align="char" char=".">&#x2212;0.04</td>
<td align="char" char=".">0.19</td>
<td align="char" char=".">0.02</td>
<td align="char" char=".">&#x2212;0.07</td>
<td align="char" char=".">&#x2212;0.01</td>
<td align="char" char=".">&#x2212;0.05</td>
<td align="char" char=".">&#x2212;0.01</td>
<td align="char" char=".">&#x2212;0.04</td>
<td align="char" char=".">&#x2212;0.01</td>
</tr>
<tr>
<td align="left">RNN-FCM</td>
<td align="char" char=".">1,251.78</td>
<td align="char" char=".">0.03</td>
<td align="char" char=".">361.94</td>
<td align="char" char=".">0.09</td>
<td align="char" char=".">-11.67</td>
<td align="char" char=".">-0.07</td>
<td align="char" char=".">6.76</td>
<td align="char" char=".">0.02</td>
<td align="char" char=".">0.04</td>
<td align="char" char=".">0.02</td>
<td align="char" char=".">&#x2212;0.14</td>
<td align="char" char=".">&#x2212;0.07</td>
<td align="char" char=".">&#x2212;0.24</td>
<td align="char" char=".">&#x2212;0.13</td>
<td align="char" char=".">0.05</td>
<td align="char" char=".">0.04</td>
</tr>
<tr>
<td align="left">RF-FCM</td>
<td align="char" char=".">3,846.31</td>
<td align="char" char=".">0.09</td>
<td align="char" char=".">336.27</td>
<td align="char" char=".">0.06</td>
<td align="char" char=".">23.60</td>
<td align="char" char=".">0.91</td>
<td align="char" char=".">112.88</td>
<td align="char" char=".">1.13</td>
<td align="char" char=".">0.11</td>
<td align="char" char=".">0.16</td>
<td align="char" char=".">0.29</td>
<td align="char" char=".">0.33</td>
<td align="char" char=".">0.20</td>
<td align="char" char=".">0.53</td>
<td align="char" char=".">&#x2212;0.04</td>
<td align="char" char=".">&#x2212;0.07</td>
</tr>
<tr>
<td align="left">DDR-RNN</td>
<td align="char" char=".">4,505.00</td>
<td align="char" char=".">0.10</td>
<td align="char" char=".">345.50</td>
<td align="char" char=".">0.03</td>
<td align="char" char=".">1.53</td>
<td align="char" char=".">0.02</td>
<td align="char" char=".">&#x2212;4.57</td>
<td align="char" char=".">&#x2212;0.02</td>
<td align="char" char=".">0.07</td>
<td align="char" char=".">0.09</td>
<td align="char" char=".">-0.02</td>
<td align="char" char=".">-0.02</td>
<td align="char" char=".">0.08</td>
<td align="char" char=".">0.15</td>
<td align="char" char=".">&#x3c;0.01</td>
<td align="char" char=".">&#x2212;0.08</td>
</tr>
<tr>
<td align="left">FDRNN</td>
<td align="char" char=".">1,536.00</td>
<td align="char" char=".">0.04</td>
<td align="char" char=".">731.73</td>
<td align="char" char=".">0.07</td>
<td align="char" char=".">2.80</td>
<td align="char" char=".">0.04</td>
<td align="char" char=".">&#x2212;9.38</td>
<td align="char" char=".">&#x2212;0.03</td>
<td align="char" char=".">0.08</td>
<td align="char" char=".">0.10</td>
<td align="char" char=".">0.05</td>
<td align="char" char=".">0.04</td>
<td align="char" char=".">&#x2212;0.08</td>
<td align="char" char=".">&#x2212;0.10</td>
<td align="char" char=".">0.05</td>
<td align="char" char=".">0.12</td>
</tr>
<tr>
<td align="left">QF-PGN</td>
<td align="char" char=".">3,244.35</td>
<td align="char" char=".">0.07</td>
<td align="char" char=".">3,133.76</td>
<td align="char" char=".">
<bold>1.88</bold>
</td>
<td align="char" char=".">1.94</td>
<td align="char" char=".">0.05</td>
<td align="char" char=".">138.34</td>
<td align="char" char=".">
<bold>2.00</bold>
</td>
<td align="char" char=".">&#x2212;0.08</td>
<td align="char" char=".">&#x2212;0.11</td>
<td align="char" char=".">0.28</td>
<td align="char" char=".">0.37</td>
<td align="char" char=".">&#x2212;0.03</td>
<td align="char" char=".">&#x2212;0.05</td>
<td align="char" char=".">0.17</td>
<td align="char" char=".">0.50</td>
</tr>
<tr>
<td align="left">QF-TraderNet Lite</td>
<td align="char" char=".">2,779.64</td>
<td align="char" char=".">
<bold>0.17</bold>
</td>
<td align="char" char=".">155.66</td>
<td align="char" char=".">0.04</td>
<td align="char" char=".">1.56</td>
<td align="char" char=".">0.04</td>
<td align="char" char=".">82.40</td>
<td align="char" char=".">0.54</td>
<td align="char" char=".">
<bold>0.58</bold>
</td>
<td align="char" char=".">
<bold>1.69</bold>
</td>
<td align="char" char=".">
<bold>0.61</bold>
</td>
<td align="char" char=".">
<bold>1.31</bold>
</td>
<td align="char" char=".">0.20</td>
<td align="char" char=".">
<bold>0.65</bold>
</td>
<td align="char" char=".">0.02</td>
<td align="char" char=".">0.03</td>
</tr>
<tr>
<td align="left">QF-TraderNet Ultra</td>
<td align="char" char=".">
<bold>8,100.51</bold>
</td>
<td align="char" char=".">
<bold>0.17</bold>
</td>
<td align="char" char=".">
<bold>4,428.00</bold>
</td>
<td align="char" char=".">1.52</td>
<td align="char" char=".">
<bold>31.24</bold>
</td>
<td align="char" char=".">
<bold>1.49</bold>
</td>
<td align="char" char=".">
<bold>164.38</bold>
</td>
<td align="char" char=".">1.44</td>
<td align="char" char=".">
<bold>0.64</bold>
</td>
<td align="char" char=".">
<bold>1.16</bold>
</td>
<td align="char" char=".">
<bold>0.92</bold>
</td>
<td align="char" char=".">
<bold>1.31</bold>
</td>
<td align="char" char=".">
<bold>0.57</bold>
</td>
<td align="char" char=".">
<bold>1.11</bold>
</td>
<td align="char" char=".">
<bold>0.36</bold>
</td>
<td align="char" char=".">
<bold>0.97</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Bold values indicating the best performance in terms of corresponding metrics.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>The result of <sc>Market</sc> denotes that the market is in a downtrend with high volatility in the evaluating interval, due to the recent global economic fluctuation. The price range in testing is not fully covered in training data in some datasets (crude oil and AUDUSD), which tests the models in an unseen environment. Under these testing conditions, our QFTN-U trained with CPR achieves higher CPR and SR than other comparisons, except the SR in S&#x26;P500 and Crude Oil. QFTN-L is also comparable to the baselines. It signifies the profitability and robustness of our QF-TraderNet.</p>
<p>Moreover, QFTN-L, QFTN-U, and the PGN models yield significantly higher CPR and SR than other RL traders without QPL-based actions (DDR-RNN and FDRNN). The ablation study in <xref ref-type="table" rid="T2">Table&#x20;2</xref> also presents the contribution of each component in detail (<sc>Supervised</sc> counts from the average of <sc>Rf</sc> and <sc>Fcm</sc>), where the QPL actions dramatically contribute to the Sharpe Ratio of our full model. These demonstrates the benefit of trading with QPL to gain considerable profitability and efficient risk-control ability.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Ablation study for QF-TraderNet.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Models</th>
<th align="center">Avg. Sharpe%</th>
<th align="center">Impact</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Full Model</td>
<td align="char" char=".">1.15</td>
<td align="char" char=".">&#x2212;</td>
</tr>
<tr>
<td align="left">QFTN-L: Limit <italic>A</italic> to 3</td>
<td align="char" char=".">0.56</td>
<td align="char" char=".">&#x2212;0.59 (&#x2212;51%)</td>
</tr>
<tr>
<td align="left">PGN: - without LSTM</td>
<td align="char" char=".">0.59</td>
<td align="char" char=".">&#x2212;0.56 (&#x2212;49%)</td>
</tr>
<tr>
<td align="left">DDR-RNN: - without QPL</td>
<td align="char" char=".">0.03</td>
<td align="char" char=".">&#x2212;1.12 (&#x2212;97%)</td>
</tr>
<tr>
<td align="left">Supervised: - without RL</td>
<td align="char" char=".">0.19</td>
<td align="char" char=".">&#x2212;0.96 (&#x2212;83%)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The backtesting results in <xref ref-type="table" rid="T3">Table&#x20;3</xref> shows the good generalization of the QFTN-U. It is the only strategy for earning a positive profit on almost all datasets, which is because the day-trading strategy are less affected by the market trend, compared with other strategies in long, neutral, and short setting. We also find that the performance of our model in FOREX datasets is significantly better than others. FOREX contains more noise and fluctuations, which indicates the advantages of our models in highly fluctuated products.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Summary for net profit in the backtesting.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left"/>
<th align="center">USDCHF</th>
<th align="center">HSI</th>
<th align="center">S&#x26;P500</th>
<th align="center">XAGUSD</th>
<th align="center">GBPUSD</th>
<th align="center">EURUSD</th>
<th align="center">AUDUSD</th>
<th align="center">OILUSe</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Market</td>
<td align="char" char=".">&#x2212;156.43</td>
<td align="char" char=".">&#x2212;1,505.9</td>
<td align="char" char=".">&#x2212;175</td>
<td align="char" char=".">19.29</td>
<td align="char" char=".">&#x2212;28.07</td>
<td align="char" char=".">&#x2212;214.95</td>
<td align="char" char=".">&#x2212;477.53</td>
<td align="char" char=".">&#x2212;7,228.5</td>
</tr>
<tr>
<td align="left">FCM</td>
<td align="char" char=".">&#x2212;4,779.2</td>
<td align="char" char=".">&#x2212;5,585.6</td>
<td align="char" char=".">&#x2212;5,656.2</td>
<td align="char" char=".">&#x2212;4,575.5</td>
<td align="char" char=".">&#x2212;3,939.4</td>
<td align="char" char=".">&#x2212;3,685.9</td>
<td align="char" char=".">&#x2212;2,230.1</td>
<td align="char" char=".">3,008.8</td>
</tr>
<tr>
<td align="left">RF</td>
<td align="char" char=".">&#x2212;5,051.6</td>
<td align="char" char=".">
<bold>10,589</bold>
</td>
<td align="char" char=".">&#x2212;3,302.9</td>
<td align="char" char=".">15,536</td>
<td align="char" char=".">&#x2212;284.62</td>
<td align="char" char=".">&#x2212;1,229.5</td>
<td align="char" char=".">&#x2212;1,366.4</td>
<td align="char" char=".">66,743</td>
</tr>
<tr>
<td align="left">DDR-RNN</td>
<td align="char" char=".">&#x2212;2,727.0</td>
<td align="char" char=".">&#x2212;2,309.5</td>
<td align="char" char=".">&#x2212;3,979.3</td>
<td align="char" char=".">
<bold>35,248</bold>
</td>
<td align="char" char=".">&#x2212;2,298.4</td>
<td align="char" char=".">&#x2212;2,132.2</td>
<td align="char" char=".">&#x2212;3,780.9</td>
<td align="char" char=".">&#x2212;2097.6</td>
</tr>
<tr>
<td align="left">FDRNN</td>
<td align="char" char=".">&#x2212;4,543.2</td>
<td align="char" char=".">6,204.1</td>
<td align="char" char=".">&#x2212;3,791.8</td>
<td align="char" char=".">18,960</td>
<td align="char" char=".">&#x2212;1,619.6</td>
<td align="char" char=".">&#x2212;2,331.1</td>
<td align="char" char=".">&#x2212;3,249.2</td>
<td align="char" char=".">4,145.3</td>
</tr>
<tr>
<td align="left">QF-PGN</td>
<td align="char" char=".">&#x2212;5,024.0</td>
<td align="char" char=".">&#x2212;4,598.6</td>
<td align="char" char=".">3,316.0</td>
<td align="char" char=".">&#x2212;4,203.6</td>
<td align="char" char=".">&#x2212;2,341.2</td>
<td align="char" char=".">&#x2212;3,987.8</td>
<td align="char" char=".">&#x2212;2043.4</td>
<td align="char" char=".">
<bold>79,433</bold>
</td>
</tr>
<tr>
<td align="left">QFTN-U</td>
<td align="char" char=".">
<bold>588.81</bold>
</td>
<td align="char" char=".">&#x2212;4,598.6</td>
<td align="char" char=".">
<bold>10,089</bold>
</td>
<td align="char" char=".">24,602</td>
<td align="char" char=".">
<bold>2,499.3</bold>
</td>
<td align="char" char=".">
<bold>399.49</bold>
</td>
<td align="char" char=".">
<bold>538.54</bold>
</td>
<td align="char" char=".">57,689</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Bold values indicating the best performance in terms of corresponding metrics.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4-4">
<title>4.4&#x20;QPL-Inspired Intraday Trading Model Analysis</title>
<p>We analyze the decision of the QPL-based intraday models in <xref ref-type="table" rid="T4">Table&#x20;4</xref> as two classifications: 1) predict the optimal QPL to settle; 2) predict the profitable QPL (the QPLs having the same trading direction with the optimal one) to settle. Noticeably, the action space for PGN and QFTN-L is {&#x2b;1 QPL, Neutral, -1 QPL}, which means that these two classification tasks for them are actually the same. QFTN-7 might have multiple ground truths, as the payoff might be the same while settlement in varied QPLs, thus we only report the accuracy. <xref ref-type="table" rid="T4">Table&#x20;4</xref> indicates two points: 1) comparing with PGN, our QFTN-L with LSTM as feature extraction has higher accuracy in the optimal QPL selection. The contribution of LSTM to our model can also be proved in the ablation study in <xref ref-type="table" rid="T2">Table&#x20;2</xref>. 2) QFTN-U has less accuracy in optimal QPL prediction compared with QFTN-L, due to the larger action space brings difficulties in decision. Nevertheless, QFTN-U earns higher CPR and SR. We visualize the reward in the training process and the actions made in testing as shown in <xref ref-type="fig" rid="F6">Figure&#x20;6</xref>. We analyze that the better performance of QFTN-U is due to the more accurate judgment of trading direction (see their accuracy in the trading direction classification). In addition, QFTN-U can explore its policy in a broader range. When the agent perceives changes in the market environment confidently, it can select the QPL farther than the ground state as the target price for order closing, rather than only the first positive or negative QPL, thereby obtaining more potential payoff, although the action might not be optimal. For instance, if the price is in a substantial increase, agents acquire higher rewards by closing orders at &#x2b;3 QPL rather than the only positive QPL in QFTN-L&#x2019;s candidate decisions. According to <xref ref-type="fig" rid="F6">Figure&#x20;6</xref>, the trading directions made by two QFTNs are usually the same, but QFTN-U tends to enlarge the levels of selected QPL to obtain more profit. However, the Ultra model needs more training episodes to converge normally (GBPUSD, EURUSD, and OILUSe, etc.). Additionally, the Lite model suffers from the local optimal trap on some datasets (AUDUSD and HSI), in which our model tends to select the same action consistently, e.g., the Lite model keeps delivering a short trade with uniform TP setting in the -1 QPL for AUDUSD.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Decision classification metrices.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left"/>
<th colspan="4" align="center">
<italic>Optimal QPL Prediction</italic>
</th>
<th colspan="4" align="center">
<italic>Trading Direction Prediction</italic>
</th>
</tr>
<tr>
<th align="left"/>
<th align="center">Acc.</th>
<th align="center">P</th>
<th align="center">R</th>
<th align="center">F1</th>
<th align="center">Acc.</th>
<th align="center">P</th>
<th align="center">R</th>
<th align="center">F1</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Pgn (3x)</td>
<td align="char" char=".">0.34</td>
<td align="center">0.25</td>
<td align="center">0.25</td>
<td align="center">0.37</td>
<td align="char" char=".">0.34</td>
<td align="char" char=".">0.25</td>
<td align="char" char=".">0.25</td>
<td align="char" char=".">0.37</td>
</tr>
<tr>
<td align="left">Qftn-L (3x)</td>
<td align="char" char=".">0.56</td>
<td align="center">0.54</td>
<td align="center">0.50</td>
<td align="center">0.50</td>
<td align="char" char=".">0.56</td>
<td align="char" char=".">0.54</td>
<td align="char" char=".">0.50</td>
<td align="char" char=".">0.50</td>
</tr>
<tr>
<td align="left">Qftn-U (7x)</td>
<td align="char" char=".">0.48</td>
<td align="center">&#x2014;</td>
<td align="center">&#x2014;</td>
<td align="center">&#x2014;</td>
<td align="char" char=".">
<bold>0.80</bold>
</td>
<td align="char" char=".">
<bold>0.78</bold>
</td>
<td align="char" char=".">
<bold>0.78</bold>
</td>
<td align="char" char=".">
<bold>0.82</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Bold values indicating the best performance in terms of corresponding metrics.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Training curves for different settings in action space&#x20;size.</p>
</caption>
<graphic xlink:href="frai-04-749878-g006.tif"/>
</fig>
</sec>
<sec id="s4-5">
<title>4.5 Increasing the Size of Action Space</title>
<p>In this section, we compare the average CPR and SR among 8 datasets versus different settings of the action space size in <xref ref-type="fig" rid="F7">Figure&#x20;7</xref>. We observe that when the size of the action space is less than 7, increasing this parameter has a positive effect on system performance. Especially, <xref ref-type="fig" rid="F5">Figure&#x20;5</xref> shows that our lite model fails in the HSI dataset but the ultra one achieves strong performance. We argue this is because the larger action space can potentially contribute to trading with complex strategies. However, when the number of candidate actions continues to increase, SR and CPR decrease after <italic>A</italic>&#x20;&#x3d; 7. We analyze as that the action space of the daytrade model should cover the optimal settlement QPL (global ground truth) within the daily price range ideally. Therefore, if the QPL that brings the maximum reward is not in the model&#x2019;s action space, enlarging the action space will be more possible to capture the global ground truth. However, if the action space has covered the ground truth already, it is meaningless to continue to expand the action space. On the contrary, a large number of candidate actions can make the decision to be more difficult. We report the results for each dataset in the supplementary.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Effects of the different settings in action space&#x20;size.</p>
</caption>
<graphic xlink:href="frai-04-749878-g007.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<title>5 Conclusion and Future Work</title>
<p>In this paper, we investigated the Quantum Finance Theory&#x2019;s application in building an end-to-end day-trade RL trader. With a QPL inspired probabilistic loss-and-profit control for the order settlement, our model substantiate the profitability and robustness in the intraday trading task. Experiments reveal our QF-TraderNet outperforms other baselines. To perform intraday trading, we assumed the ground state in <italic>t</italic>-th day is available for QF-TraderNet in this work. One interesting future work will be combining QF-TraderNet with the state-of-the-art forecasters to perform real-time trading by a predictor-trader framework in which a forecaster predicts the opening price in <italic>t</italic>-th day for our QF-TraderNet to perform trading.</p>
</sec>
</body>
<back>
<sec id="s6">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/Supplementary Material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>YQ: Conceptualization, Methodology, Implementation and Experiment, Validation, Formal analysis. Writing and Editing. YQ: Implementation and Experiment, Editing. YY: Visualization. Implementation and Experiment. ZC: Implementation and Experiment. RL: Supervision, Reviewing and Editing.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This paper was supported by Research Grant R202008 of Beijing Normal University-Hong Kong Baptist University United International College (UIC) and Key Laboratory for Artificial Intelligence and Multi-Model Data Processing of Department of Education of Guangdong Province.</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>The authors highly appreciate the provision of computing equipment and facilities from the Division of Science and Technology of Beijing Normal University-Hong Kong Baptist University United International College (UIC). The authors also wish to thank Quantum Finance Forecast Center of UIC for the R&#x26;D supports and the provision of the platform <ext-link ext-link-type="uri" xlink:href="http://qffc.org">qffc.org</ext-link> for system testing and evaluation.</p>
</ack>
<fn-group>
<fn id="fn1">
<label>1</label>
<p>
<bold>
<italic>r</italic>
</bold> in here denotes the reward of RL agent, rather than the previous price return <italic>r</italic>(<italic>t</italic>) in the QPL evaluation</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Bian</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Xing</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>T.-Y.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Investment Behaviors Can Tell what inside: Exploring Stock Intrinsic Properties for Stock Trend Prediction</article-title>,&#x201d; in <conf-name>Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery &#x26; Data Mining</conf-name>, <conf-loc>Anchorage, AK</conf-loc>, <conf-date>August 4&#x2013;8, 2019</conf-date>, <fpage>2376</fpage>&#x2013;<lpage>2384</lpage>. </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Pan</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Jia</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Trading Strategy of Structured Mutual Fund Based on Deep Learning Network</article-title>. <source>Expert Syst. Appl.</source> <volume>183</volume>, <fpage>115390</fpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2021.115390</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Colby</surname>
<given-names>R. W.</given-names>
</name>
<name>
<surname>Meyers</surname>
<given-names>T. A.</given-names>
</name>
</person-group> (<year>1988</year>). <source>The Encyclopedia of Technical Market Indicators</source>. <publisher-loc>Homewood, IL</publisher-loc>: <publisher-name>Dow Jones-Irwin</publisher-name>. </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dempster</surname>
<given-names>M. A. H.</given-names>
</name>
<name>
<surname>Leemans</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>An Automated Fx Trading System Using Adaptive Reinforcement Learning</article-title>. <source>Expert Syst. Appl.</source> <volume>30</volume>, <fpage>543</fpage>&#x2013;<lpage>552</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2005.10.012</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Deng</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Bao</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Kong</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Ren</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Dai</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Deep Direct Reinforcement Learning for Financial Signal Representation and Trading</article-title>. <source>IEEE Trans. Neural Netw. Learn. Syst.</source> <volume>28</volume>, <fpage>653</fpage>&#x2013;<lpage>664</lpage>. <pub-id pub-id-type="doi">10.1109/TNNLS.2016.2522401</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Gers</surname>
<given-names>F. A.</given-names>
</name>
<name>
<surname>Schmidhuber</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Cummins</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>Learning to Forget: Continual Prediction with Lstm</article-title>. <source>Neural Comput.</source> <volume>12</volume> (<issue>10</issue>), <fpage>2451</fpage>&#x2013;<lpage>2471</lpage>. <pub-id pub-id-type="doi">10.1162/089976600300015015</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Giudici</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Pagnottoni</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Polinesi</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Network Models to Enhance Automated Cryptocurrency Portfolio Management</article-title>. <source>Front. Artif. Intell.</source> <volume>3</volume>, <fpage>22</fpage>. <pub-id pub-id-type="doi">10.3389/frai.2020.00022</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Giudici</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Leach</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Pagnottoni</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Libra or Librae? Basket Based Stablecoins to Mitigate Foreign Exchange Volatility Spillovers</article-title>. <source>Finance Res. Lett.</source>, <fpage>102054</fpage>. <pub-id pub-id-type="doi">10.1016/j.frl.2021.102054</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>D.-j.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Hoi</surname>
<given-names>S. C. H.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Robust Median Reversion Strategy for Online Portfolio Selection</article-title>. <source>IEEE Trans. Knowl. Data Eng.</source> <volume>28</volume>, <fpage>2480</fpage>&#x2013;<lpage>2493</lpage>. <pub-id pub-id-type="doi">10.1109/tkde.2016.2563433</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>R. S.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Chaotic Type-2 Transient-Fuzzy Deep Neuro-Oscillatory Network (Ct2tfdnn) for Worldwide Financial Prediction</article-title>. <source>IEEE Trans. Fuzzy Syst.</source> <volume>28</volume> (<issue>4</issue>), <fpage>731</fpage>&#x2013;<lpage>745</lpage>. <pub-id pub-id-type="doi">10.1109/tfuzz.2019.2914642</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2020</year>). <source>Quantum Finance: Intelligent Forecast and Trading Systems</source>. <publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer</publisher-name>. </citation>
</ref>
<ref id="B12">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Bian</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Qin</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>T.-Y.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Individualized Indicator for All: Stock-wise Technical Indicator Optimization with Stock Embedding</article-title>,&#x201d; in <conf-name>Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery &#x26; Data Mining</conf-name>, <conf-loc>Anchorage, AK</conf-loc>, <conf-date>August 4&#x2013;8, 2019</conf-date>, <fpage>894</fpage>&#x2013;<lpage>902</lpage>. </citation>
</ref>
<ref id="B13">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Marques</surname>
<given-names>N. C.</given-names>
</name>
<name>
<surname>Gomes</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2010</year>). &#x201c;<article-title>Maximus-ai: Using Elman Neural Networks for Implementing a Slmr Trading Strategy</article-title>,&#x201d; in <conf-name>International Conference on Knowledge Science, Engineering and Management</conf-name>, <conf-loc>Belfast, United Kingdom</conf-loc>, <conf-date>September 1&#x2013;3, 2010</conf-date> (<publisher-name>Springer</publisher-name>), <fpage>579</fpage>&#x2013;<lpage>584</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-642-15280-1_55</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Meng</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J.-W.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Quantum Spatial-Periodic Harmonic Model for Daily price-limited Stock Markets</article-title>. <source>Physica A: Stat. Mech. its Appl.</source> <volume>438</volume>, <fpage>154</fpage>&#x2013;<lpage>160</lpage>. <pub-id pub-id-type="doi">10.1016/j.physa.2015.06.041</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Mohan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Mullapudi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sammeta</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Vijayvergia</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Anastasiu</surname>
<given-names>D. C.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Stock price Prediction Using News Sentiment Analysis</article-title>,&#x201d; in <conf-name>2019 IEEE Fifth International Conference on Big Data Computing Service and Applications (BigDataService)</conf-name>, <conf-loc>Newark, CA</conf-loc>, <conf-date>April 4&#x2013;9, 2019</conf-date>, <fpage>205</fpage>&#x2013;<lpage>208</lpage>. <pub-id pub-id-type="doi">10.1109/BigDataService.2019.00035</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Moody</surname>
<given-names>J.&#x20;E.</given-names>
</name>
<name>
<surname>Saffell</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>1998</year>). &#x201c;<article-title>Reinforcement Learning for Trading</article-title>,&#x201d; in <source>Advances in Neural Information Processing Systems</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>, <fpage>917</fpage>&#x2013;<lpage>923</lpage>. </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Moody</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Saffell</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>Learning to Trade via Direct Reinforcement</article-title>. <source>IEEE Trans. Neural Netw.</source> <volume>12</volume>, <fpage>875</fpage>&#x2013;<lpage>889</lpage>. <pub-id pub-id-type="doi">10.1109/72.935097</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Neely</surname>
<given-names>C. J.</given-names>
</name>
<name>
<surname>Rapach</surname>
<given-names>D. E.</given-names>
</name>
<name>
<surname>Tu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Forecasting the Equity Risk Premium: the Role of Technical Indicators</article-title>. <source>Manage. Sci.</source> <volume>60</volume>, <fpage>1772</fpage>&#x2013;<lpage>1791</lpage>. <pub-id pub-id-type="doi">10.1287/mnsc.2013.1838</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pagnottoni</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Neural Network Models for Bitcoin Option Pricing</article-title>. <source>Front. Artif. Intell.</source> <volume>2</volume>, <fpage>5</fpage>. <pub-id pub-id-type="doi">10.3389/frai.2019.00005</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Peralta</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Zareei</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>A Network Approach to Portfolio Selection</article-title>. <source>J.&#x20;Empirical Finance</source> <volume>38</volume>, <fpage>157</fpage>&#x2013;<lpage>180</lpage>. <pub-id pub-id-type="doi">10.1016/j.jempfin.2016.06.003</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pichler</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Poledna</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Thurner</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Systemic Risk-Efficient Asset Allocations: Minimization of Systemic Risk as a Network Optimization Problem</article-title>. <source>J.&#x20;Financial Stab.</source> <volume>52</volume>, <fpage>100809</fpage>. <pub-id pub-id-type="doi">10.1016/j.jfs.2020.100809</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Resta</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Pagnottoni</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>De Giuli</surname>
<given-names>M. E.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Technical Analysis on the Bitcoin Market: Trading Opportunities or Investors&#x27; Pitfall?</article-title> <source>Risks</source> <volume>8</volume>, <fpage>44</fpage>. <pub-id pub-id-type="doi">10.3390/risks8020044</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Sutton</surname>
<given-names>R. S.</given-names>
</name>
<name>
<surname>Barto</surname>
<given-names>A. G.</given-names>
</name>
</person-group> (<year>2018</year>). <source>Reinforcement Learning: An Introduction</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT press</publisher-name>. </citation>
</ref>
<ref id="B24">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Sutton</surname>
<given-names>R. S.</given-names>
</name>
<name>
<surname>McAllester</surname>
<given-names>D. A.</given-names>
</name>
<name>
<surname>Singh</surname>
<given-names>S. P.</given-names>
</name>
<name>
<surname>Mansour</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2000</year>). &#x201c;<article-title>Policy Gradient Methods for Reinforcement Learning with Function Approximation</article-title>,&#x201d; in <source>Advances in Neural Information Processing Systems</source>, <fpage>1057</fpage>&#x2013;<lpage>1063</lpage>. </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tran</surname>
<given-names>D. T.</given-names>
</name>
<name>
<surname>Iosifidis</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Kanniainen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Gabbouj</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Temporal Attention-Augmented Bilinear Network for Financial Time-Series Data Analysis</article-title>. <source>IEEE Trans. Neural Netw. Learn. Syst.</source> <volume>30</volume>, <fpage>1407</fpage>&#x2013;<lpage>1418</lpage>. <pub-id pub-id-type="doi">10.1109/TNNLS.2018.2869225</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vella</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Ng</surname>
<given-names>W. L.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>A Dynamic Fuzzy Money Management Approach for Controlling the Intraday Risk-Adjusted Performance of Ai Trading Algorithms</article-title>. <source>Intell. Sys. Acc. Fin. Mgmt.</source> <volume>22</volume>, <fpage>153</fpage>&#x2013;<lpage>178</lpage>. <pub-id pub-id-type="doi">10.1002/isaf.1359</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Fang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Niu</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Financial Time Series Prediction Using Elman Recurrent Random Neural Networks</article-title>. <source>Comput. Intell. Neurosci.</source> <volume>2016</volume>, <fpage>14</fpage>. <pub-id pub-id-type="doi">10.1155/2016/4742515</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Xiong</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Alphastock: A Buying-Winners-And-Selling-Losers Investment Strategy Using Interpretable Deep Reinforcement Attention Networks</article-title>,&#x201d; in <conf-name>Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery &#x26; Data Mining</conf-name>, <conf-loc>Anchorage, AK</conf-loc>, <conf-date>August 4&#x2013;8, 2019</conf-date>, <fpage>1900</fpage>&#x2013;<lpage>1908</lpage>. </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Yue</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Rao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Boris</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>A Deep Learning Framework for Financial Time Series Using Stacked Autoencoders and Long-Short Term Memory</article-title>. <source>Plos One</source> <volume>12</volume>, <fpage>e0180944</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0180944</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wold</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Esbensen</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Geladi</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>1987</year>). <article-title>Principal Component Analysis</article-title>. <source>Chemometrics Intell. Lab. Syst.</source> <volume>2</volume>, <fpage>37</fpage>&#x2013;<lpage>52</lpage>. <pub-id pub-id-type="doi">10.1016/0169-7439(87)80084-9</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Xiong</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>X.-Y.</given-names>
</name>
<name>
<surname>Zhong</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Walid</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2018</year>). <source>Practical Deep Reinforcement Learning Approach for Stock Trading</source>. <comment>arXiv preprint arXiv:1811.07522</comment>. </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ye</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>J.&#x20;P.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Non-classical Oscillator Model for Persistent Fluctuations in Stock Markets</article-title>. <source>Physica A: Stat. Mech. its Appl.</source> <volume>387</volume>, <fpage>1255</fpage>&#x2013;<lpage>1263</lpage>. <pub-id pub-id-type="doi">10.1016/j.physa.2007.10.050</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>