<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Water</journal-id>
<journal-title>Frontiers in Water</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Water</abbrev-journal-title>
<issn pub-type="epub">2624-9375</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frwa.2021.740044</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Water</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Deep Learning for Isotope Hydrology: The Application of Long Short-Term Memory to Estimate High Temporal Resolution of the Stable Isotope Concentrations in Stream and Groundwater</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Sahraei</surname> <given-names>Amir</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1195970/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Houska</surname> <given-names>Tobias</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1184789/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Breuer</surname> <given-names>Lutz</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/124565/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Institute for Landscape Ecology and Resources Management (ILR), Research Centre for BioSystems, Land Use and Nutrition (iFZ), Justus Liebig University Giessen</institution>, <addr-line>Giessen</addr-line>, <country>Germany</country></aff>
<aff id="aff2"><sup>2</sup><institution>Centre for International Development and Environmental Research (ZEU), Justus Liebig University Giessen</institution>, <addr-line>Giessen</addr-line>, <country>Germany</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Scott Thomas Allen, University of Nevada, Reno, United States</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Catie Finkenbiner, Oregon State University, United States; Si-Liang Li, Tianjin University, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Amir Sahraei <email>amirhossein.sahraei&#x00040;umwelt.uni-giessen.de</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Water and Critical Zone, a section of the journal Frontiers in Water</p></fn></author-notes>
<pub-date pub-type="epub">
<day>10</day>
<month>09</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>3</volume>
<elocation-id>740044</elocation-id>
<history>
<date date-type="received">
<day>12</day>
<month>07</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>19</day>
<month>08</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2021 Sahraei, Houska and Breuer.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Sahraei, Houska and Breuer</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract><p>Recent advances in laser spectroscopy has made it feasible to measure stable isotopes of water in high temporal resolution (i.e., sub-daily). High-resolution data allow the identification of fine-scale, short-term transport and mixing processes that are not detectable at coarser resolutions. Despite such advantages, operational routine and long-term sampling of stream and groundwater sources in high temporal resolution is still far from being common. Methods that can be used to interpolate infrequently measured data at multiple sampling sites would be an important step forward. This study investigates the application of a Long Short-Term Memory (LSTM) deep learning model to predict complex and non-linear high-resolution (3 h) isotope concentrations of multiple stream and groundwater sources under different landuse and hillslope positions in the Schwingbach Environmental Observatory (SEO), Germany. The main objective of this study is to explore the prediction performance of an LSTM that is trained on multiple sites, with a set of explanatory data that are more straightforward and less expensive to measure compared to the stable isotopes of water. The explanatory data consist of meteorological data, catchment wetness conditions, and natural tracers (i.e., water temperature, pH and electrical conductivity). We analyse the model&#x00027;s sensitivity to different input data and sequence lengths. To ensure an efficient model performance, a Bayesian optimization approach is employed to optimize the hyperparameters of the LSTM. Our main finding is that the LSTM allows for predicting stable isotopes of stream and groundwater by using only short-term sequence (6 h) of measured water temperature, pH and electrical conductivity. The best performing LSTM achieved, on average of all sampling sites, an RMSE of 0.7&#x02030;, MAE of 0.4&#x02030;, <italic>R</italic><sup>2</sup> of 0.9 and NSE of 0.7. The LSTM can be utilized to predict and interpolate the continuous isotope concentration time series either for data gap filling or in case where no continuous data acquisition is feasible. This is very valuable in practice because measurements of these tracers are still much cheaper than stable isotopes of water and can be continuously conducted with relatively minor maintenance.</p></abstract>
<kwd-group>
<kwd>machine learning</kwd>
<kwd>deep learning</kwd>
<kwd>long short-term memory</kwd>
<kwd>sensitivity analysis</kwd>
<kwd>Bayesian hyperparameter optimization</kwd>
<kwd>stable isotopes of water</kwd>
<kwd>high-resolution data</kwd>
<kwd>isotope hydrology</kwd>
</kwd-group>
<counts>
<fig-count count="10"/>
<table-count count="3"/>
<equation-count count="10"/>
<ref-count count="80"/>
<page-count count="20"/>
<word-count count="11139"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>Catchment hydrological processes are complex and it is challenging to comprehend how the catchment responds to precipitation (Uhlenbrook et al., <xref ref-type="bibr" rid="B64">2002</xref>; Zhou et al., <xref ref-type="bibr" rid="B78">2021</xref>). Stable isotopes of water (&#x003B4;<sup>2</sup>H and &#x003B4;<sup>18</sup>O) have been widely employed as conservative natural tracers in catchment hydrology to shed light on the hydrological processes. Such tracers have proved to be valuable tools to investigate the origin and formation of recharged water, surface-groundwater interactions, mixing processes between various water sources, and differentiation of evaporation and evapotranspiration (Kendall and McDonnell, <xref ref-type="bibr" rid="B24">2012</xref>; Orlowski et al., <xref ref-type="bibr" rid="B49">2016</xref>). Particularly at the catchment scale, the stable isotopes of water have been used to differentiate runoff components via hydrograph separation techniques (Klaus and McDonnell, <xref ref-type="bibr" rid="B30">2013</xref>), to estimate mean transit times (McGuire and McDonnell, <xref ref-type="bibr" rid="B42">2006</xref>), to identify flow pathways (Tetzlaff et al., <xref ref-type="bibr" rid="B63">2015</xref>), to explore groundwater recharge rates (Koeniger et al., <xref ref-type="bibr" rid="B31">2016</xref>), to understand soil water mixing processes (Sprenger et al., <xref ref-type="bibr" rid="B61">2016</xref>) and to improve hydrological model simulations (Windhorst et al., <xref ref-type="bibr" rid="B71">2014</xref>).</p>
<p>Recent advances in laser spectroscopy has made it feasible to measure stable isotopes in high temporal resolution (i.e., sub-daily) and <italic>in situ</italic>. High-resolution data allow the identification of fine-scale, short-term transport and mixing processes that are not detectable at coarser resolutions (Birkel et al., <xref ref-type="bibr" rid="B7">2012</xref>). Previously, studies using sub-daily isotope data focused on single precipitation events (McGlynn et al., <xref ref-type="bibr" rid="B41">2004</xref>; Wissmeier and Uhlenbrook, <xref ref-type="bibr" rid="B72">2007</xref>; Berman et al., <xref ref-type="bibr" rid="B6">2009</xref>). Whilst useful, they are still limited in providing insight into short-term response variability and mixing processes over a longer-term catchment behavior. To overcome these limitations, a few research groups recently developed automated systems for continuous monitoring of water isotopes directly in the field. von Freyberg et al. (<xref ref-type="bibr" rid="B69">2017</xref>) analyzed isotopes of precipitation and stream water every 30 min over 28 days to derive fractions of event water from hydrograph separation at eight precipitation events in the Erlenbach catchment, Switzerland. The result indicated that the high-resolution measurements allowed an in-depth comparison of event water fractions through endmember mixing analysis. Heinz et al. (<xref ref-type="bibr" rid="B21">2014</xref>) introduced the technical design of an automated system and reported a primary proof-of-concept to monitor isotopes of stream and groundwater in rice paddies in the Philippines. They concluded that the high-resolution measurements provided the foundation for insights into hydrological interactions that could not be studied previously, particularly with respect to spatially distributed sampling. Based on this setup, Mahindawansha et al. (<xref ref-type="bibr" rid="B37">2018</xref>) reported impact of seasons and crops on surface and groundwater isotope concentrations in these rice paddies. The result showed that groundwater isotopes reacted rapidly to irrigation under maize dry season, suggesting the process of preferential flow through deep roots and cracks. Quade et al. (<xref ref-type="bibr" rid="B54">2019</xref>) monitored soil water isotopic composition every 30 min during one growing season of sugar beet to partition evapotranspiration flux at the Selhausen agricultural research site, Germany. The comparison between non-destructive high-resolution and destructive coarser resolution sampling of soil water showed significant discrepancies between the isotopic compositions of evaporation led in turn to significant differences in evapotranspiration flux estimation. Sahraei et al. (<xref ref-type="bibr" rid="B57">2020</xref>) analyzed isotopes of stream, groundwater and precipitation every 20 min over approximately 5 months to investigate hydrological response behavior and role of precipitation and antecedent wetness conditions in runoff generation in the Schwingbach Environmental Observatory (SEO), Germany. The result revealed that maximum event water fractions of stream and groundwater responded rapidly to precipitation events, indicating the fast delivery of water to the stream through shallow subsurface flow pathways.</p>
<p>Despite the advantages of high-resolution water isotope data, the routine measurements are still far from being common. Long-term, high-resolution sampling of multiple sources are even less common despite the fact that such measurements are likely to provide new insight into hydrological processes and can help to constrain individual endmembers and flow pathways that contribute to runoff generation. Methods that can be used to interpolate infrequently measured data at multiple sampling sites would be an important step forward. Statistical time series methods such as simple exponential smoothing (SES), autoregressive (AR), moving average (MA) or autoregressive integrated moving average (ARIMA) have been traditionally employed to predict time series problems. The main drawback of the statistical methods is that they assume that the series are derived from linear processes and hence they might be inadequate for water isotope time series that are non-linear (Zhang et al., <xref ref-type="bibr" rid="B76">1998</xref>; Khashei et al., <xref ref-type="bibr" rid="B26">2009</xref>). Another major limitation is that the statistical methods are local models, in which the free parameters are individually estimated for each time series. It means that it is not possible to share the learning over multiple time series to extract patterns that cannot be distinguished at an individual level (Calkoen et al., <xref ref-type="bibr" rid="B8">2021</xref>). In contrast, machine learning provides a useful tool for the joint extraction of non-linear patterns form a collection of time series. The core idea is to predict isotope concentrations with a set of explanatory data that are more straightforward and less expensive to measure. If the machine learning algorithm is able to predict isotope concentrations from the explanatory data, data-driven interpolations of continuous isotope concentration time series can be acquired.</p>
<p>Machine learning is a data-driven approach that aims to give computers the ability to automatically learn and extract patterns from data (Samuel, <xref ref-type="bibr" rid="B58">1959</xref>; Goodfellow et al., <xref ref-type="bibr" rid="B18">2016</xref>). Machine learning has been increasingly applied in hydrology and earth system science in recent years, owing to its capability to efficiently simulate highly non-linear and complex systems without any a priori knowledge of underlying physical processes. Extensive reviews have been released for the application of machine learning in water resources (Lange and Sippel, <xref ref-type="bibr" rid="B34">2020</xref>; Zounemat-Kermani et al., <xref ref-type="bibr" rid="B79">2020</xref>). Artificial Neural Network (ANN) is the most commonly used machine learning algorithm for the prediction of hydrological variables (Maier and Dandy, <xref ref-type="bibr" rid="B38">2000</xref>; Maier et al., <xref ref-type="bibr" rid="B39">2010</xref>). The main advantage of ANN models is that they are universal function approximators, meaning that they can automatically fit a wide range of functions with a high accuracy level (Khashei and Bijari, <xref ref-type="bibr" rid="B25">2010</xref>). The other major advantage of ANNs is that they have an inherent generalization capability, meaning that they are able to recognize and respond to the patterns that are analogous, but not identical to those on which they have been trained (Benardos and Vosniakos, <xref ref-type="bibr" rid="B3">2007</xref>). Nevertheless, a drawback of ANNs, which have primarily been employed for the analysis of time series in the past, is that any information about the sequence of input features is lost (Kratzert et al., <xref ref-type="bibr" rid="B32">2018</xref>). Therefore, more advance machine learning models are required to efficiently handle these temporal dependencies.</p>
<p>Deep learning is an advance sub-field of machine learning that has drawn significant attention recently. Deep learning generally refers to deep neural networks with multilayer structures that can extract high-level representations from complex and high-dimensional data via a hierarchical learning process applying multiple non-linear transformations (Shen, <xref ref-type="bibr" rid="B60">2018</xref>; Zuo et al., <xref ref-type="bibr" rid="B80">2019</xref>). Long Short-Term Memory (LSTM) is the current state-of-the-art deep learning architecture that is widely adopted to simulate sequential data like time series (Gers et al., <xref ref-type="bibr" rid="B17">2002</xref>). LSTM is a type of Recurrent Neural Network (RNN) that was originally developed by Hochreiter and Schmidhuber (<xref ref-type="bibr" rid="B22">1997</xref>). Unlike the traditional RNN networks, LSTM does not suffer from exploding and vanishing gradients, which allows the network to learn long-term dependencies (Hochreiter and Schmidhuber, <xref ref-type="bibr" rid="B22">1997</xref>). This is beneficial to capture dynamics of catchment processes like storage effects, which may play an important role in hydrological processes (Kratzert et al., <xref ref-type="bibr" rid="B32">2018</xref>). Fang et al. (<xref ref-type="bibr" rid="B15">2017</xref>) successfully applied LSTM for the first time in hydrological research to predict soil moisture using meteorological forcing data, static physiographic attributes and model-simulated moisture as inputs. They concluded that the LSTM generalizes well across regions with different climates and environmental conditions. Kratzert et al. (<xref ref-type="bibr" rid="B32">2018</xref>) used LSTM to model daily runoff using the Catchment Attributes and Meteorology for Large-sample Studies (CAMELS) data set over hundreds of catchments in the USA. The result demonstrated that the LSTM showed better prediction performance than traditional RNN due to its ability of learning and storing long-term dependencies. Zhang et al. (<xref ref-type="bibr" rid="B77">2018b</xref>) compared the performance of LSTM with that of a feed-forward ANN for the prediction of water table depths in agricultural areas in northwestern China. The result revealed that the LSTM outperformed the traditional feed-forward ANN. Liu et al. (<xref ref-type="bibr" rid="B36">2019</xref>) investigated the application of LSTM to predict water quality parameters such as turbidity, dissolved oxygen and chemical oxygen demand in China. They reported that the LSTM is a feasible and effective approach for water quality prediction.</p>
<p>To the best of our knowledge, the potential of LSTM has not been investigated in the field of isotope hydrology yet; and in general, only few studies have explored the application of machine learning in this field. Cerar et al. (<xref ref-type="bibr" rid="B9">2018</xref>) compared the performance of multilayer feed-forward ANN network with that of ordinary kriging, simple and multiple linear regression for predicting the isotope composition (&#x003B4;<sup>18</sup>O) of groundwater over several locations across Slovenia. They collected 83 groundwater samples from two campaigns, first in spring and second in autumn under base flow conditions. The result showed that feed-forward ANN achieved better performance than the other three models. Sahraei et al. (<xref ref-type="bibr" rid="B56">2021</xref>) investigated the potential of Support Vector Machine (SVM) and multilayer feed-forward ANN to predict maximum event water fractions of streamflow in the Schwingbach Environmental Observatory (SEO), Germany. They found that the SVM outperformed the ANN model as it could better captured the dynamics of maximum event water fractions under distinct hydroclimatic conditions and flow regimes.</p>
<p>For the first time in the field of isotope hydrology, we investigate the application of deep leaning to predict stable isotopes of water. We apply an LSTM to estimate complex and non-linear high-resolution isotope concentrations of multiple stream and groundwater sources under different land use and hillslope positions in the Schwingbach Environmental Observatory (SEO), Germany. We use an automated <italic>in situ</italic> mobile laboratory, the Water Analysis Trailer for Environmental Research (WATER), to sample and measure high-resolution (3 h) isotope concentrations in two stream reaches and three groundwater sources. Explanatory data comprise meteorological data, catchment wetness conditions, and natural tracers, i.e., water temperature, potential of hydrogen (pH) and electrical conductivity (EC) that are more straightforward and less expensive to measure compared to the stable isotopes of water. In particular, we report on: (1) how different combinations of input data affect the prediction accuracy of the LSTM; and (2) how short-, medium- and long-term dependencies relate to the sequence length of input data in the LSTM.</p>
</sec>
<sec sec-type="materials and methods" id="s2">
<title>Materials and Methods</title>
<sec>
<title>Study Area and Data Collection</title>
<p>The study was carried out in the headwater catchment of the Schwingbach Environmental Observatory (SEO) in Hesse, Germany (<xref ref-type="fig" rid="F1">Figure 1</xref>). The catchment area is 1.03 km<sup>2</sup> with the elevation ranging from 310 m in the north to 415 m a.s.l. in the south (<xref ref-type="fig" rid="F1">Figure 1A</xref>). The climate is categorized as temperate oceanic, with a mean annual precipitation of 623 mm and a mean annual air temperature of 9.6&#x000B0;C (Deutscher Wetterdienst, Giessen-Wettenberg station, period 1969&#x02013;2019). 76% of catchment area is covered by forest that is mostly located in the east and south, 15% by farmland in the north and west and 7% by meadows alongside the stream (<xref ref-type="fig" rid="F1">Figure 1B</xref>). The soil is categorized as Cambisol, covered mainly by forests and Stagnosols under farmland. The soil texture is predominantly consists of silt and fine sand with a low clay content. Further details can be found in Orlowski et al. (<xref ref-type="bibr" rid="B49">2016</xref>).</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p><bold>(A)</bold> Schwingbach Environmental Observatory (SEO), <bold>(B)</bold> study area in the Schwingbach headwater and <bold>(C)</bold> measuring network along the stream reach of the Schwingbach headwater.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frwa-03-740044-g0001.tif"/>
</fig>
<p>An automated climate station (AQ5, Campbell Scientific Inc., Shepshed, UK) equipped with a CR1000 data logger recorded precipitation depth, air temperature, relative humidity, air pressure, solar radiation and wind speed at 5 min intervals (<xref ref-type="fig" rid="F1">Figure 1C</xref>). Six remote-controlled data loggers (A753, Adcon, Klosterneuburg, Austria), three at the toeslope (SM1, SM2 and SM3) and three at the footslope (SM4, SM5 and SM6), were connected to sensors (ECH2O 5TE, METER Environment, Pullman, USA) to automatically monitor soil moisture at 5 and 15 cm depths at 5 min intervals. A stream gauge (RBC flume, Eijkelkamp Agrisearch Equipment, Giesbeek, Netherlands) equipped with a pressure transducer (Mini-Diver, Eigenbrodt Inc., K&#x000F6;nigsmoor, Germany) automatically measured water levels at the outlet (SW2) of the catchment at 10 min intervals. The transducer measurements were calibrated against manual readings and continuous stream discharge was obtained via the calibrated stage-discharge relationship provided by the manufacturer.</p>
<p>An automated mobile laboratory, the Water Analysis Trailer for the Environmental Research (WATER), was used to automatically sample and analyse the stable isotopes of water (&#x003B4;<sup>2</sup>H and &#x003B4;<sup>18</sup>O), water temperature, potential of hydrogen (pH) and electrical conductivity (EC) for multiple water sources <italic>in situ</italic> from August 8th until December 9th in 2018 and from April 12th until October 10th in 2019. The WATER was equipped with a continuous water sampler (CWS) (A0217, Picarro Inc., Santa Clara, USA), coupled to a wavelength-scanned cavity ring-down spectrometer (WS-CRDS) (L2130-i, Picarro Inc., Santa Clara, USA) to analyse the isotopic composition of sampled water. Isotopic ratios are reported in per mill (&#x02030;) deviations from the Vienna Standard Mean Ocean Water (VSMOW). We only used &#x003B4;<sup>2</sup>H time series in the LSTM modeling, as &#x003B4;<sup>18</sup>O had a similar variation but lower precision (precision 0.23&#x02030; for &#x003B4;<sup>18</sup>O and 0.57&#x02030; for &#x003B4;<sup>2</sup>H). The high-resolution &#x003B4;<sup>2</sup>H data were verified with water samples that were manually collected from the sampling sites on a weekly basis. A multi-parameter water quality probe (YSI600R, YSI Inc., Yellow Springs, USA) was installed on the sampling board of the WATER to measure temperature, pH and EC of sampled water. For a detailed description of the WATER and the sampling setup refer to Sahraei et al. (<xref ref-type="bibr" rid="B57">2020</xref>).</p>
<p>We measured the stable isotope composition (&#x003B4;<sup>2</sup>H), water temperature, pH and EC of two stream water reaches (SW1 and SW2) and three groundwater sources (GW1, GW2 and GW3) (<xref ref-type="fig" rid="F1">Figure 1C</xref>). SW1 was sampled approximately 145 m upstream of the WATER at the edge of farmland and SW2 was sampled at the outlet of the catchment next to the WATER. Groundwater was sampled from piezometers made from perforated PVC tubes sealed in the upper part with bentonite clay to prevent contamination by surface water. The piezometers of GW1 and GW2 were located at the toeslope on the farmland and meadow, respectively. The piezometer of GW3 was located at the footslope at the edge of the forest. The sampling schedule of the WATER allowed measuring isotopic composition, water temperature, pH and EC for each of the stream water reaches at 1.5-h and for each of the groundwater sources at 3-h intervals.</p>
</sec>
<sec>
<title>Data Pre-processing</title>
<p>As the input features for our LSTM model, we selected a set of explanatory data that are more straightforward and less expensive to measure compared to the stable isotopes of water. The input features were categorized into meteorological, catchment wetness and natural tracer variables (<xref ref-type="table" rid="T1">Table 1</xref>). Our collected data set contained different sampling frequencies. To obtain a uniform frequency, which matches the frequency of the output variable, we aggregated all of the observed data to 3-h intervals by averaging, except the precipitation, which was aggregated to 3-h intervals by summing up. The winter period (10th December 2018&#x02013;11th April 2019) was excluded from the data set because the WATER did not sample during this period due to the freezing weather conditions. When gaps occurred over small timescales due to faulty sensors or equipment maintenance, we used linear interpolation to estimate the missing values. The proportions of observations, which were gap filled through linear interpolation, accounted for at most 3% (73 observations) of total length of the time series for each of the features. We split the observations of all the features into train (70%, 1,700 observations), validation (15%, 365 observations) and test (15%, 365 observations) sets, while maintaining the temporal order of the observations. The train set was used to learn the internal parameters (i.e., weights and bias) of the model. The validation set was used to optimize the hyperparameter, and the test set was used to evaluate the generalization capability of the model. To stabilize the learning process and to speed up the convergence, the data needs to be normalized before feeding to the model. For this, the train set was normalized to be in the range [0&#x02013;1]. The validation and test sets were normalized using the parameters obtained from the train set normalization to avoid data leakage (Hastie et al., <xref ref-type="bibr" rid="B19">2009</xref>).</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Input and output features.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Input features</bold></th>
<th valign="top" align="left"><bold>Symbol</bold></th>
<th valign="top" align="left"><bold>Unit</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><bold>Meteorological data</bold></td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">Precipitation</td>
<td valign="top" align="left">P</td>
<td valign="top" align="left">mm</td>
</tr>
<tr>
<td valign="top" align="left">Air temperature</td>
<td valign="top" align="left">T<sub>a</sub></td>
<td valign="top" align="left">&#x000B0;C</td>
</tr>
<tr>
<td valign="top" align="left">Relative humidity</td>
<td valign="top" align="left">RH</td>
<td valign="top" align="left">%</td>
</tr>
<tr>
<td valign="top" align="left">Air pressure</td>
<td valign="top" align="left">PR</td>
<td valign="top" align="left">hPa</td>
</tr>
<tr>
<td valign="top" align="left">Solar radiation</td>
<td valign="top" align="left">SR</td>
<td valign="top" align="left">W m<sup>&#x02212;2</sup></td>
</tr>
<tr>
<td valign="top" align="left">Wind speed</td>
<td valign="top" align="left">WS</td>
<td valign="top" align="left">m s<sup>&#x02212;1</sup></td>
</tr>
<tr>
<td valign="top" align="left"><bold>Catchment wetness</bold></td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">Stream discharge</td>
<td valign="top" align="left">Q</td>
<td valign="top" align="left">l s<sup>&#x02212;1</sup></td>
</tr>
<tr>
<td valign="top" align="left">Soil moisture at 5 cm depth</td>
<td valign="top" align="left">SM<sub>5</sub></td>
<td valign="top" align="left">%</td>
</tr>
<tr>
<td valign="top" align="left">Soil moisture at 15 cm depth</td>
<td valign="top" align="left">SM<sub>15</sub></td>
<td valign="top" align="left">%</td>
</tr>
<tr>
<td valign="top" align="left"><bold>Natural tracers</bold></td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">Temperature</td>
<td valign="top" align="left">T<sub>w</sub></td>
<td valign="top" align="left">&#x000B0;C</td>
</tr>
<tr>
<td valign="top" align="left">Potential of hydrogen</td>
<td valign="top" align="left">pH</td>
<td valign="top" align="left">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Electrical conductivity</td>
<td valign="top" align="left">EC</td>
<td valign="top" align="left">&#x003BC;S cm<sup>&#x02212;1</sup></td>
</tr>
<tr>
<td valign="top" align="left"><bold>Output feature</bold></td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">Isotopic composition of water</td>
<td valign="top" align="left">&#x003B4;<sup>2</sup>H</td>
<td valign="top" align="left">&#x02030;</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The input features are categorized into meteorological, catchment wetness and natural tracer variables. The output feature is isotope composition (&#x003B4;<sup>2</sup>H) at SW1, SW2, GW1, GW2 and GW3 sampling sites</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>Long Short-Term Memory (LSTM) Model</title>
<p>The LSTM is a type of a Recurrent Neural Network (RNN) that is capable of learning long-term dependencies by overcoming the exploding and vanishing gradient problems of traditional RNN networks (Hochreiter and Schmidhuber, <xref ref-type="bibr" rid="B22">1997</xref>). The main characteristics of the LSTM are the specially designed units so called <italic>memory cell</italic> and <italic>gates</italic>. The memory cell consists of forget, input and output gates that together control the flow of information within the LSTM network. The structure of the LSTM memory cell and the algorithms in the cell are shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. The memory cell contains a specific status for each time step, the so-called cell state <italic>c</italic>, which contains the information for long-term memory. The forget gate controls which information is removed from the cell state. The input gate defines which information is updated to the cell state and the output gate specifies which information is used from the cell state. Suppose a sequence of inputs <italic>x</italic> &#x0003D; [<italic>x</italic><sub>1</sub>, . .., <italic>x</italic><sub><italic>T</italic></sub>] with <italic>T</italic> time steps, where each element <italic>x</italic><sub><italic>t</italic></sub> is a vector that contains input at time step (1 &#x02264; <italic>t</italic> &#x02264; <italic>T</italic>), the process within the LSTM cell is represented with the equations 1&#x02013;6. The LSTM cell updates six parameters at each time step. The first parameter is the forget gate parameter <italic>f</italic><sub><italic>t</italic></sub> that decides how much of the information from the previous cell state <italic>c</italic><sub><italic>t</italic>&#x02212;1</sub> needs to be forgotten by a sigmoid function (&#x003C3;) with a linear calculation of the current input <italic>x</italic><sub><italic>t</italic></sub> and the previous output <italic>h</italic><sub><italic>t</italic>&#x02212;1</sub>. <italic>W</italic>&#x00027;s and <italic>b</italic>&#x00027;s with different subscripts represent the gate-specific network weights and bias parameters for the linear calculations. The second parameter is the input gate parameter <italic>i</italic><sub><italic>t</italic></sub> that controls which new information is updated to the cell state by the sigmoid function with a linear relation on <italic>x</italic><sub><italic>t</italic></sub> and <italic>h</italic><sub><italic>t</italic>&#x02212;1</sub> as well. The new cell state candidate <italic>c</italic><sub><italic>t</italic></sub><sup>&#x02032;</sup> is calculated by a hyperbolic tangent function (<italic>tanh</italic>) with a linear relation on <italic>x</italic><sub><italic>t</italic></sub> and <italic>h</italic><sub><italic>t</italic>&#x02212;1</sub>. The cell state <italic>c</italic><sub><italic>t</italic></sub> is then updated through an element-wise multiplication &#x02299; operator. In the end, the output parameter <italic>o</italic><sub><italic>t</italic></sub> is calculated by the sigmoid function with a linear relation on <italic>x</italic><sub><italic>t</italic></sub> and <italic>h</italic><sub><italic>t</italic>&#x02212;1</sub>. The final output at the current time step <italic>h</italic><sub><italic>t</italic></sub> is the production of <italic>o</italic><sub><italic>t</italic></sub> and the <italic>tanh</italic> function value of the cell state <italic>c</italic><sub><italic>t</italic></sub>.</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M2"><mml:mtable class="eqnarray" columnalign="right"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E2"><label>(2)</label><mml:math id="M3"><mml:mtable class="eqnarray" columnalign="right"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E3"><label>(3)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="right"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>h</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E4"><label>(4)</label><mml:math id="M5"><mml:mtable class="eqnarray" columnalign="right"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02299;</mml:mo><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02299;</mml:mo><mml:mstyle displaystyle="true"><mml:msubsup><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E5"><label>(5)</label><mml:math id="M6"><mml:mtable class="eqnarray" columnalign="right"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E6"><label>(6)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="right"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02299;</mml:mo><mml:mo class="qopname">tanh</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>The architecture of LSTM memory cell.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frwa-03-740044-g0002.tif"/>
</fig>
</sec>
<sec>
<title>Hyperparameter Optimization</title>
<p>The majority of machine learning algorithms possess several settings that control the entire learning process (Goodfellow et al., <xref ref-type="bibr" rid="B18">2016</xref>). These settings are referred to as <italic>hyperparameters</italic>. The hyperparameters are exterior to the model and need to be set before the learning process (G&#x000E9;ron, <xref ref-type="bibr" rid="B16">2019</xref>). The performance and computational complexity of LSTM models strongly depend on the set of hyperparameters that determine many aspects of the algorithm&#x00027;s behavior (Nakisa et al., <xref ref-type="bibr" rid="B47">2018</xref>). Therefore, it is essential to optimize hyperparameters to boost the LSTM performance. In this study, we used a Sequential Model-Based Optimization (SMBO) search with the Tree-structured Parzen Estimator (TPE) algorithm, a Bayesian optimization approach (Bergstra et al., <xref ref-type="bibr" rid="B4">2011</xref>). Bayesian optimization is a very effective optimization algorithm that has been shown to outperform well-established methods i.e., grid and random search (Bergstra et al., <xref ref-type="bibr" rid="B5">2013</xref>; Eggensperger et al., <xref ref-type="bibr" rid="B13">2013</xref>). It develops a statistical model between the hyperparameters and the objective function and makes the assumption that there is a smooth but noisy function that maps between the hyperparameters and the objective function (Reimers and Gurevych, <xref ref-type="bibr" rid="B55">2017</xref>). Given the search history of hyperparameters and objective function, SMBO-TPE suggests hyperparameters for the next trial that are expected to improve the objective function. As the number of trials grows, the search history expands and eventually the hyperparameters become optimized. In this study, we optimized the number of hidden units (i.e., neurons), dropout rate, learning rate, number of epochs and batch size (<xref ref-type="table" rid="T2">Table 2</xref>). We ran the optimization evaluation for 1,000 trials through the search space of the hyperparameters to minimize mean squared error (MSE) on the validation set. The complete set of optimized hyperparameters can be found in <xref ref-type="supplementary-material" rid="SM1">Table S1</xref> in the Supplementary Material.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Hyperparameter set for the LSTM model.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Hyperparameter</bold></th>
<th valign="top" align="left"><bold>Choices</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Number of hidden units</td>
<td valign="top" align="left">10, 20, 30, 40, 50, 60, 70, 80, 90, 100</td>
</tr>
<tr>
<td valign="top" align="left">Dropout rate</td>
<td valign="top" align="left">0.1, 0.2, 0.3, 0.4, 0.5</td>
</tr>
<tr>
<td valign="top" align="left">Learning rate</td>
<td valign="top" align="left">0.00001, 0.0001, 0.001, 0.01, 0.1</td>
</tr>
<tr>
<td valign="top" align="left">Number of epochs</td>
<td valign="top" align="left">2, 3, 5, 7, 10, 15, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100</td>
</tr>
<tr>
<td valign="top" align="left">Batch size</td>
<td valign="top" align="left">32, 64, 128, 256</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>LSTM Model Setup</title>
<p>The architecture that we used in this study consists of an input layer with as many neurons as input features, one LSTM layer, a dropout layer and a fully connected dense layer with a single unit for the output feature. We used the Adam optimizer for the optimization of the learning process (Kingma and Ba, <xref ref-type="bibr" rid="B27">2014</xref>) and <italic>tanh</italic> and sigmoid functions as state and gate activation functions, respectively. We ran the LSTM model in a sequence-to-one mode so that an input sequence of a fixed length was used to predict a &#x003B4;<sup>2</sup>H value at the next time step. The sequence length, i.e., the look-back window, is the length of past input observations that the LSTM looks back to predict an output for the next time step.</p>
<p>The performance of deep learning models improves by making more training data available (Schmidhuber, <xref ref-type="bibr" rid="B59">2015</xref>). Building a single LSTM model that is trained and optimized on all of the sampling sites instead of training and optimizing a separate LSTM model for each of the sampling sites, allows the network to learn more general and abstract patterns of the input&#x02013;output relationship (Kratzert et al., <xref ref-type="bibr" rid="B32">2018</xref>). We therefore built a single LSTM model to predict the &#x003B4;<sup>2</sup>H values at SW1, SW2, GW1, GW2 and GW3 sampling sites. The LSTM was simultaneously trained on train sets of all sampling sites to learn the internal parameters of the model. It simultaneously used validation sets of all the sampling sites to optimize the hyperparameters and predict the &#x003B4;<sup>2</sup>H values on each test set of the sampling sites to estimate the generalization capability of the model.</p>
</sec>
<sec>
<title>LSTM Model Sensitivity to Input Features and Sequence Length</title>
<p>We examined how different combinations of input features affect the prediction performance of the LSTM model. It allows us to identify which combinations give us the most accurate predictions of isotope concentrations in the stream and groundwater sampling sites. We tested seven different input feature scenarios shown in <xref ref-type="table" rid="T3">Table 3</xref>. The first three scenarios (S1, S2 and S3) investigate the effect of meteorological data, catchment wetness conditions and natural tracers on the prediction performance separately. S4, S5 and S6 test the prediction performance when we combine the first three scenarios together and S7 examines the predication performance when we use all of the input features for training of the LSTM. We also used the labels of the sampling sites as input features in all of the seven scenarios. By using the labels of sampling sites as input features, we allow the model to know on which site it trains and hence it produces a unique output for each individual site. This is especially beneficial when using the meteorological data as input features. Since the meteorological data is the same for all sites, the model would otherwise produce the same outputs for them if we do not consider their labels as input features. We used one-hot encoding technique to encode the site labels to the numeric features (Heaton, <xref ref-type="bibr" rid="B20">2015</xref>). This technique creates unique binary columns for each of the labels so that SW1, SW2, GW1, GW2 and GW3 labels are encoded to [1, 0, 0, 0], [0, 1, 0, 0], [0, 0, 1, 0], [0, 0, 0, 1] and [0, 0, 0, 0], respectively. Moreover, we used the month of the year of the measurements&#x00027; time stamp as an input feature in all seven scenarios to represent the inherent seasonality of the data.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Input feature scenarios for the LSTM model.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Scenario label</bold></th>
<th valign="top" align="left"><bold>Input features</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">S1</td>
<td valign="top" align="left">Meteorological data</td>
</tr>
<tr>
<td valign="top" align="left">S2</td>
<td valign="top" align="left">Catchment wetness</td>
</tr>
<tr>
<td valign="top" align="left">S3</td>
<td valign="top" align="left">Natural tracers</td>
</tr>
<tr>
<td valign="top" align="left">S4</td>
<td valign="top" align="left">Meteorological data, catchment wetness</td>
</tr>
<tr>
<td valign="top" align="left">S5</td>
<td valign="top" align="left">Meteorological data, natural tracers</td>
</tr>
<tr>
<td valign="top" align="left">S6</td>
<td valign="top" align="left">Catchment wetness, natural tracers</td>
</tr>
<tr>
<td valign="top" align="left">S7</td>
<td valign="top" align="left">Meteorological data, catchment wetness, natural tracers</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Month of the year and label of sampling sites are used as input features in all of seven scenarios</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>We also examined the effect of the sequence length of input data on the prediction performance of our LSTM model. The sequence length is intrinsically connected to the underlying physical processes driving the dynamics of output variables (Duan et al., <xref ref-type="bibr" rid="B12">2020</xref>). A proper sequence length should be large enough to incorporate all historical information relevant to the prediction of the isotopic compositions. However, too large sequence lengths increase the model complexity and training time that can in turn reduce the performance (Duan et al., <xref ref-type="bibr" rid="B12">2020</xref>). Some previous studies have set the sequence length to an arbitrary number (Zhang et al., <xref ref-type="bibr" rid="B75">2018a</xref>; Le et al., <xref ref-type="bibr" rid="B35">2019</xref>; Tennant et al., <xref ref-type="bibr" rid="B62">2020</xref>), whereas some other studies have reported that the sequence length affects the model performance (Fan et al., <xref ref-type="bibr" rid="B14">2020</xref>; Meyal et al., <xref ref-type="bibr" rid="B44">2020</xref>; Xiang et al., <xref ref-type="bibr" rid="B73">2020</xref>). We therefore investigated the LSTM sensitivity to the sequence length. We defined nine different sequence lengths into three categories to represent short-, medium- and long-term dependencies. We tested sequence lengths of 6, 12 and 24 h for short-term, 72 (3), 168 (7) and 336 (14) h (days) for medium-term and 720 (30), 1,440 (60) and 2,160 (90) h (days) for long-term dependencies. We examined all of these nine sequence lengths for each of the seven input feature scenarios resulting in 63 different scenarios for the sensitivity analysis. We trained the LSTM model and optimized its hyperparameters for each of these scenarios separately to ensure a fair comparison between them.</p>
</sec>
<sec>
<title>Evaluation Metrics</title>
<p>We evaluated the prediction performance of the LSTM model using four statistical metrics (goodness-of-fit criteria). We used root mean squared error (RMSE), mean absolute error (MAE), coefficient of determination (<italic>R</italic><sup>2</sup>) and Nash&#x02013;Sutcliffe efficiency (NSE) according to equations 7&#x02013;10.</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M8"><mml:mtable class="eqnarray" columnalign="right"><mml:mtr><mml:mtd><mml:mi>R</mml:mi><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:msqrt></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E8"><label>(8)</label><mml:math id="M9"><mml:mtable class="eqnarray" columnalign="right"><mml:mtr><mml:mtd><mml:mi>M</mml:mi><mml:mi>A</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E9"><label>(9)</label><mml:math id="M10"><mml:mtable class="eqnarray" columnalign="right"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mstyle displaystyle="true"><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msqrt><mml:mrow><mml:mstyle displaystyle="true"><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:mstyle><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:msqrt><mml:msqrt><mml:mrow><mml:msup><mml:mrow><mml:mstyle displaystyle="true"><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:msqrt></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E10"><label>(10)</label><mml:math id="M11"><mml:mtable class="eqnarray" columnalign="right"><mml:mtr><mml:mtd><mml:mi>N</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="true"><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:mstyle><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:msubsup><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:mstyle><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>N</italic> is the number of observations, <italic>O</italic><sub><italic>i</italic></sub> is the observed value, <italic>P</italic><sub><italic>i</italic></sub> is the predicted value, <inline-formula><mml:math id="M12"><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:mi>O</mml:mi></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:math></inline-formula> is the mean of observed values and <inline-formula><mml:math id="M13"><mml:mover accent="false" class="mml-overline"><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mo accent="true">&#x000AF;</mml:mo></mml:mover></mml:math></inline-formula> is the mean of predicted values. RMSE and MAE explain the average deviation of predicted from observed values with smaller values indicating a better performance. <italic>R</italic><sup>2</sup> is a good indicator for the temporal correspondence between observed and predicted values ranging between [0&#x02013;1], with higher values indicating stronger correlations. NSE assesses the model&#x00027;s capability to predict values different from the mean and gives the proportion of the initial variance accounted for by the model (Nash and Sutcliffe, <xref ref-type="bibr" rid="B48">1970</xref>). The NSE is a preferred criterion in many hydrological studies as it particularly describes a model&#x00027;s capability of matching higher values, extremes and outliers. NSE ranges between [&#x02013;&#x0221E; to 1]. The closer the NSE value is to 1, the better is the match between observed and predicted values. A negative value of the NSE denotes that the mean of observed values is a better predictor than the proposed model. Following, Moriasi et al. (<xref ref-type="bibr" rid="B45">2007</xref>), model performance can be rated as: very good (0.75 &#x0003C; NSE &#x02264; 1.00), good (0.65 &#x0003C; NSE &#x02264; 0.75), satisfactory (0.50 &#x0003C; NSE &#x02264; 0.65) and unsatisfactory (NSE &#x02264; 0.50).</p>
</sec>
<sec>
<title>Setup of Numerical Experiments</title>
<p>We conducted the numerical experiments with Python 3.7 programming environment (van Rossum, <xref ref-type="bibr" rid="B66">1995</xref>) on Ubuntu 20.04 with AMD EPYC 745-core processor, 125 GB of random access memory (RAM) and NVIDIA GTX 1050 Ti graphical processing unit (GPU). The LSTM model was built with Keras 2.3.1 deep learning framework (Chollet, <xref ref-type="bibr" rid="B11">2015</xref>) on top of TensorFlow backend 2.0 (Abadi et al., <xref ref-type="bibr" rid="B1">2016</xref>). Hyperopt 0.2.5 (Bergstra et al., <xref ref-type="bibr" rid="B5">2013</xref>) and Hyperas 0.4.1 (Pumperla, <xref ref-type="bibr" rid="B53">2019</xref>) libraries were used to implement hyperparameter optimization. Scikit-Learn 0.22.2 (Pedregosa et al., <xref ref-type="bibr" rid="B52">2011</xref>), Pandas 1.0.1 (McKinney, <xref ref-type="bibr" rid="B43">2010</xref>), Numpy 1.18.1 (Van Der Walt et al., <xref ref-type="bibr" rid="B65">2011</xref>), and Scipy 1.4.1 (Virtanen et al., <xref ref-type="bibr" rid="B67">2020</xref>) libraries were used for pre-processing and data management. The Matplotlib 3.1.3 (Hunter, <xref ref-type="bibr" rid="B23">2007</xref>) and Seaborn 0.9.1 (Waskom et al., <xref ref-type="bibr" rid="B70">2020</xref>) libraries were used for the data visualization.</p>
</sec>
</sec>
<sec id="s3">
<title>Results and Discussion</title>
<sec>
<title>Temporal Dynamics</title>
<p>In the following, we briefly describe the dynamics of the observed data over the sampling period (<xref ref-type="fig" rid="F3">Figures 3</xref>&#x02013;<xref ref-type="fig" rid="F6">6</xref>). During the sampling period of 2018, precipitation is very low with respect to the long-term historical record (Deutscher Wetterdienst, Giessen-Wettenberg station, period 1969-2019, mean annual precipitation of 623 mm), with a total annual precipitation of 452 mm (<xref ref-type="fig" rid="F3">Figure 3</xref>), stating the unusual dry conditions in 2018 (Vogel et al., <xref ref-type="bibr" rid="B68">2019</xref>). More precipitation is observed during the sampling period of 2019 with an annual sum of 624 mm. Air temperature, relative humidity, air pressure, solar radiation and wind speed show strong diurnal fluctuations during the sampling period (<xref ref-type="fig" rid="F3">Figure 3</xref>). However, during the precipitation periods, these fluctuations tend to decrease particularly in case of air temperature and relative humidity, and tend to increase in case of wind speed. Air temperature, relative humidity and solar radiation represent seasonal trends, whereas air pressure and wind speed do not depict a clear trend during the sampling period.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Time series of meteorological data. From the top, we report precipitation, air temperature, relative humidity, air pressure, solar radiation and wind speed.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frwa-03-740044-g0003.tif"/>
</fig>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Time series of catchment wetness conditions. From the top, we report precipitation, discharge and soil moisture at SM1, SM2, SM3, SM4, SM5 and SM6 stations.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frwa-03-740044-g0004.tif"/>
</fig>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Time series of natural tracers at stream and groundwater sampling sites. From the top, we report precipitation, water temperature, pH and electrical conductivity (EC) at stream and groundwater sampling sites.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frwa-03-740044-g0005.tif"/>
</fig>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Time series of precipitation and stable isotope concentrations at stream and groundwater sampling sites. From the top, we report, precipitation and &#x003B4;<sup>2</sup>H values at stream and groundwater sampling sites.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frwa-03-740044-g0006.tif"/>
</fig>
<p>Stream discharge and soil moisture reflect the drought of 2018 with only few small peaks during the precipitation periods (<xref ref-type="fig" rid="F4">Figure 4</xref>). The catchment wetness increases in 2019 and stream discharge as well as soil moisture react similarly to precipitation inputs. Soil moisture of shallower depth (i.e., 5 cm) generally shows more clear responses to precipitation than at deeper depths (i.e., 15 cm) at all soil moisture stations. This response is more pronounced in the meadow (SM2, SM3 and SM6) than in the farmland (SM1 and SM4). However, soil moisture of deeper depths in the farmland generally reveal a higher responsiveness to rainfall than equivalent depths in the meadow. The shallower depths of SM2 and SM3 at the toeslope depict a similar reaction pattern to SM6 in footslope suggesting the extension of linkage between toeslope and footslope during precipitation events. On average, SM1 and SM3 show the lowest (8 &#x000B1; 4.3%, mean &#x000B1; standard deviation, for 5 cm depth, 14 &#x000B1; 4.6% for 15 cm depth) and highest (29.4 &#x000B1; 11%, 24.3 &#x000B1; 6.8%) soil moisture among the SM stations, respectively.</p>
<p><xref ref-type="fig" rid="F5">Figure 5</xref> displays that stream water temperature strongly matches the diurnal fluctuations of air temperature, whereas the groundwater temperature follows the overall long-term pattern of air temperature without any high-temporal fluctuations. SW1 and SW2 exhibit similar temperature patterns with means of 11.8 &#x000B1; 2.7 and 11.7 &#x000B1; 2.7&#x000B0;C over the sampling period, respectively. However, groundwater temperatures at the sampling sites slightly differs from each other with the mean of 11.3 &#x000B1; 1.8, 13.0 &#x000B1; 2.1 and 12.3 &#x000B1; 1.8&#x000B0;C for GW1, GW2 and GW3, respectively. The patterns of pH dynamics in stream and groundwater sources are almost identical over the sampling period (<xref ref-type="fig" rid="F5">Figure 5</xref>). The pH shows an increasing trend while the air temperature decreases in 2018 and then remains almost stable in 2019 with some fluctuations. The EC of stream and groundwater is highly responsive to precipitation inputs (<xref ref-type="fig" rid="F5">Figure 5</xref>). Rapid reductions of EC are observed upon in the course of precipitation events that are followed by a fast recovery. The EC of stream as well as groundwater is relatively high in 2018, followed by a smooth decline toward the end of the sampling period. The EC of SW1 (336.9 &#x000B1; 59.3 &#x003BC;S cm<sup>&#x02212;1</sup>) is slightly higher than that of SW2 (325.5 &#x000B1; 58.3 &#x003BC;S cm<sup>&#x02212;1</sup>), whereas the difference between the EC of groundwater sampling sites is higher with means of 228.7 &#x000B1; 61.2, 321.9 &#x000B1; 59.9 and 239.3 &#x000B1; 50 &#x003BC;S cm<sup>&#x02212;1</sup> for GW1, GW2 and GW3, respectively.</p>
<p><xref ref-type="fig" rid="F6">Figure 6</xref> presents the high temporal resolution of &#x003B4;<sup>2</sup>H at the stream and groundwater sampling sites. The &#x003B4;<sup>2</sup>H of stream and groundwater responds promptly to precipitation, and is strongly synchronized with the response patterns of EC. These rapid isotopic responses in stream as well as groundwater indicate the strong linkage between stream and subsurface flow pathways during the precipitation events (Sahraei et al., <xref ref-type="bibr" rid="B57">2020</xref>). The &#x003B4;<sup>2</sup>H values are lighter with stronger diurnal variations in 2018 compared to those during sampling period of 2019. Seasonal variations are observed during the sampling period with heavier &#x003B4;<sup>2</sup>H values in May&#x02013;June and lighter in November&#x02013;December. <xref ref-type="fig" rid="F7">Figure 7</xref> shows the distribution of the isotopic compositions in the stream and groundwater sources over the sampling period. Both stream sampling sites indicate similar variations for &#x003B4;<sup>2</sup>H. The mean values of &#x003B4;<sup>2</sup>H for SW1 and SW2 are &#x02212;60.7 &#x000B1; 2.7&#x02030; and &#x02212;60.6 &#x000B1; 2.5&#x02030;, respectively. However, the variations of &#x003B4;<sup>2</sup>H tend to increase from GW1 to GW3. On average, GW1 (&#x02212;60.1 &#x000B1; 2.1&#x02030;) and GW2 (&#x02212;60 &#x000B1; 1.9&#x02030;) at the toeslope depict slightly heavier mean &#x003B4;<sup>2</sup>H values compared to GW3 at the footslope (60.5 &#x000B1; 2.6&#x02030;).</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Boxplots of the isotopic compositions at stream and groundwater sampling sites. The average is indicated by a black square and the median by the bar separating a box.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frwa-03-740044-g0007.tif"/>
</fig>
</sec>
<sec>
<title>LSTM Model Sensitivity to Input Features and Sequence Length</title>
<p>We investigated the impact of input features and sequence length on the prediction performance of the optimized LSTM model. <xref ref-type="fig" rid="F8">Figure 8</xref> shows the heatmaps of the prediction performance using four statistical metrics for the combination of seven input feature scenarios and nine sequence lengths at the stream and groundwater sampling sites (SW1, SW2, GW1, GW2 and GW3). It is apparent that using the meteorological data (scenario S1) and catchment wetness conditions (scenario S2) alone as the input features does not lead to a good prediction performance. Although using meteorological and catchment wetness data together (scenario S4) slightly decreases prediction errors for some cases, yet the performance is not satisfying. For these scenarios, the NSE values are mainly near to zero and even negative. This emphasizes that the model is not able to predict the peak values properly. It suggests that the relations of meteorological data and catchment wetness variables with the water isotope concentrations are not strong enough to transfer enough information that is required for an adequate learning of the LSTM to predict peak values of isotope concentrations.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Heatmaps of the prediction performance using four statistical metrics for the combination of seven input feature scenarios and nine sequence lengths at stream and groundwater sampling sites. The dark blue indicates good prediction performance (low values of RMSE and MAE and high values of R<sup>2</sup> and NSE), whereas the light blue indicates poor prediction performance (high values of RMSE and MAE and low values of <italic>R</italic><sup>2</sup> and NSE).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frwa-03-740044-g0008.tif"/>
</fig>
<p>In contrast, using the natural tracers (scenario S3), i.e., water temperature, pH and EC of the sampling sites, as the input features achieves good prediction performance. It indicates that these easy-to-measure tracers are able to indirectly simulate complex interactions between controlling drivers and water isotope concentrations in the Schwingbach catchment. With a closer look to the prediction errors across the sequence lengths of this scenario, we can see that the model performs the best at all of the sampling sites when we use short sequence lengths (6 and 12 h). This implies that the most recent input data contain the most relevant information for the prediction of isotope concentrations and that there is no significant benefit using older input data. This behavior can also be noticed for the prediction of peak values, which is measured by NSE metric. The NSE values are higher for shorter sequence lengths indicating that the model can better predict the reaction of isotope signatures during precipitation events by using the recent input data. This can be associated to the rapid response characteristics of the Schwingbach catchment (Orlowski et al., <xref ref-type="bibr" rid="B50">2014</xref>, <xref ref-type="bibr" rid="B49">2016</xref>; Sahraei et al., <xref ref-type="bibr" rid="B57">2020</xref>). The responses of isotopes in stream and groundwater during precipitation events often reveal a rapid mixing of water at the sampling sites with the event water. It takes, on average, 6 h for stream water and 7&#x02013;8 h for groundwater at the toeslope (GW1 and GW2) and footslope (GW3), respectively, that the isotope signatures reach peak values. After that, they immediately return to pre-event water conditions (Sahraei et al., <xref ref-type="bibr" rid="B57">2020</xref>). This short-term response behavior together with the synchronized behavior of isotopic compositions with electrical conductivity makes it sufficient for the model to extract the dynamics of isotopic compositions. In contrast, the model performance declines when we use medium and long sequence lengths. This is possibly due to the noise generated by irrelevant inputs (Xiang et al., <xref ref-type="bibr" rid="B73">2020</xref>).</p>
<p>We combined the natural tracers with meteorological and catchment wetness variables (scenarios S5 and S6) to test if this boosts the prediction performance of the LSTM model. Comparing scenario S3 with S5, we can see that the LSTM still performs better at scenario S3 when using short sequence lengths. However, the model tends to perform better at scenario S5 for medium and long sequence lengths. Similarly, the inclusion of catchment wetness (scenario S6) provides some improvements in the model performance for medium and long sequence lengths. The physical explanation underlying the improved performance of the model when we provide longer historical meteorological and catchment wetness information, may be related to the climatological and storage characteristic of the catchment. It could well be that long-term dependencies exist between the dynamics of isotope concentrations and climatic conditions as well as soil water content of the catchment. The ability of the LSTM to learn and remember the long-term dependencies provides the opportunity to identify such storage effects within the catchments (Kratzert et al., <xref ref-type="bibr" rid="B32">2018</xref>, <xref ref-type="bibr" rid="B33">2019</xref>). A comparison between prediction errors of scenario S5 and S6 indicates that the model generally performs better at scenario S5. This suggests that the meteorological condition has a higher impact on the dynamic of isotope concentrations compared to the soil moisture state in the catchment, which is in line with the previous findings in the Schwingbach catchment (Orlowski et al., <xref ref-type="bibr" rid="B50">2014</xref>; Sahraei et al., <xref ref-type="bibr" rid="B57">2020</xref>).</p>
<p>We also tested the prediction performance for all meteorological, catchment wetness and natural tracer variables as input features (scenario S7). In general, the model performance does not improve remarkably and it even deteriorates in some cases compared to performance of model at scenario S3, S5 and S6. The inclusion of too many input features increases the model complexity. Using too complex models augments the chance of overfitting (Maier et al., <xref ref-type="bibr" rid="B39">2010</xref>). Overfitting arises when the model learns the details and noises in the training data to an extent that it is unable to generalize to new data.</p>
</sec>
<sec>
<title>Visualization of the LSTM Prediction Performance</title>
<p>We built a single LSTM model to predict the high-resolution &#x003B4;<sup>2</sup>H values at SW1, SW2, GW1, GW2 and GW3 sampling sites in the Schwingbach catchment. Ideally, we do not want to train and optimize the model specifically for each sampling site to achieve the best performance, but rather, we intend to train and optimize the model once for all of the sampling sites using only one specific input feature scenario and sequence length. Therefore, we trained and optimized the model with the input feature scenario and sequence length, at which the model performs on average the best for all of the sampling sites. According to the results presented in <xref ref-type="fig" rid="F8">Figure 8</xref>, LSTM achieves the best performance when measurements of the last 6 h of water temperature, pH and EC are used as input features (scenario S3). It achieves, on average of all sampling sites, an RMSE of 0.7&#x02030;, MAE of 0.4&#x02030;, <italic>R</italic><sup>2</sup> of 0.9, and NSE of 0.7.</p>
<p>In the following, we visualize the prediction performance of the best LSTM model on the test sets to better illustrate its generalization ability on the unseen data. <xref ref-type="fig" rid="F9">Figures 9</xref>, <xref ref-type="fig" rid="F10">10</xref> show the performance of the LSTM using a sequence length of 6 h for scenario S3 at SW1 and GW3 sampling sites, respectively. The LSTM captures the timing of hydrologic events and base flow conditions of &#x003B4;<sup>2</sup>H at stream and groundwater sites quite well, but it still underestimates the peak of the hydrograph. This is a commonly known issue with LSTMs and in general, ANN models that they cannot learn adequately the phenomenon in respect of extremes. The major reason that the LSTM may not be able to capture extreme values is the lack of a large number of extremes in the training data (Adnan et al., <xref ref-type="bibr" rid="B2">2019</xref>). The other major reason could be the fact that the range of extreme values in the training data is smaller than those of the validation and test data (Adnan et al., <xref ref-type="bibr" rid="B2">2019</xref>; Malik et al., <xref ref-type="bibr" rid="B40">2020</xref>). This results in difficulties for extrapolation in ANN models (Kisi and Aytek, <xref ref-type="bibr" rid="B28">2013</xref>; Kisi and Parmar, <xref ref-type="bibr" rid="B29">2016</xref>). Several scholars have also reported this limitation in the implementation of the LSTM in previous studies (Kratzert et al., <xref ref-type="bibr" rid="B32">2018</xref>; Chen et al., <xref ref-type="bibr" rid="B10">2020</xref>; Fan et al., <xref ref-type="bibr" rid="B14">2020</xref>; M&#x000FC;ller et al., <xref ref-type="bibr" rid="B46">2020</xref>; Xiang et al., <xref ref-type="bibr" rid="B73">2020</xref>). The prediction performances of the best LSTM at SW2, GW1 and GW2 sampling sites are provided in the <xref ref-type="supplementary-material" rid="SM1">Supplementary Figures 1&#x02013;3</xref>. The development of the LSTM model assumed stable boundary conditions with regard to land use and cover. Changes of these during the study period have not been taking place in the SEO. However, such changes could be part of long-term simulations in the LSTM model if information were available on this.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p>Comparison of observed and predicted &#x003B4;<sup>2</sup>H values by the best LSTM using a sequence length of 6 h for water temperature, pH and electrical conductivity measurements (scenario S3-6 h) with an RMSE of 0.9&#x02030;, MAE of 0.5&#x02030;, <italic>R</italic><sup>2</sup> of 0.9 and NSE of 0.6 at SW1 site.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frwa-03-740044-g0009.tif"/>
</fig>
<fig id="F10" position="float">
<label>Figure 10</label>
<caption><p>Comparison of observed and predicted &#x003B4;<sup>2</sup>H values by the best LSTM using a sequence length of 6 h for water temperature, pH and electrical conductivity measurements (scenario S3-6 h) with an RMSE of 0.4&#x02030;, MAE of 0.3&#x02030;, <italic>R</italic><sup>2</sup> of 0.7 and NSE of 0.7 at GW3 site.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frwa-03-740044-g0010.tif"/>
</fig>
<p>The power of LSTM as a deep learning model is that it can synthesize information from multiple sampling sites and situations into a single model. An LSTM trained on multiple sites under different landuse and hillslope conditions can learn different patterns of isotopic behavior. It is evident from <xref ref-type="fig" rid="F9">Figures 9, 10</xref> and <xref ref-type="supplementary-material" rid="SM1">Supplementary Figures 1&#x02013;3</xref> that the LSTM performs efficiently at the stream as well as groundwater sites. Since we trained a single LSTM on stream and groundwater sites with different landuse and hillslope positions with a range of different isotope concentrations and magnitudes of response to precipitation events, the LSTM tends to balance the error so that it minimizes the error between predicted and observed isotope concentrations for all of the sampling sites at the same time. The trained LSTM can be used to spatially estimate the isotope concentrations for multiple sites across the catchment with only simple-to-measure tracers. This leads not only to a substantial reduction of the cost of measurements, but also to an increase in the spatiotemporal knowledge of hydrological processes in the catchment.</p>
<p>One of the challenging tasks in hydrology is the spatial transferability of the models, particularly to data-scarce catchments. Most of hydrological models need to be rebuilt from scratch using newly collected data when the distribution of data is changed in the feature space. However, collecting adequate data is still challenging in many catchments due to the time consuming and expensive measurement procedures. &#x0201C;Transfer learning&#x0201D; is a powerful technique of deep learning that reduces the need and effort to collect the data in data-scarce catchments. Transfer learning is a method that transfers the knowledge obtained in the source domain to the target domain when the latter has few data (Pan and Yang, <xref ref-type="bibr" rid="B51">2010</xref>). Generally, if the deep learning model exhibits satisfying performance in the source catchment, it can be then generalized to target catchments and achieve good prediction performances without much additional data. The transfer learning is efficient in the LSTM because the hidden layers that have been trained to digest shape information, are still effective even when applied to data-scarce catchments (Shen, <xref ref-type="bibr" rid="B60">2018</xref>). In order to apply the proposed LSTM to another catchment, the last fully connected layers are trained on the new data with initial random weights (Yosinski et al., <xref ref-type="bibr" rid="B74">2014</xref>). Although the data is different from the one on which the model is already trained, the low-level features are similar. Therefore, transferring parameters from the &#x0201C;pre-trained LSTM&#x0201D; can provide the new target model with powerful feature extraction ability and reduce the data demand as well as computation or monitoring costs.</p>
<p>The prediction performance of LSTM and in general deep learning models could be enhanced when having more training data available (Schmidhuber, <xref ref-type="bibr" rid="B59">2015</xref>). Providing longer historical data from multiple years for the training process in future research may allow the model to capture long-term seasonal trends under various flow regimes and hydroclimatic conditions, and hence extract more related information that can effectively reproduce the dynamic of the isotope concentrations. It will be particularly interesting to investigate how extreme weather events impact on the LSTM outcome. For the future, we might be able to test this given the current rather wet weather patterns of 2021, which are substantially different to the study period of our current work, which considers rather dry weather conditions from 2018 and 2019.</p>
</sec>
</sec>
<sec sec-type="conclusions" id="s4">
<title>Conclusion</title>
<p>This study investigates the application of Long Short-Term Memory (LSTM) deep learning model to predict high-resolution (3 h) isotope concentrations of multiple stream and groundwater sources under different landuse and hillslope positions in the Schwingbach Environmental Observatory (SEO), Germany. The core objective of this study is to explore the prediction performance of an LSTM that is trained on multiple sites, with a set of explanatory data that are more straightforward and less expensive to measure compared to the stable isotopes of water. The explanatory data consist of meteorological data, catchment wetness conditions and natural tracers (i.e., water temperature, pH and electrical conductivity). This study further conducts a sensitivity analysis to examine how different input features and their sequence lengths affect the performance of the LSTM. The ability of the LSTM that inherently considers the impact of environmental factors such as evaporation and evapotranspiration on fractionation of water isotopes without an explicit representation of the underlying processes provides the opportunity to efficiently apply the proposed model to isotope hydrology. The result showed that the LSTM could successfully predict stable isotopes of stream and groundwater sites when using only short-term sequence (6 h) of measured water temperature, pH and electrical conductivity. The LSTM prediction can be utilized to predict and interpolate the continuous isotope concentration time series either for data gap filling or in case where no continuous data acquisition is feasible. This is very valuable in practice because measurements of these tracers are still much cheaper than stable isotopes of water and can be continuously carried out with relatively minor maintenance.</p>
<p>For the future research, we will be collecting more data to enhance the prediction performance of the LSTM. There is still room to capture the peak isotope concentrations better. New input features like groundwater table should be used to potentially improve the model performance in the future. The data-hungry nature of deep learning models is a potential barrier for applying them in data-scarce catchments. Since many catchments of potential application may lack the length of the data available in this work, the sensitivity of prediction performance to the length of the training data warrants further investigation. The use of pre-trained LSTM is a promising approach to mitigate the large demand for data in a single catchment. The future research direction also includes applying the LSTM to predict high-resolution water quality parameters such as nitrate, pH and water temperature.</p>
</sec>
<sec sec-type="data-availability" id="s5">
<title>Data Availability Statement</title>
<p>The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found below: Zenodo repository, <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5281/zenodo.5084698">https://doi.org/10.5281/zenodo.5084698</ext-link>.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>AS, TH, and LB conceptualized the research. AS performed data acquisition, data analysis and model development, and prepared the manuscript with contributions from all authors. LB and TH reviewed and revised the manuscript. LB supervised the project. All authors have read and approved the content of the manuscript.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s7">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec> </body>
<back>
<sec sec-type="supplementary-material" id="s8">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/frwa.2021.740044/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/frwa.2021.740044/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.PDF" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Abadi</surname> <given-names>M.</given-names></name> <name><surname>Agarwal</surname> <given-names>A.</given-names></name> <name><surname>Barham</surname> <given-names>P.</given-names></name> <name><surname>Brevdo</surname> <given-names>E.</given-names></name> <name><surname>Chen</surname> <given-names>Z.</given-names></name> <name><surname>Citro</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2016</year>). <source>TensorFlow: Large-scale machine learning on heterogeneous distributed systems. arXiv Preprint</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1603.04467">https://arxiv.org/abs/1603.04467</ext-link>.</citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Adnan</surname> <given-names>R. M.</given-names></name> <name><surname>Liang</surname> <given-names>Z.</given-names></name> <name><surname>Heddam</surname> <given-names>S.</given-names></name> <name><surname>Zounemat-Kermani</surname> <given-names>M.</given-names></name> <name><surname>Kisi</surname> <given-names>O.</given-names></name> <name><surname>Li</surname> <given-names>B.</given-names></name></person-group> (<year>2019</year>). <article-title>Least square support vector machine and multivariate adaptive regression splines for streamflow prediction in mountainous basin using hydro-meteorological data as inputs</article-title>. <source>J. Hydrol.</source> <volume>586</volume>:<fpage>124371</fpage>. <pub-id pub-id-type="doi">10.1016/j.jhydrol.2019.124371</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Benardos</surname> <given-names>P. G.</given-names></name> <name><surname>Vosniakos</surname> <given-names>G. C.</given-names></name></person-group> (<year>2007</year>). <article-title>Optimizing feedforward artificial neural network architecture</article-title>. <source>Eng. Appl. Artif. Intell.</source> <volume>20</volume>, <fpage>365</fpage>&#x02013;<lpage>382</lpage>. <pub-id pub-id-type="doi">10.1016/j.engappai.2006.06.005</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bergstra</surname> <given-names>J.</given-names></name> <name><surname>Bardenet</surname> <given-names>R.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>K&#x000E9;gl</surname> <given-names>B.</given-names></name></person-group> (<year>2011</year>). <article-title>Algorithms for hyper-parameter optimization</article-title>. <source>Adv. Neural Inform. Process. Syst.</source> <volume>24</volume>, <fpage>1</fpage>&#x02013;<lpage>9</lpage>.</citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bergstra</surname> <given-names>J.</given-names></name> <name><surname>Yamins</surname> <given-names>D.</given-names></name> <name><surname>Cox</surname> <given-names>D. D.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;Making a Science of Model Search: Hyperparameter optimization in hundreds of dimensions for vision architectures,&#x0201D;</article-title> in <source>International Conference on Machine Learning</source>, <fpage>115</fpage>&#x02013;<lpage>123</lpage>.</citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Berman</surname> <given-names>E. S. F.</given-names></name> <name><surname>Gupta</surname> <given-names>M.</given-names></name> <name><surname>Gabrielli</surname> <given-names>C.</given-names></name> <name><surname>Garland</surname> <given-names>T.</given-names></name> <name><surname>McDonnell</surname> <given-names>J. J.</given-names></name></person-group> (<year>2009</year>). <article-title>High-frequency field-deployable isotope analyzer for hydrological applications</article-title>. <source>Water Resour. Res.</source> <volume>45</volume>, <fpage>1</fpage>&#x02013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1029/2009WR008265</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Birkel</surname> <given-names>C.</given-names></name> <name><surname>Soulsby</surname> <given-names>C.</given-names></name> <name><surname>Tetzlaff</surname> <given-names>D.</given-names></name> <name><surname>Dunn</surname> <given-names>S.</given-names></name> <name><surname>Spezia</surname> <given-names>L.</given-names></name></person-group> (<year>2012</year>). <article-title>High-frequency storm event isotope sampling reveals time-variant transit time distributions and influence of diurnal cycles</article-title>. <source>Hydrol. Process.</source> <volume>26</volume>, <fpage>308</fpage>&#x02013;<lpage>316</lpage>. <pub-id pub-id-type="doi">10.1002/hyp.8210</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Calkoen</surname> <given-names>F.</given-names></name> <name><surname>Luijendijk</surname> <given-names>A.</given-names></name> <name><surname>Rivero</surname> <given-names>C. R.</given-names></name> <name><surname>Kras</surname> <given-names>E.</given-names></name> <name><surname>Baart</surname> <given-names>F.</given-names></name></person-group> (<year>2021</year>). <article-title>Traditional vs. Machine-learning methods for forecasting sandy shoreline evolution using historic satellite-derived shorelines</article-title>. <source>Remote Sens.</source> <volume>13</volume>, <fpage>1</fpage>&#x02013;<lpage>21</lpage>. <pub-id pub-id-type="doi">10.3390/rs13050934</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cerar</surname> <given-names>S.</given-names></name> <name><surname>Mezga</surname> <given-names>K.</given-names></name> <name><surname>&#x0017D;ibret</surname> <given-names>G.</given-names></name> <name><surname>Urbanc</surname> <given-names>J.</given-names></name> <name><surname>Komac</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>Comparison of prediction methods for oxygen-18 isotope composition in shallow groundwater</article-title>. <source>Sci. Total Environ.</source> <volume>631&#x02013;632</volume>, <fpage>358</fpage>&#x02013;<lpage>368</lpage>. <pub-id pub-id-type="doi">10.1016/j.scitotenv.2018.03.033</pub-id><pub-id pub-id-type="pmid">29529429</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Huang</surname> <given-names>J.</given-names></name> <name><surname>Han</surname> <given-names>Z.</given-names></name> <name><surname>Gao</surname> <given-names>H.</given-names></name> <name><surname>Liu</surname> <given-names>M.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>The importance of short lag-time in the runoff forecasting model based on long short-term memory</article-title>. <source>J. Hydrol.</source> <volume>589</volume>:<fpage>125359</fpage>. <pub-id pub-id-type="doi">10.1016/j.jhydrol.2020.125359</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Chollet</surname> <given-names>F.</given-names></name></person-group> (<year>2015</year>). <source>Keras</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://keras.io">https://keras.io</ext-link> (accessed August 25, 2021).</citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Duan</surname> <given-names>S.</given-names></name> <name><surname>Ullrich</surname> <given-names>P.</given-names></name> <name><surname>Shu</surname> <given-names>L.</given-names></name></person-group> (<year>2020</year>). <article-title>Using convolutional neural networks for streamflow projection in California</article-title>. <source>Front. Water</source> <volume>2</volume>:<fpage>28</fpage>. <pub-id pub-id-type="doi">10.3389/frwa.2020.00028</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Eggensperger</surname> <given-names>K.</given-names></name> <name><surname>Feurer</surname> <given-names>M.</given-names></name> <name><surname>Hutter</surname> <given-names>F.</given-names></name> <name><surname>Bergstra</surname> <given-names>J.</given-names></name> <name><surname>Snoek</surname> <given-names>J.</given-names></name> <name><surname>Holger</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>&#x0201C;Towards an empirical foundation for assessing bayesian optimization of hyperparameters,&#x0201D;</article-title> in <source>NIPS Workshop on Bayesian Optimazation in Theory and Practice</source>, <volume>vol 10</volume>, <fpage>3</fpage>.</citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fan</surname> <given-names>H.</given-names></name> <name><surname>Jiang</surname> <given-names>M.</given-names></name> <name><surname>Xu</surname> <given-names>L.</given-names></name> <name><surname>Zhu</surname> <given-names>H.</given-names></name> <name><surname>Cheng</surname> <given-names>J.</given-names></name> <name><surname>Jiang</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Comparison of long short term memory networks and the hydrological model in runoff simulation</article-title>. <source>Water</source> <volume>12</volume>, <fpage>1</fpage>&#x02013;<lpage>17</lpage>. <pub-id pub-id-type="doi">10.3390/w12010175</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fang</surname> <given-names>K.</given-names></name> <name><surname>Shen</surname> <given-names>C.</given-names></name> <name><surname>Kifer</surname> <given-names>D.</given-names></name> <name><surname>Yang</surname> <given-names>X.</given-names></name></person-group> (<year>2017</year>). <article-title>Prolongation of SMAP to spatiotemporally seamless coverage of continental U.S. using a deep learning neural network</article-title>. <source>Geophys. Res. Lett.</source> <volume>44</volume>, <fpage>11,030</fpage>&#x02013;<lpage>11,039</lpage>. <pub-id pub-id-type="doi">10.1002/2017GL075619</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>G&#x000E9;ron</surname> <given-names>A.</given-names></name></person-group> (<year>2019</year>). <source>Hands-On Machine Learning With Scikit-Learn, Keras, And TensorFlow: Concepts, Tools, And Techniques To Build Intelligent Systems</source>. <publisher-loc>Sebastopol, CA</publisher-loc>: <publisher-name>O&#x00027;Reilly Media</publisher-name>.</citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gers</surname> <given-names>F. A.</given-names></name> <name><surname>Eck</surname> <given-names>D.</given-names></name> <name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>2002</year>). <article-title>&#x0201C;Applying LSTM to time series predictable through time-window approaches,&#x0201D;</article-title> in <source>Neural Nets WIRN Vietri-01</source> (<publisher-loc>London</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>193</fpage>&#x02013;<lpage>200</lpage>. <pub-id pub-id-type="doi">10.1007/978-1-4471-0219-9_20</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Goodfellow</surname> <given-names>I.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Courville</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <source>Deep Learning</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation>
</ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hastie</surname> <given-names>T.</given-names></name> <name><surname>Tibshirani</surname> <given-names>R.</given-names></name> <name><surname>Friedman</surname> <given-names>J.</given-names></name></person-group> (<year>2009</year>). <source>The Elements of Statistical Learning: Data Mining, Inference, and Prediction</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer Science and Business Media</publisher-name>. <pub-id pub-id-type="doi">10.1007/978-0-387-84858-7</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Heaton</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <source>Artificial Intelligence for Humans, Volume 3: Deep Learning and Neural Networks</source>. <publisher-loc>Chesterfield, MO</publisher-loc>: <publisher-name>Heaton Research, Inc</publisher-name>.</citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Heinz</surname> <given-names>E.</given-names></name> <name><surname>Kraft</surname> <given-names>P.</given-names></name> <name><surname>Buchen</surname> <given-names>C.</given-names></name> <name><surname>Frede</surname> <given-names>H.-G.</given-names></name> <name><surname>Aquino</surname> <given-names>E.</given-names></name> <name><surname>Breuer</surname> <given-names>L.</given-names></name></person-group> (<year>2014</year>). <article-title>Set up of an automatic water quality sampling system in irrigation agriculture</article-title>. <source>Sensors</source> <volume>14</volume>, <fpage>212</fpage>&#x02013;<lpage>228</lpage>. <pub-id pub-id-type="doi">10.3390/s140100212</pub-id><pub-id pub-id-type="pmid">24366178</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hochreiter</surname> <given-names>S.</given-names></name> <name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>1997</year>). <article-title>Long short-term memory</article-title>. <source>Neural Comput.</source> <volume>9</volume>, <fpage>1735</fpage>&#x02013;<lpage>1780</lpage>. <pub-id pub-id-type="doi">10.1162/neco.1997.9.8.1735</pub-id><pub-id pub-id-type="pmid">9377276</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hunter</surname> <given-names>J. D.</given-names></name></person-group> (<year>2007</year>). <article-title>Matplotlib: A 2D graphics environment</article-title>. <source>IEEE Ann. Hist. Comput.</source> <volume>9</volume>, <fpage>90</fpage>&#x02013;<lpage>95</lpage>. <pub-id pub-id-type="doi">10.1109/MCSE.2007.55</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kendall</surname> <given-names>C.</given-names></name> <name><surname>McDonnell</surname> <given-names>J. J.</given-names></name></person-group> (<year>2012</year>). <source>Isotope Tracers in Catchment Hydrology</source>. <publisher-loc>Amsterdam</publisher-loc>: <publisher-name>Elsevier</publisher-name>.</citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khashei</surname> <given-names>M.</given-names></name> <name><surname>Bijari</surname> <given-names>M.</given-names></name></person-group> (<year>2010</year>). <article-title>An artificial neural network (p, d, q) model for timeseries forecasting</article-title>. <source>Expert Syst. Appl.</source> <volume>37</volume>, <fpage>479</fpage>&#x02013;<lpage>489</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2009.05.044</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khashei</surname> <given-names>M.</given-names></name> <name><surname>Bijari</surname> <given-names>M.</given-names></name> <name><surname>Raissi Ardali</surname> <given-names>G. A.</given-names></name></person-group> (<year>2009</year>). <article-title>Improvement of auto-regressive integrated moving average models using fuzzy logic and Artificial Neural Networks (ANNs)</article-title>. <source>Neurocomputing</source> <volume>72</volume>, <fpage>956</fpage>&#x02013;<lpage>967</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2008.04.017</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D. P.</given-names></name> <name><surname>Ba</surname> <given-names>J. L.</given-names></name></person-group> (<year>2014</year>). <article-title>Adam: a method for stochastic optimization</article-title>. <source>arXiv Preprint</source> Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1412.6980">https://arxiv.org/abs/1412.6980</ext-link>.</citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kisi</surname> <given-names>O.</given-names></name> <name><surname>Aytek</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Explicit neural network in suspended sediment load estimation</article-title>. <source>Neural Netw. World</source> <volume>23</volume>, <fpage>587</fpage>&#x02013;<lpage>607</lpage>. <pub-id pub-id-type="doi">10.14311/NNW.2013.23.035</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kisi</surname> <given-names>O.</given-names></name> <name><surname>Parmar</surname> <given-names>K. S.</given-names></name></person-group> (<year>2016</year>). <article-title>Application of least square support vector machine and multivariate adaptive regression spline models in long term prediction of river water pollution</article-title>. <source>J. Hydrol.</source> <volume>534</volume>, <fpage>104</fpage>&#x02013;<lpage>112</lpage>. <pub-id pub-id-type="doi">10.1016/j.jhydrol.2015.12.014</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Klaus</surname> <given-names>J.</given-names></name> <name><surname>McDonnell</surname> <given-names>J. J.</given-names></name></person-group> (<year>2013</year>). <article-title>Hydrograph separation using stable isotopes: review and evaluation</article-title>. <source>J. Hydrol.</source> <volume>505</volume>, <fpage>47</fpage>&#x02013;<lpage>64</lpage>. <pub-id pub-id-type="doi">10.1016/j.jhydrol.2013.09.006</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koeniger</surname> <given-names>P.</given-names></name> <name><surname>Gaj</surname> <given-names>M.</given-names></name> <name><surname>Beyer</surname> <given-names>M.</given-names></name> <name><surname>Himmelsbach</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <article-title>Review on soil water isotope-based groundwater recharge estimations</article-title>. <source>Hydrol. Process.</source> <volume>30</volume>, <fpage>2817</fpage>&#x02013;<lpage>2834</lpage>. <pub-id pub-id-type="doi">10.1002/hyp.10775</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kratzert</surname> <given-names>F.</given-names></name> <name><surname>Klotz</surname> <given-names>D.</given-names></name> <name><surname>Brenner</surname> <given-names>C.</given-names></name> <name><surname>Schulz</surname> <given-names>K.</given-names></name> <name><surname>Herrnegger</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>Rainfall-runoff modelling using Long Short-Term Memory (LSTM) networks</article-title>. <source>Hydrol. Earth Syst. Sci.</source> <volume>22</volume>, <fpage>6005</fpage>&#x02013;<lpage>6022</lpage>. <pub-id pub-id-type="doi">10.5194/hess-22-6005-2018</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kratzert</surname> <given-names>F.</given-names></name> <name><surname>Klotz</surname> <given-names>D.</given-names></name> <name><surname>Shalev</surname> <given-names>G.</given-names></name> <name><surname>Klambauer</surname> <given-names>G.</given-names></name> <name><surname>Hochreiter</surname> <given-names>S.</given-names></name> <name><surname>Nearing</surname> <given-names>G.</given-names></name></person-group> (<year>2019</year>). <article-title>Towards learning universal, regional, and local hydrological behaviors via machine learning applied to large-sample datasets</article-title>. <source>Hydrol. Earth Syst. Sci.</source> <volume>23</volume>, <fpage>5089</fpage>&#x02013;<lpage>5110</lpage>. <pub-id pub-id-type="doi">10.5194/hess-23-5089-2019</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lange</surname> <given-names>H.</given-names></name> <name><surname>Sippel</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Machine learning applications in hydrology,&#x0201D;</article-title> in <source>Forest-Water Interactions</source>, eds. D. F. Levia, D. E. Carlyle-Moses, S. Iida, and B. Michalzik (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>233</fpage>&#x02013;<lpage>257</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-26086-6_10</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Le</surname> <given-names>X.</given-names></name> <name><surname>Ho</surname> <given-names>H. V.</given-names></name> <name><surname>Lee</surname> <given-names>G.</given-names></name> <name><surname>Jung</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>Application of Long Short-Term Memory (LSTM) neural network for flood forecasting</article-title>. <source>Water</source> <volume>11</volume>:<fpage>1387</fpage>. <pub-id pub-id-type="doi">10.3390/w11071387</pub-id></citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>P.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Sangaiah</surname> <given-names>A.</given-names></name> <name><surname>Xie</surname> <given-names>Y.</given-names></name> <name><surname>Yin</surname> <given-names>X.</given-names></name></person-group> (<year>2019</year>). <article-title>Analysis and prediction of water quality using LSTM deep neural networks in IoT environment</article-title>. <source>Sustainability</source> <volume>11</volume>:<fpage>2058</fpage>. <pub-id pub-id-type="doi">10.3390/su11072058</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mahindawansha</surname> <given-names>A.</given-names></name> <name><surname>Breuer</surname> <given-names>L.</given-names></name> <name><surname>Chamorro</surname> <given-names>A.</given-names></name> <name><surname>Kraft</surname> <given-names>P.</given-names></name></person-group> (<year>2018</year>). <article-title>High-frequency water isotopic analysis using an automatic water sampling system in rice-based cropping systems</article-title>. <source>Water</source> <volume>10</volume>:<fpage>1327</fpage>. <pub-id pub-id-type="doi">10.3390/w10101327</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maier</surname> <given-names>H. R.</given-names></name> <name><surname>Dandy</surname> <given-names>G. C.</given-names></name></person-group> (<year>2000</year>). <article-title>Neural networks for the prediction and forecasting of water resources variables: a review of modelling issues and applications</article-title>. <source>Environ. Model. Softw.</source> <volume>15</volume>, <fpage>101</fpage>&#x02013;<lpage>124</lpage>. <pub-id pub-id-type="doi">10.1016/S1364-8152(99)00007-9</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maier</surname> <given-names>H. R.</given-names></name> <name><surname>Jain</surname> <given-names>A.</given-names></name> <name><surname>Dandy</surname> <given-names>G. C.</given-names></name> <name><surname>Sudheer</surname> <given-names>K. P.</given-names></name></person-group> (<year>2010</year>). <article-title>Methods used for the development of neural networks for the prediction of water resource variables in river systems: current status and future directions</article-title>. <source>Environ. Model. Softw.</source> <volume>25</volume>, <fpage>891</fpage>&#x02013;<lpage>909</lpage>. <pub-id pub-id-type="doi">10.1016/j.envsoft.2010.02.003</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Malik</surname> <given-names>A.</given-names></name> <name><surname>Tikhamarine</surname> <given-names>Y.</given-names></name> <name><surname>Souag-Gamane</surname> <given-names>D.</given-names></name> <name><surname>Kisi</surname> <given-names>O.</given-names></name> <name><surname>Pham</surname> <given-names>Q. B.</given-names></name></person-group> (<year>2020</year>). <article-title>Support vector regression optimized by meta-heuristic algorithms for daily streamflow prediction</article-title>. <source>Stoch. Environ. Res. Risk Assess.</source> <volume>34</volume>, <fpage>1755</fpage>&#x02013;<lpage>1773</lpage>. <pub-id pub-id-type="doi">10.1007/s00477-020-01874-1</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McGlynn</surname> <given-names>B. L.</given-names></name> <name><surname>McDonnell</surname> <given-names>J. J.</given-names></name> <name><surname>Seibert</surname> <given-names>J.</given-names></name> <name><surname>Kendall</surname> <given-names>C.</given-names></name></person-group> (<year>2004</year>). <article-title>Scale effects on headwater catchment runoff timing, flow sources, and groundwater-streamflow relations</article-title>. <source>Water Resour. Res.</source> <fpage>40</fpage>. <pub-id pub-id-type="doi">10.1029/2003WR002494</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McGuire</surname> <given-names>K. J.</given-names></name> <name><surname>McDonnell</surname> <given-names>J. J.</given-names></name></person-group> (<year>2006</year>). <article-title>A review and evaluation of catchment transit time modeling</article-title>. <source>J. Hydrol.</source> <volume>330</volume>, <fpage>543</fpage>&#x02013;<lpage>563</lpage>. <pub-id pub-id-type="doi">10.1016/j.jhydrol.2006.04.020</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>McKinney</surname> <given-names>W.</given-names></name></person-group> (<year>2010</year>). <article-title>&#x0201C;Data structures for statistical computing in python,&#x0201D;</article-title> in <source>Proceedings of the 9th Python in Science Conference</source> (<publisher-loc>Austin, TX</publisher-loc>), <fpage>56</fpage>&#x02013;<lpage>61</lpage>. <pub-id pub-id-type="doi">10.25080/Majora-92bf1922-00a</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Meyal</surname> <given-names>A. Y.</given-names></name> <name><surname>Versteeg</surname> <given-names>R.</given-names></name> <name><surname>Alper</surname> <given-names>E.</given-names></name> <name><surname>Johnson</surname> <given-names>D.</given-names></name></person-group> (<year>2020</year>). <article-title>Automated cloud based long short-term memory neural network based SWE prediction</article-title>. <source>Front. Water</source> <volume>2</volume>:<fpage>574917</fpage>. <pub-id pub-id-type="doi">10.3389/frwa.2020.574917</pub-id></citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moriasi</surname> <given-names>D. N.</given-names></name> <name><surname>Arnold</surname> <given-names>J. G.</given-names></name> <name><surname>Van Liew</surname> <given-names>M. W.</given-names></name> <name><surname>Bingner</surname> <given-names>R. L.</given-names></name> <name><surname>Harmel</surname> <given-names>R. D.</given-names></name> <name><surname>Veith</surname> <given-names>T. L.</given-names></name></person-group> (<year>2007</year>). <article-title>Model evaluation guidelines for systematic quantification of accuracy in watershed simulations</article-title>. <source>Trans. Am. Soc. Agric. Biol. Eng.</source> <volume>50</volume>, <fpage>885</fpage>&#x02013;<lpage>900</lpage>. <pub-id pub-id-type="doi">10.13031/2013.23153</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>M&#x000FC;ller</surname> <given-names>J.</given-names></name> <name><surname>Park</surname> <given-names>J.</given-names></name> <name><surname>Sahu</surname> <given-names>R.</given-names></name> <name><surname>Varadharajan</surname> <given-names>C.</given-names></name> <name><surname>Arora</surname> <given-names>B.</given-names></name> <name><surname>Faybishenko</surname> <given-names>B.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Surrogate optimization of deep neural networks for groundwater predictions</article-title>. <source>J. Glob. Optim.</source> <volume>81</volume>, <fpage>203</fpage>&#x02013;<lpage>231</lpage>. <pub-id pub-id-type="doi">10.1007/s10898-020-00912-0</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nakisa</surname> <given-names>B.</given-names></name> <name><surname>Rastgoo</surname> <given-names>M. N.</given-names></name> <name><surname>Rakotonirainy</surname> <given-names>A.</given-names></name> <name><surname>Maire</surname> <given-names>F.</given-names></name> <name><surname>Chandran</surname> <given-names>V.</given-names></name></person-group> (<year>2018</year>). <article-title>Long short term memory hyperparameter optimization for a neural network based emotion recognition framework</article-title>. <source>IEEE Access</source> <volume>6</volume>, <fpage>49325</fpage>&#x02013;<lpage>49338</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2018.2868361</pub-id></citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nash</surname> <given-names>J. E.</given-names></name> <name><surname>Sutcliffe</surname> <given-names>J. V.</given-names></name></person-group> (<year>1970</year>). <article-title>River flow forecasting through conceptual models part I&#x02014;A discussion of principles</article-title>. <source>J. Hydrol.</source> <volume>10</volume>, <fpage>282</fpage>&#x02013;<lpage>290</lpage>. <pub-id pub-id-type="doi">10.1016/0022-1694(70)90255-6</pub-id></citation>
</ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Orlowski</surname> <given-names>N.</given-names></name> <name><surname>Kraft</surname> <given-names>P.</given-names></name> <name><surname>Pferdmenges</surname> <given-names>J.</given-names></name> <name><surname>Breuer</surname> <given-names>L.</given-names></name></person-group> (<year>2016</year>). <article-title>Exploring water cycle dynamics by sampling multiple stable water isotope pools in a developed landscape in Germany</article-title>. <source>Hydrol. Earth Syst. Sci.</source> <volume>20</volume>, <fpage>3873</fpage>&#x02013;<lpage>3894</lpage>. <pub-id pub-id-type="doi">10.5194/hess-20-3873-2016</pub-id></citation>
</ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Orlowski</surname> <given-names>N.</given-names></name> <name><surname>Lauer</surname> <given-names>F.</given-names></name> <name><surname>Kraft</surname> <given-names>P.</given-names></name> <name><surname>Frede</surname> <given-names>H. G.</given-names></name> <name><surname>Breuer</surname> <given-names>L.</given-names></name></person-group> (<year>2014</year>). <article-title>Linking spatial patterns of groundwater table dynamics and streamflow generation processes in a small developed catchment</article-title>. <source>Water</source> <volume>6</volume>, <fpage>3085</fpage>&#x02013;<lpage>3117</lpage>. <pub-id pub-id-type="doi">10.3390/w6103085</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pan</surname> <given-names>S. J.</given-names></name> <name><surname>Yang</surname> <given-names>Q.</given-names></name></person-group> (<year>2010</year>). <article-title>A survey on transfer learning</article-title>. <source>IEEE Trans. Knowl. Data Eng.</source> <volume>22</volume>, <fpage>1345</fpage>&#x02013;<lpage>1359</lpage>. <pub-id pub-id-type="doi">10.1109/TKDE.2009.191</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pedregosa</surname> <given-names>F.</given-names></name> <name><surname>Varoquaux</surname> <given-names>G.</given-names></name> <name><surname>Gramfort</surname> <given-names>A.</given-names></name> <name><surname>Michel</surname> <given-names>V.</given-names></name> <name><surname>Thirion</surname> <given-names>B.</given-names></name> <name><surname>Grisel</surname> <given-names>O.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Scikit-learn: machine learning in python</article-title>. <source>J. Mach. Learn. Res.</source> <volume>12</volume>, <fpage>2825</fpage>&#x02013;<lpage>2830</lpage>.</citation>
</ref>
<ref id="B53">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Pumperla</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <source>Hyperas</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://pypi.org/project/hyperas/">https://pypi.org/project/hyperas/</ext-link> (accessed May 10, 2021).</citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Quade</surname> <given-names>M.</given-names></name> <name><surname>Klosterhalfen</surname> <given-names>A.</given-names></name> <name><surname>Graf</surname> <given-names>A.</given-names></name> <name><surname>Br&#x000FC;ggemann</surname> <given-names>N.</given-names></name> <name><surname>Hermes</surname> <given-names>N.</given-names></name> <name><surname>Vereecken</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>In-situ monitoring of soil water isotopic composition for partitioning of evapotranspiration during one growing season of sugar beet (Beta vulgaris)</article-title>. <source>Agric. For. Meteorol.</source> <volume>266&#x02013;267</volume>, <fpage>53</fpage>&#x02013;<lpage>64</lpage>. <pub-id pub-id-type="doi">10.1016/j.agrformet.2018.12.002</pub-id></citation>
</ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Reimers</surname> <given-names>N.</given-names></name> <name><surname>Gurevych</surname> <given-names>I.</given-names></name></person-group> (<year>2017</year>). <article-title>Optimal hyperparameters for deep LSTM-networks for sequence labeling tasks</article-title>. arXiv Preprint. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1707.06799">https://arxiv.org/abs/1707.06799</ext-link>.</citation>
</ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sahraei</surname> <given-names>A.</given-names></name> <name><surname>Chamorro</surname> <given-names>A.</given-names></name> <name><surname>Kraft</surname> <given-names>P.</given-names></name> <name><surname>Breuer</surname> <given-names>L.</given-names></name></person-group> (<year>2021</year>). <article-title>Application of machine learning models to predict maximum event water fractions in streamflow</article-title>. <source>Front. Water</source> <volume>3</volume>:<fpage>652100</fpage>. <pub-id pub-id-type="doi">10.3389/frwa.2021.652100</pub-id></citation>
</ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sahraei</surname> <given-names>A.</given-names></name> <name><surname>Kraft</surname> <given-names>P.</given-names></name> <name><surname>Windhorst</surname> <given-names>D.</given-names></name> <name><surname>Breuer</surname> <given-names>L.</given-names></name></person-group> (<year>2020</year>). <article-title>High-resolution, in situ monitoring of stable isotopes of water revealed insight into hydrological behavior</article-title>. <source>Water</source> <volume>12</volume>:<fpage>565</fpage>. <pub-id pub-id-type="doi">10.3390/w12020565</pub-id></citation>
</ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Samuel</surname> <given-names>A. L.</given-names></name></person-group> (<year>1959</year>). <article-title>Some studies in machine learning using the game of checkers</article-title>. <source>IBM J. Res. Dev.</source> <volume>3</volume>, <fpage>210</fpage>&#x02013;<lpage>299</lpage>. <pub-id pub-id-type="doi">10.1147/rd.33.0210</pub-id></citation>
</ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep Learning in neural networks: an overview</article-title>. <source>Neural Netw.</source> <volume>61</volume>, <fpage>85</fpage>&#x02013;<lpage>117</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2014.09.003</pub-id><pub-id pub-id-type="pmid">25462637</pub-id></citation></ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shen</surname> <given-names>C.</given-names></name></person-group> (<year>2018</year>). <article-title>A transdisciplinary review of deep learning research and its relevance for water resources scientists</article-title>. <source>Water Resour. Res.</source> <volume>54</volume>, <fpage>8558</fpage>&#x02013;<lpage>8593</lpage>. <pub-id pub-id-type="doi">10.1029/2018WR022643</pub-id></citation>
</ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sprenger</surname> <given-names>M.</given-names></name> <name><surname>Leistert</surname> <given-names>H.</given-names></name> <name><surname>Gimbel</surname> <given-names>K.</given-names></name> <name><surname>Weiler</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Illuminating hydrological processes at the soil-vegetation- atmosphere interface with water stable isotopes</article-title>. <source>Rev. Geophys.</source> <volume>54</volume>, <fpage>674</fpage>&#x02013;<lpage>704</lpage>. <pub-id pub-id-type="doi">10.1002/2015RG000515</pub-id></citation>
</ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tennant</surname> <given-names>C.</given-names></name> <name><surname>Larsen</surname> <given-names>L.</given-names></name> <name><surname>Bellugi</surname> <given-names>D.</given-names></name> <name><surname>Moges</surname> <given-names>E.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Ma</surname> <given-names>H.</given-names></name></person-group> (<year>2020</year>). <article-title>The utility of information flow in formulating discharge forecast models: a case study from an arid snow-dominated catchment</article-title>. <source>Water Resour. Res.</source> <volume>56</volume>, <fpage>1</fpage>&#x02013;<lpage>21</lpage>. <pub-id pub-id-type="doi">10.1029/2019WR024908</pub-id></citation>
</ref>
<ref id="B63">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tetzlaff</surname> <given-names>D.</given-names></name> <name><surname>Buttle</surname> <given-names>J.</given-names></name> <name><surname>Carey</surname> <given-names>S. K.</given-names></name> <name><surname>Mcguire</surname> <given-names>K.</given-names></name> <name><surname>Laudon</surname> <given-names>H.</given-names></name> <name><surname>Soulsby</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>Tracer-based assessment of flow paths, storage and runoff generation in northern catchments: a review</article-title>. <source>Hydrol. Process.</source> <volume>29</volume>, <fpage>3475</fpage>&#x02013;<lpage>3490</lpage>. <pub-id pub-id-type="doi">10.1002/hyp.10412</pub-id></citation>
</ref>
<ref id="B64">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Uhlenbrook</surname> <given-names>S.</given-names></name> <name><surname>Frey</surname> <given-names>M.</given-names></name> <name><surname>Leibundgut</surname> <given-names>C.</given-names></name> <name><surname>Maloszewski</surname> <given-names>P.</given-names></name></person-group> (<year>2002</year>). <article-title>Hydrograph separations in a mesoscale mountainous basin at event and seasonal timescales</article-title>. <source>Water Resour. Res.</source> <volume>38</volume>, <fpage>31-1</fpage>&#x02013;<lpage>31-14</lpage>. <pub-id pub-id-type="doi">10.1029/2001WR000938</pub-id></citation>
</ref>
<ref id="B65">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Van Der Walt</surname> <given-names>S.</given-names></name> <name><surname>Colbert</surname> <given-names>S. C.</given-names></name> <name><surname>Varoquaux</surname> <given-names>G.</given-names></name></person-group> (<year>2011</year>). <article-title>The NumPy array: a structure for efficient numerical computation</article-title>. <source>Comput. Sci. Eng.</source> <volume>13</volume>, <fpage>22</fpage>&#x02013;<lpage>30</lpage>. <pub-id pub-id-type="doi">10.1109/MCSE.2011.37</pub-id></citation>
</ref>
<ref id="B66">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>van Rossum</surname> <given-names>G.</given-names></name></person-group> (<year>1995</year>). <source>Python Tutorial</source>. Amsterdam. Centrum voor Wiskunde en Informatica (CWI).</citation>
</ref>
<ref id="B67">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Virtanen</surname> <given-names>P.</given-names></name> <name><surname>Gommers</surname> <given-names>R.</given-names></name> <name><surname>Oliphant</surname> <given-names>T. E.</given-names></name> <name><surname>Haberland</surname> <given-names>M.</given-names></name> <name><surname>Reddy</surname> <given-names>T.</given-names></name> <name><surname>Cournapeau</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>SciPy 1.0: fundamental algorithms for scientific computing in Python</article-title>. <source>Nat. Methods</source> <volume>17</volume>, <fpage>261</fpage>&#x02013;<lpage>272</lpage>. <pub-id pub-id-type="doi">10.1038/s41592-019-0686-2</pub-id><pub-id pub-id-type="pmid">32094914</pub-id></citation></ref>
<ref id="B68">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vogel</surname> <given-names>M. M.</given-names></name> <name><surname>Zscheischler</surname> <given-names>J.</given-names></name> <name><surname>Wartenburger</surname> <given-names>R.</given-names></name> <name><surname>Dee</surname> <given-names>D.</given-names></name> <name><surname>Seneviratne</surname> <given-names>S. I.</given-names></name></person-group> (<year>2019</year>). <article-title>Concurrent 2018 hot extremes across northern hemisphere due to human-induced climate change</article-title>. <source>Earth&#x00027;s Futur.</source> <volume>7</volume>, <fpage>692</fpage>&#x02013;<lpage>703</lpage>. <pub-id pub-id-type="doi">10.1029/2019EF001189</pub-id><pub-id pub-id-type="pmid">31598535</pub-id></citation></ref>
<ref id="B69">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>von Freyberg</surname> <given-names>J.</given-names></name> <name><surname>Studer</surname> <given-names>B.</given-names></name> <name><surname>Kirchner</surname> <given-names>J. W.</given-names></name></person-group> (<year>2017</year>). <article-title>A lab in the field: high-frequency analysis of water quality and stable isotopes in stream water and precipitation</article-title>. <source>Hydrol. Earth Syst. Sci.</source> <volume>21</volume>, <fpage>1721</fpage>&#x02013;<lpage>1739</lpage>. <pub-id pub-id-type="doi">10.5194/hess-21-1721-2017</pub-id></citation>
</ref>
<ref id="B70">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Waskom</surname> <given-names>M.</given-names></name> <name><surname>Botvinnik</surname> <given-names>O.</given-names></name> <name><surname>Gelbart</surname> <given-names>M.</given-names></name> <name><surname>Ostblom</surname> <given-names>J.</given-names></name> <name><surname>Hobson</surname> <given-names>P.</given-names></name> <name><surname>Lukauskas</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Seaborn: statistical data visualization</article-title>. <source>Astrophys. Source Code Libr</source>. <fpage>1</fpage>&#x02013;<lpage>4</lpage>. <pub-id pub-id-type="doi">10.5281/zenodo.592845</pub-id></citation>
</ref>
<ref id="B71">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Windhorst</surname> <given-names>D.</given-names></name> <name><surname>Kraft</surname> <given-names>P.</given-names></name> <name><surname>Timbe</surname> <given-names>E.</given-names></name> <name><surname>Frede</surname> <given-names>H. G.</given-names></name> <name><surname>Breuer</surname> <given-names>L.</given-names></name></person-group> (<year>2014</year>). <article-title>Stable water isotope tracing through hydrological models for disentangling runoff generation processes at the hillslope scale</article-title>. <source>Hydrol. Earth Syst. Sci.</source> <volume>18</volume>, <fpage>4113</fpage>&#x02013;<lpage>4127</lpage>. <pub-id pub-id-type="doi">10.5194/hess-18-4113-2014</pub-id></citation>
</ref>
<ref id="B72">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wissmeier</surname> <given-names>L.</given-names></name> <name><surname>Uhlenbrook</surname> <given-names>S.</given-names></name></person-group> (<year>2007</year>). <article-title>Distributed, high-resolution modelling of 18O signals in a meso-scale catchment</article-title>. <source>J. Hydrol.</source> <volume>332</volume>, <fpage>497</fpage>&#x02013;<lpage>510</lpage>. <pub-id pub-id-type="doi">10.1016/j.jhydrol.2006.08.003</pub-id></citation>
</ref>
<ref id="B73">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xiang</surname> <given-names>Z.</given-names></name> <name><surname>Yan</surname> <given-names>J.</given-names></name> <name><surname>Demir</surname> <given-names>I.</given-names></name></person-group> (<year>2020</year>). <article-title>A rainfall-runoff model with LSTM-based sequence-to-sequence learning</article-title>. <source>Water Environ. Res.</source> <fpage>56</fpage>. <pub-id pub-id-type="doi">10.1029/2019WR025326</pub-id></citation>
</ref>
<ref id="B74">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yosinski</surname> <given-names>J.</given-names></name> <name><surname>Clune</surname> <given-names>J.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Lipson</surname> <given-names>H.</given-names></name></person-group> (<year>2014</year>). <article-title>How transferable are features in deep neural networks?</article-title> arXiv Preprint. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1411.1792">https://arxiv.org/abs/1411.1792</ext-link>.<pub-id pub-id-type="pmid">30935654</pub-id></citation></ref>
<ref id="B75">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>D.</given-names></name> <name><surname>Lin</surname> <given-names>J.</given-names></name> <name><surname>Peng</surname> <given-names>Q.</given-names></name> <name><surname>Wang</surname> <given-names>D.</given-names></name> <name><surname>Yang</surname> <given-names>T.</given-names></name> <name><surname>Sorooshian</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2018a</year>). <article-title>Modeling and simulating of reservoir operation using the artificial neural network, support vector regression, deep learning algorithm</article-title>. <source>J. Hydrol.</source> <volume>565</volume>, <fpage>720</fpage>&#x02013;<lpage>736</lpage>. <pub-id pub-id-type="doi">10.1016/j.jhydrol.2018.08.050</pub-id></citation>
</ref>
<ref id="B76">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>G.</given-names></name> <name><surname>Eddy Patuwo</surname> <given-names>B. Y.</given-names></name> <name><surname>Hu</surname> <given-names>M.</given-names></name></person-group> (<year>1998</year>). <article-title>Forecasting with artificial neural networks: the state of the art</article-title>. <source>Int. J. Forecast.</source> <volume>14</volume>, <fpage>35</fpage>&#x02013;<lpage>62</lpage>. <pub-id pub-id-type="doi">10.1016/S0169-2070(97)00044-7</pub-id></citation>
</ref>
<ref id="B77">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Zhu</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ye</surname> <given-names>M.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name></person-group> (<year>2018b</year>). <article-title>Developing a Long Short-Term Memory (LSTM) based model for predicting water table depth in agricultural areas</article-title>. <source>J. Hydrol.</source> <volume>561</volume>, <fpage>918</fpage>&#x02013;<lpage>929</lpage>. <pub-id pub-id-type="doi">10.1016/j.jhydrol.2018.04.065</pub-id></citation>
</ref>
<ref id="B78">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>J.</given-names></name> <name><surname>Liu</surname> <given-names>G.</given-names></name> <name><surname>Meng</surname> <given-names>Y.</given-names></name> <name><surname>Xia</surname> <given-names>C.</given-names></name> <name><surname>Chen</surname> <given-names>K.</given-names></name> <name><surname>Chen</surname> <given-names>Y.</given-names></name></person-group> (<year>2021</year>). <article-title>Using stable isotopes as tracer to investigate hydrological condition and estimate water residence time in a plain region, Chengdu, China</article-title>. <source>Sci. Rep.</source> <volume>11</volume>:<fpage>2812</fpage>. <pub-id pub-id-type="doi">10.1038/s41598-021-82349-3</pub-id><pub-id pub-id-type="pmid">33531607</pub-id></citation></ref>
<ref id="B79">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zounemat-Kermani</surname> <given-names>M.</given-names></name> <name><surname>Matta</surname> <given-names>E.</given-names></name> <name><surname>Cominola</surname> <given-names>A.</given-names></name> <name><surname>Xia</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>Q.</given-names></name> <name><surname>Liang</surname> <given-names>Q.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Neurocomputing in surface water hydrology and hydraulics: a review of two decades retrospective, current status and future prospects</article-title>. <source>J. Hydrol.</source> <pub-id pub-id-type="doi">10.1016/j.jhydrol.2020.125085</pub-id></citation>
</ref>
<ref id="B80">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zuo</surname> <given-names>R.</given-names></name> <name><surname>Xiong</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Carranza</surname> <given-names>E. J. M.</given-names></name></person-group> (<year>2019</year>). <article-title>Deep learning and its application in geochemical mapping</article-title>. <source>Earth-Sci. Rev.</source> <volume>192</volume>, <fpage>1</fpage>&#x02013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1016/j.earscirev.2019.02.023</pub-id></citation>
</ref>
</ref-list>
</back>
</article>