<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Energy Res.</journal-id>
<journal-title>Frontiers in Energy Research</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Energy Res.</abbrev-journal-title>
<issn pub-type="epub">2296-598X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">752185</article-id>
<article-id pub-id-type="doi">10.3389/fenrg.2021.752185</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Energy Research</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Accurate and Rapid Forecasts for Geologic Carbon Storage <italic>via</italic> Learning-Based Inversion-Free Prediction</article-title>
<alt-title alt-title-type="left-running-head">Lu et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Rapid Geologic Carbon Storage Forecasts</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Lu</surname>
<given-names>Dan</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1406459/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Painter</surname>
<given-names>Scott L.</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/873525/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Azzolina</surname>
<given-names>Nicholas A.</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Burton-Kelly</surname>
<given-names>Matthew</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Jiang</surname>
<given-names>Tao</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Williamson</surname>
<given-names>Cody</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Computational Sciences and Engineering Division, Oak Ridge National Laboratory</institution>, <addr-line>Oak Ridge</addr-line>, <addr-line>TN</addr-line>, <country>United States</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Environmental Sciences Division, Oak Ridge National Laboratory</institution>, <addr-line>Oak Ridge</addr-line>, <addr-line>TN</addr-line>, <country>United&#x20;States</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>Energy and Environmental Research Center, University of North Dakota</institution>, <addr-line>Grand Forks</addr-line>, <addr-line>ND</addr-line>, <country>United&#x20;States</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/189946/overview">Greeshma Gadikota</ext-link>, Cornell University, United&#x20;States</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1097712/overview">Hussein Hoteit</ext-link>, King Abdullah University of Science and Technology, Saudi Arabia</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/178072/overview">Christine Doughty</ext-link>, Lawrence Berkeley National Laboratory, United&#x20;States</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Dan Lu, <email>lud1@ornl.gov</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Carbon Capture, Utilization and Storage, a section of the journal Frontiers in Energy Research</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>12</day>
<month>01</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>9</volume>
<elocation-id>752185</elocation-id>
<history>
<date date-type="received">
<day>02</day>
<month>08</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>26</day>
<month>11</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Lu, Painter, Azzolina, Burton-Kelly, Jiang and Williamson.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Lu, Painter, Azzolina, Burton-Kelly, Jiang and Williamson</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Carbon capture and storage (CCS) is one approach being studied by the U.S. Department of Energy to help mitigate global warming. The process involves capturing CO<sub>2</sub> emissions from industrial sources and permanently storing them in deep geologic formations (storage reservoirs). However, CCS projects generally target &#x201c;green field sites,&#x201d; where there is often little characterization data and therefore large uncertainty about the petrophysical properties and other geologic attributes of the storage reservoir. Consequently, ensemble-based approaches are often used to forecast multiple realizations prior to CO<sub>2</sub> injection to visualize a range of potential outcomes. In addition, monitoring data during injection operations are used to update the pre-injection forecasts and thereby improve agreement between forecasted and observed behavior. Thus, a system for generating accurate, timely forecasts of pressure buildup and CO<sub>2</sub> movement and distribution within the storage reservoir and for updating those forecasts <italic>via</italic> monitoring measurements becomes crucial. This study proposes a learning-based prediction method that can accurately and rapidly forecast spatial distribution of CO<sub>2</sub> concentration and pressure with uncertainty quantification without relying on traditional inverse modeling. The machine learning techniques include dimension reduction, multivariate data analysis, and Bayesian learning. The outcome is expected to provide CO<sub>2</sub> storage site operators with an effective tool for timely and informative decision making based on limited simulation and monitoring&#x20;data.</p>
</abstract>
<kwd-group>
<kwd>carbon capture and storage</kwd>
<kwd>saline formations</kwd>
<kwd>machine learning</kwd>
<kwd>Bayesian inference</kwd>
<kwd>dimension reduction</kwd>
<kwd>accurate and rapid forecasts</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Carbon capture and storage (CCS) has been proposed as a strategy to reduce greenhouse gas emissions entering the atmosphere from stationary sources and thereby help to mitigate the global climate crisis (<xref ref-type="bibr" rid="B19">Pacala and Socolow, 2004</xref>; <xref ref-type="bibr" rid="B1">Alcalde et&#x20;al., 2018</xref>). For example, the Intergovernmental Panel on Climate Change (IPCC) estimated that capturing CO<sub>2</sub> at a modern conventional power plant could reduce CO<sub>2</sub> emissions to the atmosphere by approximately 80&#x2013;90% compared to a plant that does not have the technology to capture carbon (IPCC report, <xref ref-type="bibr" rid="B16">Metz et&#x20;al. (2005)</xref>). Once the CO<sub>2</sub> has been captured, it must be permanently stored and isolated from the atmosphere, and carbon storage in geological formations is a proven method to store CO<sub>2</sub> at significant (commercial) scales, e.g., one million metric tons per year or greater. The U.S. Department of Energy&#x2019;s (DOE) National Energy Technology Laboratory (NETL) has been working with Regional Carbon Sequestration Partnerships through the Carbon Storage Program to identify prospective sites within the United&#x20;States for the geologic storage of CO<sub>2</sub> (DOE-NETL about the carbon storage program, 2021). Since 2007, NETL has published several assessments of CO<sub>2</sub> storage resource potential in geologic formations and terrestrial sinks in the United&#x20;States, considering the following geologic formations as viable targets for CO<sub>2</sub> storage: saline formations, coal seams, conventional hydrocarbon reservoirs, basalt formations, and unconventional oil and gas formations including shales and tight sands (DOE-NETL carbon storage Atlas, 2015).</p>
<p>The present work focuses on storage in saline reservoirs, which provide significantly larger storage capacity, are globally more ubiquitous (<xref ref-type="bibr" rid="B14">Ji and Zhu, 2015</xref>) and have few competing uses than hydrocarbon reservoirs. Although depleted oil and gas reservoirs may provide important intermediate-scale storage, any CCS-activity, at a scale sufficient to impact the carbon problem (e.g., billions of metric tons), will necessarily involve large-scale CO<sub>2</sub> injections into deep saline aquifers (e.g., multiple projects inject one million metric tons per year or greater). However, CCS projects generally target &#x201c;green field sites,&#x201d; where there is often little characterization data and therefore large uncertainty about the petrophysical properties and other geologic attributes of the storage reservoir (<xref ref-type="bibr" rid="B4">Brandt et&#x20;al., 2014</xref>; <xref ref-type="bibr" rid="B6">Celia et&#x20;al., 2015</xref>). Uncertainty associated with predicting subsurface response to CO<sub>2</sub> injection is a key challenge to project developers seeking to secure financing, permits, and social license to inject CO<sub>2</sub> into the storage reservoir (<xref ref-type="bibr" rid="B17">Namhata et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B8">Chen et&#x20;al., 2020</xref>). Due to the inherent uncertainty about the storage reservoir, ensemble-based approaches are often used to forecast multiple realizations prior to CO<sub>2</sub> injection to visualize a range of potential outcomes. In addition, monitoring data during injection operations are used to update the pre-injection forecasts and thereby improve agreement between forecasted and observed behavior. Today, forecasting the subsurface response to CO<sub>2</sub> injection requires detailed three-dimensional (3D) geologic models coupled with numerical reservoir simulation, which are labor- and time-intensive and require specialists with backgrounds in petrophysics, geology, and reservoir engineering. Providing CO<sub>2</sub> storage site operators and regulators with rapid forecasting tools for timely decision making is essential to addressing these challenges to CCS project development and management. Delivering on this need requires transformational changes in how we predict subsurface responses to CO<sub>2</sub> injection and update those predictions using monitoring measurements.</p>
<p>Different methods have been employed to forecast geological carbon storage scenarios including analytical solutions and numerical simulations. Analytical methods are useful in providing quick evaluations with minimum input data and they are free from numerical artifacts (<xref ref-type="bibr" rid="B7">Celia et&#x20;al., 2005</xref>; <xref ref-type="bibr" rid="B12">Guo et&#x20;al., 2014</xref>; <xref ref-type="bibr" rid="B21">Qiao et&#x20;al., 2021</xref>). Numerical simulations, on the other hand, have been widely used in large-scale projects (<xref ref-type="bibr" rid="B20">Pawar et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B13">Humez et&#x20;al., 2011</xref>). However, numerical approaches (e.g., compositional reservoir simulation) usually require significant computational time and detailed geological data and measurements that may not always be available. In this study, we use machine learning techniques to address some of the challenges of numerical methods.</p>
<p>The conventional numerical method for predicting CO<sub>2</sub> distribution in a reservoir relies heavily on inverse modeling (history matching, calibration) to constrain uncertain parameters in complex reservoir simulation models (<xref ref-type="bibr" rid="B2">Bianco et&#x20;al., 2007</xref>; <xref ref-type="bibr" rid="B18">Oliver and Chen, 2011</xref>; <xref ref-type="bibr" rid="B10">Doughty and Oldenburg, 2020</xref>). This inversion-based prediction approach has limitations for rapid integration of observation data and providing timely decision support due to the following reasons: 1) Model inversion is computationally expensive and can require tens of thousands of expensive reservoir model simulations. Not all these simulations can be conducted concurrently and thus cannot take full advantage of contemporary parallel computing resources (each forward simulation may be parallelizable and a set of forward simulations may be conducted concurrently, but most inverse methods are essentially iterative and cannot achieve full parallelism). 2) Model inversion can be numerically ill-posed resulting in poor predictions when the number of parameters is greater than the number of independent observations, which is usually the case in geological carbon storage simulation. 3) Model inversion needs to be repeated when incorporating new observations. 4) Reservoir simulation models are based on geologic models, which may artificially constrain simulations and are slow and expensive to update with new field (as opposed to operational)&#x20;data.</p>
<p>To address these challenges, our research aims to develop machine learning (ML) techniques with a potential to provide significant improvements to the conventional history matching-based forecasts, thus enhancing the timeliness and accuracy of information provided to the operator. This paper describes our methods and analyzes their performance in predicting the CO<sub>2</sub> plume and pressure distribution in the storage reservoir at a commercial-scale storage project. Our project is part of a large initiative called SMART (Science-informed Machine Learning for Accelerating Real Time Decisions in Subsurface Applications) funded by U.S. Department of Energy with the goal to enable better decisions in CO<sub>2</sub> storage operations.</p>
</sec>
<sec id="s2">
<title>2 Materials and Methods</title>
<p>We propose a Learning-based Inversion-free Prediction (LIP) framework that produces fast prediction with uncertainty quantification <italic>via</italic> integrating observations, based on parallel forward simulations. The observations can be streaming measurements that are obtained from point locations continuously or near-continuously in time such as pressure and CO<sub>2</sub> saturation data from a well, and they can also be a saturation distribution data from a time-lapse 3D seismic survey (4D seismic survey). In this study, we consider the former, the point data discrete in locations but continuous in time. The key idea of the LIP framework is to circumvent the challenge of inverse modeling by precomputing an ensemble of unconstrained forward simulations and then using ML methods to learn the relationship between simulated observation and prediction variables. Once the ML model has learned the relationship, it can be used to update the prediction of future system behavior from its prior distribution to the posterior distribution by integrating actual observed data. When additional observations are available, we retrain the ML model to update the observation-prediction relationship by extracting the corresponding simulated observation and prediction variable samples from the prior sample set. Because the ML model training is very fast (a few seconds by using LIP) and the incorporation of new observations does not require extra reservoir simulations (by extracting the simulated samples from the prior sample set), the LIP method enables rapid data assimilation and timely decision support. The new observations can be the transient data from the same location/well or can be the data from different locations or even different types of data. As long as these observation variables have been simulated in the forward model runs, there is no need in the LIP framework to perform additional forward simulations when incorporating the new observations.</p>
<p>The key of LIP is to establish an observation-prediction relationship from prior samples in a reduced dimension to be able to estimate posterior prediction distributions for given observations. Specifically, LIP consists of four steps:<list list-type="simple">
<list-item>
<p>1 Generating prior samples of observation and prediction variables by running forward models based on the prior distribution of model parameters;</p>
</list-item>
<list-item>
<p>2 Dimension reduction of the simulated observations and predictions;</p>
</list-item>
<list-item>
<p>3 Establishing a statistical relationship between observation and prediction in the reduced dimension;</p>
</list-item>
<list-item>
<p>4 Using Bayesian inference to calculate the posterior distribution of the prediction based on the statistical model and by integrating the observed&#x20;data.</p>
</list-item>
</list>
</p>
<p>Steps 1-3 correspond to the training stage, where the observation-prediction relationship in the reduced dimension is learned from unconstrained forward simulations. Step 4 corresponds to the prediction stage, where the posterior distribution of the prediction is deduced from the observed data after back transformation to its original high-dimensional space. The LIP method can be generally applied to geological carbon storage problems. In this study, a clastic shelf model was considered as the geological model for a demonstration because the clastic shelf environment exhibits the greatest CO<sub>2</sub> storage rate in the model comparison study of <xref ref-type="bibr" rid="B3">Bosshart et&#x20;al. (2018)</xref>.</p>
<sec id="s2-1">
<title>2.1 Model Description and Generation of Prior Samples</title>
<p>To meet the goal of producing results relevant to commercial-scale CCS operations and meanwhile being able to perform the model simulations in a reasonable time, a 3D model domain was designed at a resolution of 211 by 211 with 30 layers, i.e.,&#x20;1,335,630 grid cells in total, where each cell has a size of 500 feet long, 500 feet wide, and 10 feet thick. The model has a flat structure and the storage formation is 4,000 feet deep. The model contains three facies: high-quality reservoir, low-quality reservoir, and cap rock. The top two layers of the model were assigned cap rock (shale) facies and they were given shale porosity and permeability values based upon previous work by <xref ref-type="bibr" rid="B5">Cavanagh and Wildgust (2011)</xref>. In this study, we considered the uncertainty of porosity and permeability in the reservoir and generated their realizations in the following&#x20;way.</p>
<p>All the realizations have the same rock facies geometry, but the porosity-permeability distributions differ. We sampled the porosity-permeability parameter space to generate the realizations. The Energy &#x26; Environmental Research Center at the University of North Dakota maintains an Average Global database (AGD) of paired porosity and permeability measurements for a host of lithologies, facies, and depositional environments (currently 26,700 &#x2b; measurements) (<xref ref-type="bibr" rid="B11">Gorecki et&#x20;al., 2009</xref>). The current work used a clastic shelf depositional environment and porosity-permeability paired samples specific to that environment. We first generated porosity realizations using Gaussian random function simulation with a variogram of 5,000 feet in the major and minor directions and 20 feet in the vertical direction. Permeability realizations were generated from the porosity-permeability cross-plots based on the derived relationship between these two variables from the AGD. We created 100 geological realizations (i.e.,&#x20;geomodels) using Schlumberger&#x2019;s Petrel software suite that reflect the variation in porosity and permeability. The ensemble can be envisioned as a stratified sample, where the number of realizations is proportional to the probability distribution for porosity. For example, we sampled seven percentiles of p05, p10, p25, p50, p75, p90, and p95 to represent the low to high porosity/permeability and for each percentile we have the following number of realizations, 10, 15, 22, 23, 16, 9, and 5, respectively. <xref ref-type="fig" rid="F1">Figure&#x20;1</xref> shows the porosity and permeability fields of layer three for one realization from the p25, p50, and p75 geomodels. We can see that these geomodels have a large variation in porosity and permeability.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>One realization of porosity (top) and permeability (bottom) fields of layer three for geomodels p25, p50, and p75. These three models have the same rock type geometry; the models p25, p50, and p75 correspond to low, mid, high porosity, respectively. The porosity color scale is a fraction from 0 to 0.4 (0&#x2013;40% porosity) and the permeability color scale is the logarithm (base 10) of the permeability in millidarcys</p>
</caption>
<graphic xlink:href="fenrg-09-752185-g001.tif"/>
</fig>
<p>For each geomodel, we performed a full equation-of-state compositional simulation (physics including convective and dispersive flow, residual gas trapping, CO<sub>2</sub> dissolution in aqueous phase, thermal capability) for 10&#xa0;years using CMG-GEM (v2019), which is a reservoir simulator for compositional, chemical and unconventional reservoir modelling. The model was simulated using closed lateral and vertical boundaries. The temperature and pressure regimes for the simulation at 4000-foot depth were 120.778&#x2009;4&#xb0;F and 1,800.93&#xa0;psi, respectively. The temperature and pressure were defined for each layer. They followed a linear pressure gradient of 0.43&#xa0;psi/ft and a linear temperature gradient of 0.019&#x2009;67&#xb0;F/ft. Four injection wells&#x2014;located regularly at the grid cells (71, 71, 3&#x2013;30) (71, 141, 3&#x2013;30), (141, 141, 3&#x2013;30), and (141, 71, 3&#x2013;30), respectively&#x2014;inject CO<sub>2</sub> into the reservoir with a target mass injection rate of two million metric tons per year across all four wells. We simulated 10&#xa0;years of injection with a 20 million metric tons CO<sub>2</sub> injection target (2 million metric tons per year &#xd7; 10&#xa0;years), to represent the earliest years of a commercial-scale storage project, during which CO<sub>2</sub> plumes are expected be the least predictable (i.e.,&#x20;the greatest rates of change per unit time). Maximum bottom hole pressure (BHP) constraint for each injector was Pf &#x3d; 0.7&#xa0;psi/ft. If BHP &#x2264; Pf, then CO<sub>2</sub> will continue to flow into the formation. However, if BHP <inline-formula id="inf1">
<mml:math id="m1">
<mml:mo>&#x3e;</mml:mo>
</mml:math>
</inline-formula> Pf, then the injection rate for that well will slow down. We saved the simulation outputs of CO<sub>2</sub> plume and pressure distribution of the entire 3D domain at 32&#x20;time steps, i.e.,&#x20;monthly for years 1 and 2 and then annually for years 3&#x2013;10. Each simulation takes about 7&#x2013;10&#xa0;h on average using 4&#x2013;10 cores on an Intel Xeon Scalable (Cascade Lakes) CPU. The lengthy simulation time makes it really difficult, if not impossible, to enable conventional inversion-based history matching for timely forecasting.</p>
<p>In this study, we use the LIP method to predict the CO<sub>2</sub> plume and pressure distribution in layer 3 (the top model layer of the storage reservoir immediately below the cap rock layer) after 10&#xa0;years of injection based on the saturation and pressure observations from the four injection wells in layer 3. For example, the prediction variables for pressure is the CO<sub>2</sub> pressure distribution at each of the 211 by 211 grid cells in layer 3 (44&#x2009;,521 variables in total) at year 10, and the observation variables are the four time-series of pressure-transients at the four wells in layer 3. We performed five case studies, depending on the duration of the observation period and thus the look-ahead period. We summarized the five cases in <xref ref-type="table" rid="T1">Table&#x20;1</xref>. Specifically, we forecasted pressure distribution in year 10 from the perspective of year 1, 2, 5, 7 and 9. In each case, we used only the data (both the simulated and observed data of observation variables) available up to that time, which corresponds to varying the look-ahead period from 9&#xa0;years (i.e.,&#x20;the perspective of year 1 looking ahead to year 10) to 1&#xa0;year (i.e.,&#x20;the perspective of year 9 looking ahead to year 10). For example, in Case I, we used 1&#x20;year of pressure-transient observations (12&#x20;time steps&#xd7;4 wells &#x3d; 48 observation variables) in layer three to predict the pressure distribution at each of the 211 by 211 grid cells in layer three at year 10; and in Case V, we used 9&#xa0;years of observations (31&#x20;time steps&#xd7;4 wells &#x3d; 124 observation variables) to predict the pressure distribution in year 10. These five case studies were designed to evaluate LIP&#x2019;s accuracy, efficiency, and capacity to incorporate streaming observations to improve prediction.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Definition of the five case studies and the corresponding LIP method&#x2019;s prediction performance. In the five cases, the prediction variables are the same which are the target we want to predict and the observation variables are different depending on the duration of the observation period. We investigate LIP&#x2019;s predictive capability (measured by mean absolute error (MAE)) in incorporating different number of observation&#x20;data.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th colspan="6" align="left">Prediction variable: CO<sub>2</sub> pressure distribution at each grid cell in layer three at year 10</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="2" align="left">Observation variable: CO<sub>2</sub> pressure observations from the 4 injection wells in layer 3 with different duration of observation period</td>
<td align="center">Case I</td>
<td align="center">Case II</td>
<td align="center">Case III</td>
<td align="center">Case IV</td>
<td align="center">Case V</td>
</tr>
<tr>
<td align="center">1&#xa0;year of observations</td>
<td align="center">2&#xa0;years of observations</td>
<td align="center">5&#xa0;years of observations</td>
<td align="center">7&#xa0;years of observations</td>
<td align="center">9&#xa0;years of observations</td>
</tr>
<tr>
<td align="left">MAE of LIP predicted pressure (synthetic &#x201c;truth&#x201d; p50)</td>
<td align="center">14.85&#xa0;psi</td>
<td align="center">11.59&#xa0;psi</td>
<td align="center">10.77&#xa0;psi</td>
<td align="center">8.05&#xa0;psi</td>
<td align="center">7.77&#xa0;psi</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>These are challenging applications because of the large uncertain domain and the limited 100 geomodels and simulation samples. To evaluate LIP performance, we took one geomodel as a synthetic &#x201c;truth&#x201d; and used the other 99 geomodels to learn the observation-prediction relationship in above Steps 2&#x2013;3. The corresponding pressure plume of the synthetic geomodel served as the reference against which we assessed our prediction results. To investigate the robustness of the LIP method for predicting the CO<sub>2</sub> plume and pressure field with different patterns, we made three choices of the synthetic geomodels corresponding to low, mid, and high porosity, i.e.,&#x20;picking one realization from the p25, p50, and p75 geomodels, respectively. For each synthetic case, we used the selected geomodel as reference and the other 99 geomodels for learning. In the similar manner to predict the pressure, we used the CO<sub>2</sub> saturation data in the four wells to predict the CO<sub>2</sub> plume in layer 3 after 10&#xa0;years of injection. <xref ref-type="fig" rid="F2">Figure&#x20;2</xref> shows the 100 samples of the CO<sub>2</sub> pressure and saturation profile in the four wells over 10&#xa0;years at the 32&#x20;time steps where we highlighted the three samples chosen as the synthetic observed data in the three synthetic cases (each dot in the highlighted line represents one time step). The figure indicates that the difference of the pressure and saturation among the samples is fairly large and our selected synthetic &#x201c;truth&#x201d; has a good representation of the low, mid, and high pressure/saturation. The small number of training data (99 geomodels) and the limited and non-smooth observation-transient data (monthly in first 2&#xa0;years and annually in last 8&#xa0;years) make the prediction problem rather challenging. In the following subsections of <xref ref-type="sec" rid="s2">Section 2</xref>, we explain the key Steps 2-4 of the LIP method to solve this problem. In <xref ref-type="sec" rid="s3">Section 3</xref>, we demonstrate how this problem was addressed using&#x20;LIP.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>The 100 simulation outputs of CO<sub>2</sub> pressure and saturation profile in the four wells over 10&#xa0;years at 32&#x20;time steps (identified by dots in the colored lines) where the first 24&#x20;time steps are monthly data and the last eight time steps are annual data. The highlighted three colored lines, which correspond to one realization of geomodels p25, p50, and p75 shown in <xref ref-type="fig" rid="F1">Figure&#x20;1</xref>, were used as &#x201c;synthetic&#x201d; observed&#x20;data.</p>
</caption>
<graphic xlink:href="fenrg-09-752185-g002.tif"/>
</fig>
</sec>
<sec id="s2-2">
<title>2.2 Dimension Reduction</title>
<p>The prediction variable (denote as <bold>h</bold>) here is a spatial distribution and the observation variables (denote as <bold>d</bold>) are four time series, which have spatial and temporal correlations, respectively. When the variable dimensions are highly correlated with each other, multicollinearity occurs (<xref ref-type="bibr" rid="B9">Daoud, 2017</xref>). Multicollinearity results in numerical issues during model fitting and degrades predictive performance of the statistical model. One solution for addressing multicollinearity is dimension reduction. Dimension reduction identifies degrees of freedom that capture most of the variance in the data. Therefore, performing statistical analysis in the reduced dimension removes the multicollinearity and facilitates the model fitting. Additionally, dimension reduction reduces the number of variables and thus reduces the required number of samples, which improves the computational efficiency and enhances the model reliability.</p>
<p>We use principal component analysis (PCA) for dimension reduction. PCA is a multivariate analysis technique that applies an orthogonal transformation to convert a set of samples of possibly correlated variables into a set of values of uncorrelated variables, called principal components. Typically, the first a few components of the PCA decomposition explain most of the variance of the data. PCA is commonly used for dimensionality reduction by projecting each data point onto only the first few principal components to obtain lower-dimensional data while preserving as much of the data&#x2019;s variation as possible. The first principal component can equivalently be defined as a direction that maximizes the variance of the projected data. The <italic>i</italic>
<sup>
<italic>th</italic>
</sup> principal component can be taken as a direction orthogonal to the first <italic>i</italic>&#x20;&#x2212; 1 principal component that maximizes the variance of the projected&#x20;data.</p>
<p>Since our observation variables are from multiple sources (i.e.,&#x20;four injection wells), we use a mixed PCA to pool data together and generate a reduced dimensional projection of the combined data. First, a standard PCA is performed on each of the data source (i.e.,&#x20;the pressure transient from each injection well) to obtain the largest singular values. Next, each data source is normalized according to its first singular value; this accounts for any difference in scales amongst the data sources. Last, the normalized data inputs are concatenated and the standard PCA is applied to this final matrix. After dimension reduction, we obtain observation variables <bold>d</bold>
<sup>
<italic>f</italic>
</sup> and prediction variables <bold>h</bold>
<sup>
<italic>f</italic>
</sup>, respectively, in the reduced dimension. PCA is a bijective operation, so the original high-dimensional variable can be recovered uniquely by undoing the projection.</p>
</sec>
<sec id="s2-3">
<title>2.3 Establishing the Statistical Relationship</title>
<p>The relationship between <bold>d</bold>
<sup>
<italic>f</italic>
</sup> and <bold>h</bold>
<sup>
<italic>f</italic>
</sup> in the reduced dimension can be nonlinear which challenges the statistical model learning. We first use canonical correlation analysis (CCA) (<xref ref-type="bibr" rid="B22">Yang et&#x20;al., 2021</xref>) to linearize the relationship and simplify the model fitting. CCA is a multivariate analysis method that can be applied to transform the relationships between pairs of vector variables into a set of independent linearized relationships between pairs of scalar variables. The resulting linear combinations are denoted as <bold>d</bold>
<sup>
<italic>c</italic>
</sup> and <bold>h</bold>
<sup>
<italic>c</italic>
</sup>, and called the canonical variates of <bold>d</bold>
<sup>
<italic>f</italic>
</sup> and <bold>h</bold>
<sup>
<italic>f</italic>
</sup>. The canonical transformation is found through the eigen-decomposition of the sample covariance matrix and this CCA transformation is reversible. If <bold>d</bold>
<sup>
<italic>c</italic>
</sup> and <bold>h</bold>
<sup>
<italic>c</italic>
</sup> in the canonical space are nearly linearly correlated, a linear model can be used to simulate their relationship. If after CCA, the relationship of <bold>d</bold>
<sup>
<italic>c</italic>
</sup> and <bold>h</bold>
<sup>
<italic>c</italic>
</sup> is still not quite linear, we can use advanced ML models such as neural networks for regression.</p>
</sec>
<sec id="s2-4">
<title>2.4 Bayesian Inference of the Prediction</title>
<p>We use Bayesian inference to estimate predictions. But unlike the traditional workflow which uses Bayesian methods to quantify uncertainties of model parameters first and then infer prediction uncertainties (<xref ref-type="bibr" rid="B15">Lu et&#x20;al., 2017</xref>), we use Bayesian methods to calculate the posterior distribution of the predictions directly. Based on Bayes&#x2019; rule, the posterior distribution of a prediction variable <bold>h</bold> for some observed data <bold>d</bold>
<sub>
<italic>obs</italic>
</sub> is<disp-formula id="e1">
<mml:math id="m2">
<mml:mi>p</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x221d;</mml:mo>
<mml:mi>L</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mi>p</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(1)</label>
</disp-formula>where <italic>p</italic>(<bold>h</bold>) is the prior distribution and <italic>L</italic> (<bold>h</bold>&#x7c;<bold>d</bold>
<sub>
<italic>obs</italic>
</sub>) is the likelihood function. PCA and CCA enable reducing a set of high-dimensional variables (<bold>d</bold>, <bold>h</bold>) to a set of low-dimensional and linearly correlated variables (<bold>d</bold>
<sup>
<italic>c</italic>
</sup>, <bold>h</bold>
<sup>
<italic>c</italic>
</sup>). We first estimate the posterior distribution <inline-formula id="inf2">
<mml:math id="m3">
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> and then transform <bold>h</bold>
<sup>
<italic>c</italic>
</sup> back to its original space <bold>h</bold>. In the canonical space, <inline-formula id="inf3">
<mml:math id="m4">
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> can be estimated by<disp-formula id="e2">
<mml:math id="m5">
<mml:mi>p</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x221d;</mml:mo>
<mml:mi>L</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mi>p</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(2)</label>
</disp-formula>
</p>
<p>We use a linear model <italic>G</italic> to simulate the relationship between <bold>d</bold>
<sup>
<italic>c</italic>
</sup> and <bold>h</bold>
<sup>
<italic>c</italic>
</sup>, i.e.,&#x20;<bold>d</bold>
<sup>
<italic>c</italic>
</sup> &#x3d; <italic>G</italic>
<bold>h</bold>
<sup>
<italic>c</italic>
</sup>. By assuming a Gaussian likelihood as commonly done in the CCS community (<xref ref-type="bibr" rid="B18">Oliver and Chen, 2011</xref>), <inline-formula id="inf4">
<mml:math id="m6">
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> can be formulated as<disp-formula id="e3">
<mml:math id="m7">
<mml:mi>L</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(3)</label>
</disp-formula>where <inline-formula id="inf5">
<mml:math id="m8">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> is the covariance matrix of the observation error. In this work, we are considering a synthetic case where the observed data is from one geomodel, so <inline-formula id="inf6">
<mml:math id="m9">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> here is calculated as the covariance of the residuals from the linear model fitting.</p>
<p>Through normal score transformation based on the sample mean <inline-formula id="inf7">
<mml:math id="m10">
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and the sample covariance <inline-formula id="inf8">
<mml:math id="m11">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> calculated from the prior samples of <bold>h</bold>
<sup>
<italic>c</italic>
</sup>, we obtain a Gaussian prior of <bold>h</bold>
<sup>
<italic>c</italic>
</sup> in the transformed space. Since the prior and the likelihood of <bold>h</bold>
<sup>
<italic>c</italic>
</sup> are both Gaussian, its posterior is also Gaussian and the posterior mean <inline-formula id="inf9">
<mml:math id="m12">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula> and posterior covariance <inline-formula id="inf10">
<mml:math id="m13">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold">C</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> can be analytically estimated by<disp-formula id="e4">
<mml:math id="m14">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msub>
<mml:msup>
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msub>
<mml:msup>
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>G</mml:mi>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(4)</label>
</disp-formula>
<disp-formula id="e5">
<mml:math id="m15">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold">C</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mi>G</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>.</mml:mo>
</mml:math>
<label>(5)</label>
</disp-formula>
</p>
<p>An advantage of the Gaussian process regression is that a Gaussian distribution is uniquely defined by its mean and covariance and sampling a Gaussian distribution is straightforward. Then, based on <xref ref-type="disp-formula" rid="e4">Eqs 4</xref>, <xref ref-type="disp-formula" rid="e5">5</xref>, we generate posterior samples of <bold>h</bold>
<sup>
<italic>c</italic>
</sup> directly. By undoing the normal score transformation followed by the back transformation of CCA, we obtain posterior samples of <bold>h</bold>
<sup>
<italic>f</italic>
</sup>. Next, after back transformation of PCA, we obtain the posterior samples of prediction quantity <bold>h</bold> in its original space. Based on these <bold>h</bold> samples, we then estimate posterior prediction distribution.</p>
</sec>
</sec>
<sec id="s3">
<title>3 Results</title>
<p>In this section, we present the results of applying the LIP framework to the synthetic simulation cases to illustrate the LIP method and evaluate its prediction performance. We start with the most data-constrained case (Case V in <xref ref-type="table" rid="T1">Table&#x20;1</xref>) using 9&#xa0;years of pressure transient data from the wells to predict the pressure distribution in year 10. Then, we discuss the results of additional cases (Case I &#x2013; Case IV in <xref ref-type="table" rid="T1">Table&#x20;1</xref>) by incorporating 1, 2, 5, and 7&#xa0;years of observations to assess the sensitivity of the LIP prediction performance to the available monitoring data and to evaluate the capability of LIP to incorporate additional observations for improving the prediction. Lastly, we show the results of applying the LIP framework to the CO<sub>2</sub> plume prediction.</p>
<p>In the following, we discuss the results using 9&#xa0;years of pressure observations. We first use the synthetic case of p50 to demonstrate the LIP method, and then we analyze the prediction performance in detail for the three synthetic cases. In the LIP framework, after we generate the model simulation data from the geomodels, we perform the dimension reduction of the observation and prediction variables based on the simulation samples. <xref ref-type="fig" rid="F3">Figure&#x20;3</xref> shows the scree plots of the PCA (a line plot of cumulative variance versus number of principal components used to determine the number of principal components to keep in the PCA). We can see that the dimensions of both observation and prediction variables can be greatly reduced by keeping the first few components with a little information loss. Here, the observation variables are 9&#xa0;years of pressure data from the four wells (i.e.,&#x20;31&#x20;&#xd7; 4&#x20;&#x3d; 124 variables), and the prediction variables are the pressure distribution in each grid cell of layer three at year 10 (i.e.,&#x20;211&#x20;&#xd7; 211&#x20;&#x3d; 44&#x2009;,521 variables). <xref ref-type="fig" rid="F3">Figures 3A,B</xref> indicate that the first ten principal components capture 99% of variation in the observation variables <bold>d</bold> and that the first ten principal components capture 98% of variation in the prediction variables <bold>h</bold>. Based on these results from the dimension reduction step, for both variables we keep their first ten principal components and then establish the statistical relationship of <bold>d</bold> and <bold>h</bold> in their reduced ten dimensional space. <xref ref-type="fig" rid="F3">Figures 3C, D</xref>, using one realization as a demonstration, indicate that keeping the first ten principal components we are able to recover the target pressure field with minor difference from its original pressure distribution where the mean absolute error is about 1.6&#xa0;psi.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Scree plots of <bold>(A)</bold> observation variables <bold>d</bold> and <bold>(B)</bold> prediction variables <bold>h</bold> in PCA. <bold>(A)</bold> indicates that 10 principal components (PCs) can capture over 99% variation of the observation variable <bold>d</bold>; <bold>(B)</bold> shows that 10&#xa0;PCs can capture about 98% variation of the prediction variable <bold>h</bold>. <bold>(C)</bold> CO<sub>2</sub> pressure field from the original data and <bold>(D)</bold> the recovered pressure field from inverse PCA using the preserved 10&#xa0;PCs. The x- and <italic>y</italic>-axes for the pressure fields are the model x- and y-coordinates and the color map is pressure in psi from 1800 to 2000&#xa0;psi.</p>
</caption>
<graphic xlink:href="fenrg-09-752185-g003.tif"/>
</fig>
<p>Next, in the reduced observation-prediction dimensions, we perform the CCA for linear transformation. The scatter plot of <xref ref-type="fig" rid="F4">Figure&#x20;4</xref> indicates that after applying CCA, the canonical variates <bold>d</bold>
<sup>
<italic>c</italic>
</sup> and <bold>h</bold>
<sup>
<italic>c</italic>
</sup> have a strong linear correlation with coefficients of 0.99 and 0.92 for the first two principal components, respectively. The coefficients for the remaining eight principal components are also high, above 0.8 (results are not shown here). This suggests that a linear regression model can be established to simulate the relationship of <bold>d</bold>
<sup>
<italic>c</italic>
</sup> and <bold>h</bold>
<sup>
<italic>c</italic>
</sup>. In this study, both observation and prediction variables are the same type of quantity (i.e.,&#x20;CO<sub>2</sub> pressure) with smooth variation, so it is not very surprising that they show strong linear correlation&#x20;here.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Scatter plots of the canonical variates for the observation variables (<bold>d</bold>
<sup>
<italic>c</italic>
</sup>) and the prediction variables (<bold>h</bold>
<sup>
<italic>c</italic>
</sup>) for the first principal component (left) and the second principal component (right), together with the observation data (<bold>d</bold>
<sub>
<italic>obs</italic>
</sub>) in the canonical space after CCA.</p>
</caption>
<graphic xlink:href="fenrg-09-752185-g004.tif"/>
</fig>
<p>Lastly, we use Bayesian inference to calculate the mean and variance of the Gaussian posterior distribution of the prediction variables in the transformed space, <inline-formula id="inf11">
<mml:math id="m16">
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold">h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold">d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>b</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, according to <xref ref-type="disp-formula" rid="e4">Eqs 4</xref>, <xref ref-type="disp-formula" rid="e5">5</xref>. With the calculated mean and variance, we draw posterior samples from this Gaussian distribution, and then do a series of back transformations to transform those posterior samples in the space <bold>h</bold>
<sup>
<italic>c</italic>
</sup> back into their original space <bold>h</bold>. We start by undoing the normal score transformation, then the canonical back transformation, and lastly the PCA back transformation into the original&#x20;space.</p>
<p>The final prediction results of <bold>h</bold> are summarized in <xref ref-type="fig" rid="F5">Figure&#x20;5</xref>. Although in this p50-case the prior mean is already similar to the synthetic &#x201c;truth&#x201d; in capturing the pressure field patterns (due to the way we generated the porosity and permeability realizations where the p50-geomodel corresponds to 50% percentile of the porosity probability distribution), the LIP method, by incorporating the observations from the four wells, still greatly improves the prediction accuracy. The resulted posterior mean pressure field is more like the synthetic &#x201c;truth&#x201d; compared to the prior mean, with a coefficient of determination (<italic>R</italic>
<sup>2</sup>) of 0.99, and the mean absolute error (MAE) of the posterior mean is 7.77 psi which is about one fourth of the MAE of the prior mean of 26.58&#xa0;psi. Especially in the region of pressure buildup around the four wells, the posterior estimation accurately captures the high buildup pressure in the two wells on the right hand side and it also delineates the pressure movement and front more precisely compared to the prior estimate, which results in a uniformly small posterior error in the entire domain.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Evaluation of LIP-predicted CO<sub>2</sub> pressure after 10-years of injection based on 9&#xa0;years of pressure observations in four wells. Top, left-right: synthetic &#x201c;truth&#x201d; of CO<sub>2</sub> pressure distribution (psi) in year 10 for model p50; cross-plot of the synthetic &#x201c;truth&#x201d; and LIP-predicted pressure distribution; mean pressure distribution (psi) from the prior samples; and LIP-estimated posterior mean after incorporating 9&#xa0;years of observations; Bottom: absolute prediction error and the standard deviation (std) from the prior samples and LIP-generated posterior samples.</p>
</caption>
<graphic xlink:href="fenrg-09-752185-g005.tif"/>
</fig>
<p>Following the similar steps, we applied the LIP method to the other two synthetic cases. <xref ref-type="fig" rid="F6">Figure&#x20;6</xref> shows the results for the p25 case. The figure indicates that the prior mean pressure map is dramatically deviated from the synthetic &#x201c;truth&#x201d; in this case. Because of the low porosity and permeability of the p25 geomodel, the pressure is relatively large, up to 2,800&#xa0;psi around the injection wells. The prior mean significantly underestimates the pressure with a MAE of 108&#xa0;psi and the prior estimate does not capture the pressure movement. On the other hand, the posterior mean produced by LIP not only accurately delineates the pressure front, but also identifies that the two wells at the bottom have larger pressure buildup, resulting in a smaller MAE of 42.27&#xa0;psi. Compared to the prior, the posterior mean pressure field is much more like the synthetic &#x201c;truth&#x201d; with a <italic>R</italic>
<sup>2</sup> of 0.96 and the posterior error field is also much smaller. <xref ref-type="fig" rid="F7">Figure&#x20;7</xref> summarizes the results of the p75 case. Due to the high porosity and permeability of the p75 geomodel, the pressure is relatively small in this case, below 2,100&#xa0;psi. As the prior mean is an average of the other 99 geomodels, it significantly overestimates the pressure with a MAE of 75.4&#xa0;psi. The LIP method, after effectively incorporating the observations from the four wells, dramatically reduces the MAE to 4.76&#xa0;psi which is only 6.3% of the prior MAE. Furthermore, the posterior mean pressure field is very similar to the synthetic &#x201c;truth&#x201d; with a <italic>R</italic>
<sup>2</sup> of 0.98 resulting in uniformly small posterior errors in the entire domain.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Evaluation of LIP-predicted CO<sub>2</sub> pressure after 10-years of injection based on 9&#xa0;years of pressure observations in four wells. Top, left-right: synthetic &#x201c;truth&#x201d; of CO<sub>2</sub> pressure distribution (psi) in year 10 for model p25; cross-plot of the synthetic &#x201c;truth&#x201d; and LIP-predicted pressure distribution; mean pressure distribution (psi) from the prior samples; and LIP-estimated posterior mean after incorporating 9&#xa0;years of observations; Bottom: absolute prediction error and the standard deviation (std) from the prior samples and LIP-generated posterior samples.</p>
</caption>
<graphic xlink:href="fenrg-09-752185-g006.tif"/>
</fig>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Evaluation of LIP-predicted CO<sub>2</sub> pressure after 10-years of injection based on 9&#xa0;years of pressure observations in four wells. Top, left-right: synthetic &#x201c;truth&#x201d; of CO<sub>2</sub> pressure distribution (psi) in year 10 for model p75; cross-plot of the synthetic &#x201c;truth&#x201d; and LIP-predicted pressure distribution; mean pressure distribution (psi) from the prior samples; and LIP-estimated posterior mean after incorporating 9&#xa0;years of observations; Bottom: absolute prediction error and the standard deviation (std) from the prior samples and LIP-generated posterior samples.</p>
</caption>
<graphic xlink:href="fenrg-09-752185-g007.tif"/>
</fig>
<p>Although these three synthetic cases show dramatically different pressure distributions and patterns, e.g., within the region of pressure buildup caused by CO<sub>2</sub> injection, the difference between the cases of p25 and p75 is approximately 100&#x2013;800&#xa0;psi, all the cases indicate that the LIP method greatly improves estimation accuracy compared to the prior mean. Additionally, LIP significantly reduces the predictive uncertainty by producing a smaller posterior standard deviation field than the prior standard deviation field, which gives not only an accurate but also a confident forecasting. As shown in the last plots of <xref ref-type="fig" rid="F5">Figures 5</xref>, <xref ref-type="fig" rid="F6">6</xref>, <xref ref-type="fig" rid="F7">7</xref>, the posterior standard deviation of the pressure field is close to zero in the entire domain. The resulted accurate and credible prediction of the CO<sub>2</sub> pressure distribution in the reservoir is critical for risk assessment and to inform decisions made by site operators.</p>
<p>To evaluate the LIP&#x2019;s ability to incorporate additional observations for prediction improvement and to investigate the sensitivity of prediction performance to the number of observations, we designed a series of numerical experiments (Case I &#x2013; Case V in <xref ref-type="table" rid="T1">Table&#x20;1</xref>) where we incorporate differing numbers of years of pressure data from the wells. The results of incorporating 1, 2, 5, and 7&#xa0;years of observed data to predict pressure distribution in year 10 are presented in <xref ref-type="fig" rid="F8">Figure&#x20;8</xref>. As shown in the figure, incorporating more years of observations produces a posterior mean pressure field that gets asymptotically closer to the synthetic &#x201c;truth&#x201d; in <xref ref-type="fig" rid="F5">Figure&#x20;5</xref>. The MAE, as summarized in <xref ref-type="table" rid="T1">Table&#x20;1</xref>, gradually decreases from 14.85&#xa0;psi (incorporating 1&#xa0;year of data), to 11.59&#xa0;psi (incorporating 2&#xa0;years of data), to 10.77&#xa0;psi (incorporating 5&#xa0;years of data), to 8.05&#xa0;psi (incorporating 7&#xa0;years of data), and finally to 7.77 when incorporating 9&#xa0;years of observations for forecasting. In incorporation of only 1&#xa0;year of data, the posterior mean is already able to capture the major patterns and movement of the pressure field; additional years of data gradually refine the detail of the predicted pressure map. This indicates that the LIP method can effectively extract the information from the limited simulation data for learning the observation-prediction relationship and incorporates the observed data for updating the prediction from the unconstraint prior estimation to more accurate posterior estimation.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>LIP estimated posterior mean of the CO<sub>2</sub> pressure distribution in year 1after considering 1, 2, 5, and 7&#xa0;years of pressure observations (obs) in four wells. The synthetic &#x201c;truth&#x201d; is in <xref ref-type="fig" rid="F5">Figure&#x20;5</xref>.</p>
</caption>
<graphic xlink:href="fenrg-09-752185-g008.tif"/>
</fig>
<p>Note that the incorporation of these additional observations in LIP does not require extra reservoir simulations. LIP incorporates new observation data by performing the analysis in Steps 2-4 of <xref ref-type="sec" rid="s2">Section 2</xref> based on the corresponding observation variable simulations from the prior sample set. The statistical analysis in Steps 2-4 is very fast which takes a few seconds in this study. The ability of LIP to rapidly generate new forecasts promises fast integration of streaming observations for timely forecasts in field operation. Furthermore, the additional data are not necessarily the transient data from the same well with a longer period of observations, they can also come from other wells and can be different types of measurements. As long as these additional observation variables have been simulated in the forward runs, there is no need to perform extra simulations when incorporating the new&#x20;data.</p>
<p>In addition to predicting the pressure distribution, we also applied the LIP method to predict the CO<sub>2</sub> plume in the storage reservoir. <xref ref-type="fig" rid="F9">Figure&#x20;9</xref> shows the prediction results of CO<sub>2</sub> distribution in year 10 after incorporating 9&#xa0;years of observations from the four wells for the three synthetic cases. The three cases show dramatically different CO<sub>2</sub> distributions, e.g., within the footprint of the CO<sub>2</sub> plume, the difference in gas saturation between the cases of p25 and p75 is approximately &#xb1;0.35. Despite the significantly diverse CO<sub>2</sub> distributions, the posterior mean produced by LIP can still capture their major patterns. The prior mean shows that the footprints of CO<sub>2</sub> around the four wells are similar to each other, however, the posterior mean from LIP depicts that the CO<sub>2</sub> plume is actually different around different wells and the resulting patterns are much closer to the synthetic &#x201c;truth&#x201d;. Additionally, the prior samples display a large standard deviation around the wells. After effectively incorporating the observations, the posterior standard deviation is greatly reduced, showing more confident prediction. Although LIP improved the prediction accuracy and credibility by producing a better posterior mean and a smaller standard deviation than the prior estimation, its prediction of CO<sub>2</sub> plume is relatively poor compared to the prediction of pressure distributions, where the posterior pressure field is more like the synthetic &#x201c;truth&#x201d;. One reason is that CO<sub>2</sub> field is less continuous than the pressure field, i.e.,&#x20;the pressure field extends outwards from each of the four wells and forms a continuous extent that covers most of the model domain, whereas the CO<sub>2</sub> plumes around each of the four wells are smaller in extent and more irregularly shaped. So after dimensional reduction, the CO<sub>2</sub> plume may lose more information for statistical learning. Moreover, the observation-prediction relationship of CO<sub>2</sub> gas saturation is more complicated than the relationship of the pressure, and the current study is limited to 99 training samples for learning the relationship which may not be enough. Additionally, we only have observations from the four wells at 31&#x20;time steps; these limited simulation data and point observed data are far less than enough to accurately delineate the CO<sub>2</sub> footprint in such a large and heterogeneous domain.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>Evaluation of LIP-predicted CO<sub>2</sub> plume after 10-years of injection based on 9&#xa0;years of saturation observations in four wells. Three rows are results for the three different synthetic models where model p25, p50, and p75 correspond to low, mid, and high porosity, respectively. Left-right: synthetic &#x201c;truth&#x201d; of CO<sub>2</sub> plume; mean pressure distribution (psi) from the prior samples; LIP-estimated posterior mean after incorporating 9&#xa0;years of observations; and the standard deviation (std) from the prior samples and LIP-generated posterior samples.</p>
</caption>
<graphic xlink:href="fenrg-09-752185-g009.tif"/>
</fig>
</sec>
<sec id="s4">
<title>4 Discussion</title>
<p>In this paper, we provide an efficient and effective prediction method (LIP)&#x2014;using a set of machine learning techniques&#x2014;to perform accurate, timely forecasts for geological carbon storage based on a limited number of measurements and a few model simulations. We use three different synthetic cases demonstrate that the LIP method can greatly improve CO<sub>2</sub> plume and pressure prediction accuracy and reduce predictive uncertainty by effectively incorporating observations. The proposed LIP method runs very fast; after obtaining prior samples, it takes a few seconds to perform the entire process&#x2014;from dimension reduction, to canonical correlation analysis, to Bayesian inference for prediction. The prior samples are independent and can be performed completely concurrently on parallel computing platforms; with enough processors available, the generation of the hundred prior samples would only require the same wallclock time as one forward reservoir simulation. LIP is also data efficient; based on 99 prior samples, it can effectively learn the observation-prediction relationship and accurately infer the posterior prediction distributions by incorporating the observed data. LIP uses estimated observation-prediction relationship to infer predictions. In this study, we used PCA followed by CCA to build a linear relationship in the reduced canonical space and then use the Gaussian linear regression for predictions. In situations when the relationship is nonlinear and multimodal, we can use Bayesian neural networks for regression. To avoid extrapolation, LIP requires the observation data to lie inside of the prior samples. We can adjust the prior distribution and increase the prior sample size to satisfy this requirement.</p>
<p>LIP has a considerable potential to fundamentally change how timely decisions are made about CO<sub>2</sub> storage operations. Bypassing the conventional workflow of history matching and then forward simulations, LIP provides fast updating forecasts of CO<sub>2</sub> plume and pressure distributions from streaming observations, thus providing operators with early warning of off-normal behavior and more time to implement mitigation measures. In our future work, we will apply LIP to real measurement data from the field, and deploy it to CO<sub>2</sub> storage operators for fast decision making.</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>The raw data supporting the conclusions of this article will be made available by the authors, without undue reservation.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>DL developed the algorithms, planned and implemented the numerical experiments, and led manuscript preparation. SP contributed to the research plan, interpretation of results, and manuscript preparation. NA, MB-K, TJ, and CW generated the geomodels, performed the reservoir simulations, and helped prepare the manuscript.</p>
</sec>
<sec id="s7">
<title>Funding</title>
<p>Primary funding support for this work is provided by the Science-informed Machine Learning to Accelerate Real Time decision making for Carbon Storage (SMART-CS) Initiative, funded by the US Department of Energy (DOE), Office of Fossil Energy. Additional support is provided by the Artificial Intelligence Initiative as part of the Laboratory Directed Research and Development Program of Oak Ridge National Laboratory, managed by UT-Battelle, LLC, for the US DOE under contract DE-AC05-00OR22725. This research is also sponsored by the Data-Driven Decision Control for Complex Systems (DnC2S) project funded by the US DOE, Office of Advanced Scientific Computing Research.</p>
</sec>
<sec id="s8">
<title>Author Disclaimer</title>
<p>The United&#x20;States Government retains and the publisher, by accepting the article for publication, acknowledges that the United&#x20;States Government retains a non-exclusive, paidup, irrevocable, world-wide license to publish, or reproduce the published form of this manuscript, or allow others to do so, for United&#x20;States Government purposes.</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of Interest</title>
<p>Authors DL and SP were employed by Oak Ridge National Laboratory.</p>
<p>The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>Part of the methodology of this manuscript has been presented at the 2019 International Conference on Data Mining Workshops, DOI: <ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://10.1109/ICDMW.2019.00049">10.1109/ICDMW.2019.00049</ext-link>. This manuscript has been co-authored by staff from UT-Battelle, LLC under Contract No. DE-AC05-00OR22725 with the U.S. Department of Energy.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Alcalde</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Flude</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Wilkinson</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Johnson</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Edlmann</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Bond</surname>
<given-names>C. E.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Estimating Geological CO2 Storage Security to Deliver on Climate Mitigation</article-title>. <source>Nat. Commun.</source> <volume>9</volume>, <fpage>2201</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-018-04423-1</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Bianco</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Cominelli</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Dovera</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Naevdal</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Valles</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2007</year>). &#x201c;<article-title>History Matching and Production Forecast Uncertainty by Means of the Ensemble Kalman Filter: A Real Field Application</article-title>,&#x201d; in <conf-name>EUROPEC/EAGE Conference and Exhibition</conf-name>, <conf-loc>London, United Kingdom</conf-loc>, <conf-date>June 2007</conf-date>. <comment>SPE-107161-MS</comment>. <pub-id pub-id-type="doi">10.2118/107161-MS</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bosshart</surname>
<given-names>N. W.</given-names>
</name>
<name>
<surname>Azzolina</surname>
<given-names>N. A.</given-names>
</name>
<name>
<surname>Ayash</surname>
<given-names>S. C.</given-names>
</name>
<name>
<surname>Peck</surname>
<given-names>W. D.</given-names>
</name>
<name>
<surname>Gorecki</surname>
<given-names>C. D.</given-names>
</name>
<name>
<surname>Ge</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Quantifying the Effects of Depositional Environment on Deep saline Formation Co2 Storage Efficiency and Rate</article-title>. <source>Int. J.&#x20;Greenhouse Gas Control.</source> <volume>69</volume>, <fpage>8</fpage>&#x2013;<lpage>19</lpage>. <pub-id pub-id-type="doi">10.1016/j.ijggc.2017.12.006</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Brandt</surname>
<given-names>A. R.</given-names>
</name>
<name>
<surname>Heath</surname>
<given-names>G. A.</given-names>
</name>
<name>
<surname>Kort</surname>
<given-names>E. A.</given-names>
</name>
<name>
<surname>O&#x27;Sullivan</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Petron</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Jordaan</surname>
<given-names>S. M.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Methane Leaks from north American Natural Gas Systems</article-title>. <source>Science</source> <volume>343</volume>, <fpage>733</fpage>&#x2013;<lpage>735</lpage>. <pub-id pub-id-type="doi">10.1126/science.1247045</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cavanagh</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Wildgust</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Pressurization and Brine Displacement Issues for Deep saline Formation Co2 Storage</article-title>. <source>Energ. Proced.</source> <volume>4</volume>, <fpage>4814</fpage>&#x2013;<lpage>4821</lpage>. <comment>10th International Conference on Greenhouse Gas Control Technologies</comment>. <pub-id pub-id-type="doi">10.1016/j.egypro.2011.02.447</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Celia</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Bachu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Nordbotten</surname>
<given-names>J.&#x20;M.</given-names>
</name>
<name>
<surname>Bandilla</surname>
<given-names>K. W.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Status of CO2storage in Deep saline Aquifers with Emphasis on Modeling Approaches and Practical Simulations</article-title>. <source>Water Resour. Res.</source> <volume>51</volume>, <fpage>6846</fpage>&#x2013;<lpage>6892</lpage>. <pub-id pub-id-type="doi">10.1002/2015wr017609</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Celia</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Bachu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Nordbotten</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Gasda</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Dahle</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2005</year>). &#x201c;<article-title>Quantitative Estimation of CO2 Leakage from Geological storageAnalytical Models, Numerical Models, and Data Needs</article-title>,&#x201d; in <source>Greenhouse Gas Control Technologies 7</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Rubin</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Keith</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Gilboy</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Wilson</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Morris</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Gale</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<publisher-loc>Oxford</publisher-loc>: <publisher-name>Elsevier Science Ltd</publisher-name>), <fpage>663</fpage>&#x2013;<lpage>671</lpage>. <pub-id pub-id-type="doi">10.1016/B978-008044704-9/50067-7</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Harp</surname>
<given-names>D. R.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Pawar</surname>
<given-names>R. J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Reducing Uncertainty in Geologic Co2 Sequestration Risk Assessment by Assimilating Monitoring Data</article-title>. <source>Int. J.&#x20;Greenhouse Gas Control.</source> <volume>94</volume>, <fpage>102926</fpage>. <pub-id pub-id-type="doi">10.1016/j.ijggc.2019.102926</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Daoud</surname>
<given-names>J.&#x20;I.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Multicollinearity and Regression Analysis</article-title>. <source>J.&#x20;Phys. Conf. Ser.</source> <volume>949</volume>, <fpage>012009</fpage>. <pub-id pub-id-type="doi">10.1088/1742-6596/949/1/012009</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Doughty</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Oldenburg</surname>
<given-names>C. M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Co2 Plume Evolution in a Depleted Natural Gas Reservoir: Modeling of Conformance Uncertainty Reduction over Time</article-title>. <source>Int. J.&#x20;Greenhouse Gas Control.</source> <volume>97</volume>, <fpage>103026</fpage>. <pub-id pub-id-type="doi">10.1016/j.ijggc.2020.103026</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gorecki</surname>
<given-names>C. D.</given-names>
</name>
<name>
<surname>Sorensen</surname>
<given-names>J.&#x20;A.</given-names>
</name>
<name>
<surname>Bremer</surname>
<given-names>J.&#x20;M.</given-names>
</name>
<name>
<surname>Knudsen</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Smith</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Steadman</surname>
<given-names>E. N.</given-names>
</name>
<etal/>
</person-group> (<year>2009</year>). &#x201c;<article-title>Development of Storage Coefficients for Determining the Effective Co2 Storage Resource in Deep saline Formations</article-title>,&#x201d; in <conf-name>SPE International Conference on CO2 Capture, Storage, and Utilization</conf-name>, <conf-loc>San Diego, CA</conf-loc>, <conf-date>November 2009</conf-date>. <pub-id pub-id-type="doi">10.2118/126444-MS</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guo</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Bandilla</surname>
<given-names>K. W.</given-names>
</name>
<name>
<surname>Doster</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Keilegavlen</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Celia</surname>
<given-names>M. A.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>A Vertically Integrated Model with Vertical Dynamics for CO2storage</article-title>. <source>Water Resour. Res.</source> <volume>50</volume>, <fpage>6269</fpage>&#x2013;<lpage>6284</lpage>. <pub-id pub-id-type="doi">10.1002/2013WR015215</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Humez</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Audigane</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Lions</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chiaberge</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Bellenfant</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Modeling of CO2 Leakage up through an Abandoned Well from Deep Saline Aquifer to Shallow Fresh Groundwaters</article-title>. <source>Transp Porous Med.</source> <volume>90</volume>, <fpage>153</fpage>&#x2013;<lpage>181</lpage>. <pub-id pub-id-type="doi">10.1007/s11242-011-9801-2</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ji</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>CO2 Storage in Deep Saline Aquifers</article-title>,&#x201d; in <source>Novel Materials for Carbon Dioxide Mitigation Technology</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Shi</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Morreale</surname>
<given-names>B.</given-names>
</name>
</person-group> (<publisher-loc>Amsterdam</publisher-loc>: <publisher-name>Elsevier</publisher-name>), <fpage>299</fpage>&#x2013;<lpage>332</lpage>. <pub-id pub-id-type="doi">10.1016/b978-0-444-63259-3.00010-0</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lu</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Ricciuto</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Walker</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Safta</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Munger</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Bayesian Calibration of Terrestrial Ecosystem Models: a Study of Advanced Markov Chain Monte Carlo Methods</article-title>. <source>Biogeosciences</source> <volume>14</volume>, <fpage>4295</fpage>&#x2013;<lpage>4314</lpage>. <pub-id pub-id-type="doi">10.5194/bg-14-4295-2017</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Metz</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Davidson</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>De Coninck</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Loos</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Meyer</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2005</year>). <source>IPCC Special Report on Carbon Dioxide Capture and Storage</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>. </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Namhata</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Oladyshkin</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Dilmore</surname>
<given-names>R. M.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Nakles</surname>
<given-names>D. V.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Probabilistic Assessment of above Zone Pressure Predictions at a Geologic Carbon Storage Site</article-title>. <source>Sci. Rep.</source> <volume>6</volume>, <fpage>1</fpage>&#x2013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1038/srep39536</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Oliver</surname>
<given-names>D. S.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Recent Progress on Reservoir History Matching: a Review</article-title>. <source>Comput. Geosci.</source> <volume>15</volume>, <fpage>185</fpage>&#x2013;<lpage>221</lpage>. <pub-id pub-id-type="doi">10.1007/s10596-010-9194-2</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pacala</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Socolow</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Stabilization Wedges: Solving the Climate Problem for the Next 50&#x20;Years with Current Technologies</article-title>. <source>Science</source> <volume>305</volume>, <fpage>968</fpage>&#x2013;<lpage>972</lpage>. <pub-id pub-id-type="doi">10.1126/science.1100103</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pawar</surname>
<given-names>R. J.</given-names>
</name>
<name>
<surname>Watson</surname>
<given-names>T. L.</given-names>
</name>
<name>
<surname>Gable</surname>
<given-names>C. W.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Numerical Simulation of Co2 Leakage through Abandoned wells: Model for an Abandoned Site with Observed Gas Migration in alberta, canada</article-title>. <source>Energ. Proced.</source> <volume>1</volume>, <fpage>3625</fpage>&#x2013;<lpage>3632</lpage>. <pub-id pub-id-type="doi">10.1016/j.egypro.2009.02.158</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Qiao</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Hoteit</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Fahs</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Semi-analytical Solution to Assess Co2 Leakage in the Subsurface through Abandoned wells</article-title>. <source>Energies</source> <volume>14</volume>, <fpage>2452</fpage>. <pub-id pub-id-type="doi">10.3390/en14092452</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Tao</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A Survey on Canonical Correlation Analysis</article-title>. <source>IEEE Trans. Knowl. Data Eng.</source> <volume>33</volume>, <fpage>2349</fpage>&#x2013;<lpage>2368</lpage>. <pub-id pub-id-type="doi">10.1109/TKDE.2019.2958342</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>