<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Earth Sci.</journal-id>
<journal-title>Frontiers in Earth Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Earth Sci.</abbrev-journal-title>
<issn pub-type="epub">2296-6463</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1084414</article-id>
<article-id pub-id-type="doi">10.3389/feart.2022.1084414</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Earth Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Rapid construction of Rayleigh wave dispersion curve based on deep learning</article-title>
<alt-title alt-title-type="left-running-head">Cui et al.</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/feart.2022.1084414">10.3389/feart.2022.1084414</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Cui</surname>
<given-names>Diyu</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/2026531/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Shi</surname>
<given-names>Lijing</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Gao</surname>
<given-names>Kai</given-names>
</name>
</contrib>
</contrib-group>
<aff>
<institution>Key Laboratory of Earthquake Engineering and Engineering Vibration</institution>, <institution>Institute of Engineering Mechanics</institution>, <institution>China Earthquake Administration</institution>, <addr-line>Harbin</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1563540/overview">Shaohuan Zu</ext-link>, Chengdu University of Technology, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1172301/overview">Qiang Guo</ext-link>, China University of Mining and Technology, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2125058/overview">Lingqian Wang</ext-link>, College of Science, China University of Petroleum, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Lijing Shi, <email>shlj@iem.ac.cn</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Solid Earth Geophysics, a section of the journal Frontiers in Earth Science</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>16</day>
<month>01</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>10</volume>
<elocation-id>1084414</elocation-id>
<history>
<date date-type="received">
<day>30</day>
<month>10</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>28</day>
<month>12</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Cui, Shi and Gao.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Cui, Shi and Gao</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>
<bold>Introduction:</bold> The dispersion curve of the Rayleigh-wave phase velocity (VR) is widely utilized to determine site shear-wave velocity (Vs) structures from a distance of a few metres to hundreds of metres, even on a ten-kilometre crustal scale. However, the traditional theoretical-analytical methods for calculating VRs of a wide frequency range are time-consuming because numerous extensive matrix multiplications, transfer matrix iterations and the root searching of the secular dispersion equation are involved. It is very difficult to model site structures with many layers and apply them to a population-based inversion algorithm for which many populations of multilayers forward modelling and many generations of iterations are essential.</p>
<p>
<bold>Method:</bold> In this study, we propose a deep learning method for constructing the VR dispersion curve in a horizontally layered site with great efficiency. A deep neural network (DNN) based on the fully connected dense neural network is designed and trained to directly learn the relationships between Vs structures and dispersion curves. First, the training and validation sets are generated randomly according to a truncated Gaussian distribution, in which the mean and variance of the Vs models are statistically analysed from different regions&#x2019; empirical relationships between soil Vs and its depth. To be the supervised dataset, the corresponding VRs are calculated by the generalized reflection-transmission (R/T) coefficient method. Then, the Bayesian optimization (BO) is designed and trained to seek the optimal architecture of the deep neural network, such as the number of neurons and hidden layers and their combinations. Once the network is trained, the dispersion curve of VR can be constructed instantaneously without building and solving the secular equation.</p>
<p>
<bold>Results and Discussion:</bold> The results show that the DNN-BO achieves a coefficient of determination (R<sup>2</sup>) and MAE for the training and validation sets of 0.98 and 8.30 and 0.97 and 8.94, respectively, which suggests that the rapid method has satisfactory generalizability and stability. The DNN-BO method accelerates the dispersion curve calculation by at least 400 times, and there is almost no increase in computation expense with an increase in soil layers.</p>
</abstract>
<kwd-group>
<kwd>phase velocity</kwd>
<kwd>Rayleigh wave</kwd>
<kwd>site structure</kwd>
<kwd>deep neural network</kwd>
<kwd>Bayesian optimization</kwd>
<kwd>deep learning</kwd>
</kwd-group>
<contract-num rid="cn001">51978635</contract-num>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Rayleigh waves are formed by the interaction of incident P and SV plane waves at the free surface and travel parallel to that surface. The phase velocity (VR) dispersion curve of the Rayleigh wave depends on the medium parameters, such as the layer thickness and the P and S velocities. Longer wavelength Rayleigh waves penetrate deeper than shorter wavelengths over layered geologies. This dispersive behavior makes Rayleigh waves a valuable tool for determining the shear-wave velocity (Vs) of individual site structures from a distance of a few meters to hundreds of meters, even at the 10-km crustal scale of a region. Determining dispersive behavior is one of the key steps in calculating the theoretical phase velocity dispersion curve of Rayleigh waves in inverted geological structures, modeling regional Rayleigh waves and synthesizing seismograms.</p>
<p>To construct the dispersion curve, it is necessary to build a secular equation based on elastic wave theory and surface boundary determinations and compute its roots as a function of its frequency. The most famous method is the Thomson&#x2013;Haskell algorithm (also called the transfer matrix method) (<xref ref-type="bibr" rid="B48">Thomson, 1950</xref>; <xref ref-type="bibr" rid="B19">Haskell, 1953</xref>). In the Thomson&#x2013;Haskell algorithm, the dispersion equation is constructed by a sequence of matrix multiplications involving terms that are transcendental functions of the material properties of the layered medium. Many researchers have modified this method throughout the years (<xref ref-type="bibr" rid="B45">Schwab and Knopoff, 1970</xref>; <xref ref-type="bibr" rid="B44">Schwab and Knopoff, 1972</xref>; <xref ref-type="bibr" rid="B1">Abo-Zena, 1979</xref>; <xref ref-type="bibr" rid="B14">Fan et al., 2002</xref>) to improve its numerical stability and efficiency by reducing the dimension of transfer matrices and simplifying the iterative multiplication.</p>
<p>Another important class of algorithms for solving Rayleigh wave eigenvalue problems is the reflection and transmission coefficients (R/T) method (<xref ref-type="bibr" rid="B27">Kennett, 1974</xref>; <xref ref-type="bibr" rid="B26">Kennett and Kerry, 1979</xref>; <xref ref-type="bibr" rid="B25">Kennett and Clarke, 1983</xref>). The Rayleigh dispersion equation for a stratified medium is iteratively established by reflection and transmission matrices using R/T coefficients. Subsequently, this method was also modified and improved by other researchers (<xref ref-type="bibr" rid="B33">Luco and Apsel, 1983</xref>; <xref ref-type="bibr" rid="B10">Chen, 1993</xref>; <xref ref-type="bibr" rid="B21">Hisada, 1994</xref>; <xref ref-type="bibr" rid="B22">1995</xref>; <xref ref-type="bibr" rid="B20">He and Chen, 2006</xref>; <xref ref-type="bibr" rid="B38">Pan et al., 2022</xref>).</p>
<p>Numerical techniques can also be used to solve the Rayleigh eigenvalue problem, including the finite difference method (<xref ref-type="bibr" rid="B8">Boore, 1972</xref>), numerical integration (<xref ref-type="bibr" rid="B47">Takeuchi and Saito, 1972</xref>), the boundary element method (<xref ref-type="bibr" rid="B36">Manolis and Beskos, 1988</xref>; <xref ref-type="bibr" rid="B6">Aung and Leong, 2010</xref>), and the spectral element method (<xref ref-type="bibr" rid="B13">Faccioli et al., 1996</xref>; <xref ref-type="bibr" rid="B29">Komatitsch and Vilotte, 1998</xref>). Although these methods have several advantages, they require more computation time.</p>
<p>In conventional methods, most efforts to calculate the dispersion curve are primarily focused on building the dispersion equation and searching for its phase velocity solution as a function of frequency. The computational efficiency is improved by reducing the dimension of the transfer matrix and/or simplifying the iteration of the dispersion equation. However, tedious matrix iteration and root searching cannot be avoided.</p>
<p>In inversion analysis, the velocity structure is estimated by minimizing the deviation between the theoretical and experimental dispersion curves. For population-based inversion algorithms, many populations of multilayer forward modeling and many generations of iterations are essential to search the optimal site parameters (shear wave velocity, overburden thickness, etc.). It is difficult to model site structures that have many layers and apply population-based inversion algorithms (<xref ref-type="bibr" rid="B39">Picozzi and Albarello, 2007</xref>; <xref ref-type="bibr" rid="B32">Lu et al., 2016</xref>; <xref ref-type="bibr" rid="B41">Poormirzaee, 2016</xref>; <xref ref-type="bibr" rid="B30">Lei et al., 2019</xref>; <xref ref-type="bibr" rid="B40">Poormirzaee and Fister, 2021</xref>). Therefore, it is necessary to develop a fast, accurate, stable dispersion curve construction method.</p>
<p>The increasing number of populations and the requirement of refining the site layers induce tremendous iteration calculation costs in the population-based inversion algorithm. Under these situations, conventional forward methods struggle to satisfy the big data processes. Deep learning (DL) techniques are one of the most potent ways to establish a good mapping relationship between the seismic signals and geophysical parameters. Once a DL model is trained, it can map the site velocity structure to the Rayleigh dispersion curve directly and quickly under similar geological conditions. The DL method has been widely applied in various geotechnical earthquake engineering applications to replace the complex conventional numerical calculation method. <xref ref-type="bibr" rid="B24">Jo et al. (2022)</xref> used a trained deep neural network (DNN) to replace the conventional history matching method, greatly reducing the computational cost of predicting well responses, such as oil production rate, water production rate, and reservoir flow direction. <xref ref-type="bibr" rid="B50">Wamriew et al. (2022)</xref> used a trained convolutional neural network to improve the efficiency and accuracy of microseismic event monitoring. Their model avoids the shortcoming that the conventional methods, which are affected by manual intervention and require substantial data preprocessing, rely on. <xref ref-type="bibr" rid="B49">Tschannen et al. (2022)</xref> accelerated the compilation speed in wavelet extraction through deep learning, which restrains the iterative adjustment parameters and potential noise. Similarly, some researchers in surface wave exploration have applied deep learning to extract dispersion curves (<xref ref-type="bibr" rid="B4">Alyousuf et al., 2018</xref>; <xref ref-type="bibr" rid="B53">Zhang et al., 2020</xref>; <xref ref-type="bibr" rid="B11">Dai et al., 2021</xref>) and invert velocity structures (<xref ref-type="bibr" rid="B2">Aleardi and Stucchi, 2021</xref>; <xref ref-type="bibr" rid="B17">Fu et al., 2021</xref>; <xref ref-type="bibr" rid="B34">Luo et al., 2022</xref>). <xref ref-type="bibr" rid="B4">Alyousuf et al. (2018)</xref> used a fully connected network to extract the fundamental-mode dispersion curve. <xref ref-type="bibr" rid="B53">Zhang et al. (2020)</xref> presented a convolution network to automatically extract dispersion curves. <xref ref-type="bibr" rid="B11">Dai et al. (2021)</xref> developed a DCNet deep learning model that can correctly extract the multimode dispersion curve from the segmentation dispersion image. <xref ref-type="bibr" rid="B17">Fu et al. (2021)</xref> applied DispINets to inverse the shear velocity from the multimode dispersion curve. <xref ref-type="bibr" rid="B2">Aleardi and Stucchi (2021)</xref> proposed a Monte Carlo-based hybrid neural network to resolve the degradation shortcomings in the inversions. <xref ref-type="bibr" rid="B35">Luo et al. (2022)</xref> used a DNN to investigate the range of the initial model and the reliability of the shear velocity structure. Through these successful applications, deep learning techniques have demonstrated their potential in constructing dispersion curves from site data with high efficiency.</p>
<p>Deep learning-based surface wave exploration research may be defined as a systematic process that consists of two essential components. One is establishing a sufficient and reasonably distributed dataset that includes all the geophysical states (<xref ref-type="bibr" rid="B3">Alwosheel et al., 2018</xref>). Another is designing a network that has high accuracy and efficiency. However, no one dataset ensures that the site structures of different regions are covered. Researchers (<xref ref-type="bibr" rid="B9">Chauhan and Dhingra, 2012</xref>; <xref ref-type="bibr" rid="B5">Assi et al., 2018</xref>; <xref ref-type="bibr" rid="B23">Itano et al., 2018</xref>) rely on the mathematical understanding of the algorithm, empirical judgment, and trial and error to determine the network architectures, which incurs high computational costs. Researchers have proposed a data-driven method to limit the initial calculation space and utilized a population-based algorithm to enhance the computational efficiency of nonlinear parameter combinations (<xref ref-type="bibr" rid="B18">Guo et al., 2021</xref>; <xref ref-type="bibr" rid="B34">Luo et al., 2022</xref>). Therefore, we use the Vs&#x2013;h empirical relationship of nine regions in China to generate the dataset and use the Bayesian optimization algorithm to search a high-precision, low-parameter network.</p>
<p>This paper proposes the rapid constructing dispersion curve method based on DNN with Bayesian optimization (DNN-BO) to improve forward modeling efficiency. The fully connected dense neural network is set as the DNN main architecture. Bayesian optimization is applied to optimize the DNN architecture by iteratively testing the potential combination. We design a method to generate the random layer site that ensures the diversity and abundance of the dataset. Finally, we discuss and analyze the accuracy and efficiency of the rapid method.</p>
</sec>
<sec id="s2">
<title>2 Dataset generation</title>
<p>The number and distribution of the dataset samples are closely related to the accuracy and generalization of the DL methods that determine the applicability of the method. The site velocity structure is a two-dimensional spatial structure composed of the shear wave velocity and the burial depth. The most straightforward method is to randomly and uniformly generate site samples for all types of sites in this space for the shear wave velocity at each depth. However, the uniform sampling method produces many invalid samples and has a high cost in the precomputing of the dispersion curves and the training of the model. Furthermore, the site velocity structure generated in this way conforms to a uniform distribution that dilutes the main site characteristics and does not reflect the natural sedimentation law. Therefore, it is crucial to construct a representative site structure dataset.</p>
<sec id="s2-1">
<title>2.1 Site characteristics</title>
<p>Sediments on the surface gradually accumulate in the progress of geophysical evolution. More sediments will settle in the upper layer, which causes the lower layer to have a greater consolidation and shear wave velocity. Consequently, in the most general site type, the shear wave velocity increases with depth. Geological tectonic movements, such as earthquakes, volcanoes, and debris flows, have affected sedimentary processes and physical states. Some soft sediments are forced to a deeper layer, and hard sediments are forced up to shallower layers, resulting in a small portion of the sites containing soft or hard interlayers. Sedimentary conditions, topographies, and other geological factors also influence the formation of site sediments, which leads to each region having specific regional characteristics.</p>
</sec>
<sec id="s2-2">
<title>2.2 Vs&#x2013;h empirical relationship</title>
<p>The uniform sampling method dilutes the main site characteristics and directly uses the drilling data in a regional area as the dataset, affecting the DL method&#x2019;s applicability. As shown in <xref ref-type="fig" rid="F1">Figure 1</xref>, we statistically analyzed the Vs&#x2013;h empirical relationship in Fujian, Harbin, and Kunming (<xref ref-type="sec" rid="s12">Supplementary Attached List S1</xref>). Based on these results, a comprehensive empirical relationship between shear wave velocity and burial depth is fitted, and the statistical relationship boundary points are used to determine the upper and lower boundaries of the dataset. In view of the site characteristics and regionality, the truncated normal distribution approach is adopted to generate the random layers dataset that utilizes the comprehensive Vs&#x2013;h relationship as the mean.<disp-formula id="e1">
<mml:math id="m2">
<mml:mrow>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>132.64</mml:mn>
<mml:mo>&#x2a;</mml:mo>
<mml:msup>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mn>0.18</mml:mn>
</mml:msup>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>
<disp-formula id="e2">
<mml:math id="m3">
<mml:mrow>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>97.23</mml:mn>
<mml:mo>&#x2a;</mml:mo>
<mml:msup>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mn>0.13</mml:mn>
</mml:msup>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>
<disp-formula id="e3">
<mml:math id="m4">
<mml:mrow>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>410.22</mml:mn>
<mml:mo>&#x2a;</mml:mo>
<mml:msup>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mn>0.15</mml:mn>
</mml:msup>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>
</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Fitting the empirical relationship between Vs and its burial depth in different regions.</p>
</caption>
<graphic xlink:href="feart-10-1084414-g001.tif"/>
</fig>
</sec>
<sec id="s2-3">
<title>2.3 Generating dataset</title>
<p>For a single site velocity structure, the Vs under each burial depth is randomly generated from a truncated normal distribution with <inline-formula id="inf2">
<mml:math id="m5">
<mml:mrow>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> as the mean, and <inline-formula id="inf3">
<mml:math id="m6">
<mml:mrow>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf4">
<mml:math id="m7">
<mml:mrow>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are used as the upper and lower bounds. The layer thickness is a random value within [1, 30] m. When Vs is more than 800&#xa0;m/s, the number of site layers will not continue to increase. Compared with the layer thickness and Vs, the compression wave velocity and density have less influence on the dispersion curve (<xref ref-type="bibr" rid="B12">Donoho, 1995</xref>). The compression wave velocity of each layer is set as twice Vs, and the density is set at 1.85&#xa0;<inline-formula id="inf5">
<mml:math id="m8">
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mo>/</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>m</mml:mi>
</mml:mrow>
<mml:mn>3</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>. This study generates 10,000 random layer site structures as input and uses the R/T method to calculate the dispersion curve as output. The dataset is divided into a training set and a validation set at a ratio of 4:1. The validation set does not participate in the training of the DL model.</p>
<p>To verify the velocity distribution at different depths, the velocity space at a depth of every meter is divided into 100 cells at 8&#xa0;m/s intervals. The probability density of each Vs cell is calculated by the percentage of sites in the total dataset. <xref ref-type="fig" rid="F2">Figure 2</xref> shows the distribution probability of shear velocity occurring at a depth of every meter in 10,000 site samples. Overall, Vs increases with burial depth, and there are two light-colored areas on the left and right. This indicates that the incremental sites make up the majority of the dataset, and a small number of sites have soft or hard interlayers. The estimated area is covered, indicating that the site velocity structure generated with a truncated normal distribution is representative and encompasses all the regions&#x2019; sites.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Velocity probability density map.</p>
</caption>
<graphic xlink:href="feart-10-1084414-g002.tif"/>
</fig>
<p>
<xref ref-type="fig" rid="F3">Figure 3</xref> shows the violin plot of the VR dispersion curve corresponding to the random layer site dataset. In the violin plot, the different colors of the violin represent the velocity distribution at different frequencies, and the width of the violin shows the data distribution at different velocities. The wider the violin at the phase velocity is, the more sites with that velocity are in the dataset. The distribution of the violin plot at each frequency is roughly a single peak, and the data distribution is relatively concentrated. The violin shapes at the 1.6&#xa0;Hz, 2.8&#xa0;Hz, and 4&#xa0;Hz frequencies are roughly similar to those at other frequencies. However, more discrete values indicate that the VR value changes significantly in this range, and the dispersion curve is prone to local fluctuation, which requires attention for the frequency range. The local fluctuation in the dispersion curve reflects a soft or hard interlayer. As a result, the frequency distribution of the dispersion curve is reselected. Within 0&#x2013;1&#xa0;Hz, the frequency step is set at .2&#xa0;Hz; within 1&#x2013;20&#xa0;Hz, the frequency step is set at .0077 in logarithmic coordinates.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Phase velocity dispersion point violin plot.</p>
</caption>
<graphic xlink:href="feart-10-1084414-g003.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="methods" id="s3">
<title>3 Methods</title>
<p>Constructing a dispersion curve from the site velocity structure is a regression task in deep learning. The commonly used DL models include DNN, CNN, and LSTM. The fully connected dense network is chosen as the main body and is usually called DNN (<xref ref-type="bibr" rid="B7">Bishop, 1995</xref>; <xref ref-type="bibr" rid="B37">Nabian and Meidani, 2018</xref>). It has a simple architecture and fast learning features. As shown in <xref ref-type="fig" rid="F4">Figure 4</xref>, this paper proposes a network architecture called DNN-BO. The DNN builds the mapping relationship between the site velocity structure and the dispersion curve. Bayesian optimization (BO) is used to automatically select DNN hyperparameters, such as the number of hidden layers, the number of neurons, and the activation function. A complete DNN architecture includes an input layer, some hidden layers, and an output layer. Before inputting the data, the random layer sites sample through the equivalent shear velocity calculation method to make their features correspond one-to-one with the input neurons. The dataset is normalized and input into the network. In the hidden layer, the site velocity structure features are extracted through the neurons connected in pairs. BO updates the DNN architecture based on the cross-validation method objective function minimization. The search strategy outputs the optimal network architecture until the stop condition or iteration count is reached. Finally, the stability, generalization, and efficiency of the rapid method are verified by the validation set.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Model building flow chart.</p>
</caption>
<graphic xlink:href="feart-10-1084414-g004.tif"/>
</fig>
<sec id="s3-1">
<title>3.1 Data processing</title>
<p>In the dataset, the number of site layers and layer thicknesses vary over a wide range. It is better to set a constant node size for the DL model. Therefore, the site layers are converted by the equivalent shear wave velocity method to meet the size of the DNN nodes. As a consequence, the equivalent shear wave velocity method is adopted (Eqs <xref ref-type="disp-formula" rid="e4">4</xref>&#x2013;<xref ref-type="disp-formula" rid="e6">6</xref>) to make Vs and the layer thickness correspond to the number of input nodes. In this way, the total number of site layers is standardized at 20. When the dataset contains more than 20 site layers, we prioritize searching for two adjacent layers with a minimum Vs difference. Then, the equivalent shear velocity method is used to calculate the Vs of the original two adjacent layers, and the sum of the two layer thicknesses is used to replace the original thickness. The aforementioned steps are repeated until the dataset contains 20 site layers. When the dataset has fewer than 20 site layers, we search for the thickest layer and divide it into two layers. After the division, the Vs of each layer remains unchanged, and the layer thickness is half the original. The aforementioned process is repeated until the dataset contains 20 site layers.<disp-formula id="e4">
<mml:math id="m9">
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2211;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>v</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>
<disp-formula id="e5">
<mml:math id="m10">
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2211;</mml:mo>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>
<disp-formula id="e6">
<mml:math id="m11">
<mml:mrow>
<mml:msub>
<mml:mi>v</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>/</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>where <inline-formula id="inf6">
<mml:math id="m12">
<mml:mrow>
<mml:msub>
<mml:mi>v</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the shear wave velocity of the <italic>i</italic>th layer, <inline-formula id="inf7">
<mml:math id="m13">
<mml:mrow>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the thickness of the <italic>i</italic>th layer, <inline-formula id="inf8">
<mml:math id="m14">
<mml:mrow>
<mml:msub>
<mml:mi>v</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the equivalent shear wave velocity, d represents the calculation depth, and t is the sum of the transit times of the shear wave velocity in each layer.</p>
<p>Vs is nearly 100 times larger than the layer thickness. If the original site data are directly input into the neural network, this underestimates the weight of the layer thickness and may lead to premature neural network saturation. Therefore, the max&#x2013;min normalization method is adopted to normalize the dataset into [0, 1].<disp-formula id="e7">
<mml:math id="m15">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi mathvariant="italic">min</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi mathvariant="italic">max</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi mathvariant="italic">min</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>where <inline-formula id="inf9">
<mml:math id="m16">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the eigenvalue of the column, <inline-formula id="inf10">
<mml:math id="m17">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi mathvariant="italic">min</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf11">
<mml:math id="m18">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi mathvariant="italic">max</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represent the maximum and minimum values of the eigenvalue of the column, respectively, and <inline-formula id="inf12">
<mml:math id="m19">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the normalized eigenvalue.</p>
<p>The well-trained DL model output dispersion curve value is still in the normalized scope and needs to be returned to the original scope. The denormalization formula is shown in Eq. <xref ref-type="disp-formula" rid="e8">8</xref>:<disp-formula id="e8">
<mml:math id="m20">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x223c;</mml:mo>
</mml:mover>
<mml:mo>&#x3d;</mml:mo>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x2d9;</mml:mo>
</mml:mover>
<mml:mo>&#x2a;</mml:mo>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi mathvariant="italic">max</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi mathvariant="italic">min</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi mathvariant="italic">min</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(8)</label>
</disp-formula>where <inline-formula id="inf13">
<mml:math id="m21">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x2d9;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> represents the DL model output value, <inline-formula id="inf14">
<mml:math id="m22">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi mathvariant="italic">max</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf15">
<mml:math id="m23">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi mathvariant="italic">min</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represent the maximum and minimum values of the original data, respectively, and <inline-formula id="inf16">
<mml:math id="m24">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x223c;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> represents the denormalized value.</p>
</sec>
<sec id="s3-2">
<title>3.2 Tuning the DNN hyperparameters using Bayesian optimization algorithms</title>
<p>The DL model itself contains many parameters that need to be determined manually. These parameters are collectively referred to as hyperparameters. The hyperparameters determine the complexity and computational accuracy of the model. There are many types and wide ranges of hyperparameters in the DL models. It is challenging to manually find the most suitable hyperparameter combination for the dataset. At present, hyperparameter optimization mainly depends on the mathematical understanding of the algorithm, empirical judgment, and trial and error (<xref ref-type="bibr" rid="B9">Chauhan and Dhingra, 2012</xref>; <xref ref-type="bibr" rid="B5">Assi et al., 2018</xref>; <xref ref-type="bibr" rid="B23">Itano et al., 2018</xref>; <xref ref-type="bibr" rid="B15">Feurer et al., 2019</xref>). However, these empirical methods have a high chance of missing the optimal hyperparameter combination to match the dataset and building a bloated network structure that requires too many computing resources. To find a lightweight and high-precision DL model to construct the dispersion curve, we set DNN as the main body and use the Bayesian optimization algorithm to automatically adjust the hyperparameters, including the number of neurons, number of hidden layers, activation function, learning rate, and training batch size. BO is suitable for multidimensional, high-cost valuation problems and has been widely used in maximum likelihood hyperparameter optimization (<xref ref-type="bibr" rid="B15">Feurer et al., 2019</xref>). In the geotechnical engineering and seismic engineering fields, it has been successfully applied for tunnel-boring machine performance prediction (<xref ref-type="bibr" rid="B54">Zhou et al., 2021</xref>), slope sensitivity analysis (<xref ref-type="bibr" rid="B43">Sameen et al., 2020</xref>) and non-destructive testing (<xref ref-type="bibr" rid="B31">Liang, 2019</xref>).</p>
<p>The BO algorithm is based on the evolutionary algorithm combined with the Bayesian network probability model. It contains two primary functions: the surrogate model and the acquisition function (<xref ref-type="bibr" rid="B16">Frazier, 2018</xref>). The surrogate model function fits the currently observed points to the objective function and obtains the similarity distribution between the surrogate model and the actual model. The acquisition function is used to select the next evaluated point with the least similarity in the distribution. Both the acquisition function and the surrogate function accelerate the convergence process (<xref ref-type="bibr" rid="B46">Snoek et al., 2015</xref>). The main idea of BO is to update the surrogate model by continuously adding new sample points through the acquisition function without knowing the internal structure and mathematical properties of the optimization objective function. It is not until the surrogate model is the same as the actual model or the iteration count is exhausted that the BO outputs the optimal surrogate model. This study uses a Gaussian process (GP) surrogate model that uses the mean and variance to determine the next estimated point. GP has an infinite-dimensional normal distribution extension on an infinite-dimensional random process and can consider multiple hyperparameter optimizations simultaneously (<xref ref-type="bibr" rid="B16">Frazier, 2018</xref>).</p>
<p>
<xref ref-type="fig" rid="F5">Figure 5</xref> shows the BO automatic tuning DNN steps. First, the BO initial parameters are set, which includes setting the mean of the k training loss functions as the objective function, setting the GP as the surrogate model, and setting its mean and variance as the acquisition functions. Then, BO randomly selects five DNN hyperparameter combinations to initialize the surrogate model. In each iteration, the next hyperparameter combination evaluation point is determined by the minimum acquisition function similarity. The surrogate model is updated on the objective function value. When the iteration counts are exhausted, the optimization strategy outputs the optimal DNN architecture.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Flow chart of Bayesian optimization tuning of the DNN hyperparameters.</p>
</caption>
<graphic xlink:href="feart-10-1084414-g005.tif"/>
</fig>
<p>Fivefold cross-validation is adopted to avoid overfitting and underfitting situations that significantly influence the stability of the deep learning-based dispersion curve construction (<xref ref-type="bibr" rid="B52">Wong, 2015</xref>). <xref ref-type="fig" rid="F6">Figure 6</xref> shows that the original training set is randomly divided into five equal-sized parts. One part is selected as the test set, and the other four are the training set. The objective function is the mean of the five training loss functions. The cross-validation method makes full use of every sample in the dataset. Each evaluated model will be trained five times in the BO optimization process. In this way, the generalization ability of the rapid method is reinforced.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Five-fold cross-validation.</p>
</caption>
<graphic xlink:href="feart-10-1084414-g006.tif"/>
</fig>
<p>The optimization space is closely related to the training cost and the calculational precision. A broad optimization space signifies a more comprehensive hyperparameter combination with higher accuracy and computational costs. A narrow optimization space may miss the optimal DNN architectures. It is expensive to set up an optimization space that includes all the hyperparameter combinations. Referring to deep learning-based surface wave exploration research (<xref ref-type="bibr" rid="B4">Alyousuf et al., 2018</xref>; <xref ref-type="bibr" rid="B53">Zhang et al., 2020</xref>; <xref ref-type="bibr" rid="B2">Aleardi and Stucchi, 2021</xref>; <xref ref-type="bibr" rid="B11">Dai et al., 2021</xref>; <xref ref-type="bibr" rid="B17">Fu et al., 2021</xref>; <xref ref-type="bibr" rid="B35">Luo et al., 2022</xref>), this study sets up four hidden layer blocks. The layers of each hidden layer block are between (0, 2), and the neurons are between (16, 256). The activation function is chosen among the commonly used nonlinear activation functions (&#x201c;relu,&#x201d; &#x201c;elu,&#x201d; &#x201c;selu,&#x201d; &#x201c;gelu,&#x201d; &#x201c;tanh,&#x201d; and &#x201c;swish&#x201d;).</p>
<p>The learning rate is one of the essential hyperparameters in deep learning and is used in training to update the weights and biases of the neurons and to control the loss function change rate (<xref ref-type="bibr" rid="B51">Werbos, 1988</xref>). When the learning rate is too high, the loss function changes rapidly as the training progresses, and the convergence speed is fast; however, it may jump back and forth near the optimal solution. In contrast, if the learning rate is too low, the loss function changes very slowly, and the convergence time is too long. The training result may fall into a locally optimal solution. We set the initial learning rate to (.00001, .01) and adopt an adaptive learning rate algorithm. If the loss function remains unchanged every three epochs, the learning rate will become half of the previous one.</p>
<p>The training epochs and batch size determine whether the DNN has fully learned the relationship between the site structure and the dispersion curve (<xref ref-type="bibr" rid="B42">Prechelt et al., 2012</xref>; <xref ref-type="bibr" rid="B3">Alwosheel et al., 2018</xref>). Insufficient training epochs and a large batch size may lead to an underfitting situation. Therefore, the epochs are set to 100, and the batch size is set to (4, 128).</p>
</sec>
</sec>
<sec sec-type="results" id="s4">
<title>4 Result</title>
<p>In this section, the DNN-BO model is trained with the random layer site dataset and outputs the optimal DNN architecture. The accuracy and efficiency of the rapid method are verified through the R/T method&#x2019;s calculation result. The analysis is conducted through a Lenovo R7000P 2020 laptop (R7-4800H CPU, mobile 2060 6GB GPU, and 16GB 3200 Hz DDR4 memory).</p>
<sec id="s4-1">
<title>4.1 The optimal network architecture</title>
<p>The DNN neurons in the input layer depend on the input data size, as does the output layer. After data processing, there are 39 neurons in the input layer, the VR values are used as the output, and there are 50 neurons in the output layer. This study selects the mean square error (MSE) as the loss function and Adam (<xref ref-type="bibr" rid="B28">Kingma and Ba, 2014</xref>) as the optimizer. <xref ref-type="fig" rid="F7">Figure 7A</xref> shows that after 100 iterations, BO can quickly find the DNN architecture most suitable for constructing the dispersion curve. As seen from <xref ref-type="table" rid="T1">Table 1</xref>, the optimal DNN is a three-hidden-layers simple architecture, in which the hidden layer neurons are 93, 128, and 243, respectively, and &#x201c;gelu&#x201d; is set as the activation function. <xref ref-type="fig" rid="F7">Figure 7B</xref> shows that the optimized DNN architecture is not underfitted or overfitted during training and achieves a good convergence state on both the training and validation sets.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Iterative convergence of the objective function and loss function.</p>
</caption>
<graphic xlink:href="feart-10-1084414-g008.tif"/>
</fig>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>DNN hyperparameters search space and the optimal parameters.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Parameter</th>
<th align="center">Optimization range</th>
<th align="center">Optimal parameter</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">Number of 1_hidden layers</td>
<td align="center">(0, 2)</td>
<td align="center">1</td>
</tr>
<tr>
<td align="center">Neurons in 1_hidden layers</td>
<td align="center">(16,256)</td>
<td align="center">93</td>
</tr>
<tr>
<td align="center">Number of 2_hidden layers</td>
<td align="center">(0, 2)</td>
<td align="center">1</td>
</tr>
<tr>
<td align="center">Neurons in 2_hidden layers</td>
<td align="center">(16,256)</td>
<td align="center">128</td>
</tr>
<tr>
<td align="center">Number of 3_hidden layers</td>
<td align="center">(0, 2)</td>
<td align="center">0</td>
</tr>
<tr>
<td align="center">Neurons in 3_hidden layers</td>
<td align="center">(16,256)</td>
<td align="center">0</td>
</tr>
<tr>
<td align="center">Number of 4_hidden layers</td>
<td align="center">(0, 2)</td>
<td align="center">1</td>
</tr>
<tr>
<td align="center">Neurons in 4_hidden layers</td>
<td align="center">(16,256)</td>
<td align="center">243</td>
</tr>
<tr>
<td align="center">Activation</td>
<td align="center">(&#x201c;relu,&#x201d; &#x201c;elu,&#x201d; &#x201c;selu,&#x201d; &#x201c;gelu,&#x201d; &#x201c;tanh,&#x201d; &#x201c;swish&#x201d;)</td>
<td align="center">gelu</td>
</tr>
<tr>
<td align="center">Batch size</td>
<td align="center">(4, 128)</td>
<td align="center">8</td>
</tr>
<tr>
<td align="center">Learning rate</td>
<td align="center">(.00001, .01)</td>
<td align="center">.00063</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4-2">
<title>4.2 DNN model evaluation</title>
<p>This study sets the mean square error (MSE), mean absolute error (MAE), mean absolute percent error (MAPE), and coefficient of determination (<inline-formula id="inf17">
<mml:math id="m25">
<mml:mrow>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>) as the overall accuracy evaluation criteria and adopts the ratio of the predicted VR value with the true VR value at each frequency as the local evaluation criterion. The smaller the MSE, MAE, and MAPE values are, the better the calculation effect of the model. <inline-formula id="inf18">
<mml:math id="m26">
<mml:mrow>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> represents the degree of correlation between the predicted value and the true value of the target variable. The closer the value of <inline-formula id="inf19">
<mml:math id="m27">
<mml:mrow>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is to 1, the better the degree of correlation. The calculation formulas are as follows:<disp-formula id="e10">
<mml:math id="m28">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>E</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2211;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mover accent="true">
<mml:mi>Y</mml:mi>
<mml:mo>&#x223c;</mml:mo>
</mml:mover>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(9)</label>
</disp-formula>
<disp-formula id="e11">
<mml:math id="m29">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>E</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mfenced open="|" close="|" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mover accent="true">
<mml:mi>Y</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(10)</label>
</disp-formula>
<disp-formula id="e12">
<mml:math id="m30">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>E</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mfenced open="|" close="|" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mover accent="true">
<mml:mi>Y</mml:mi>
<mml:mo>&#x223c;</mml:mo>
</mml:mover>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:msub>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mfrac>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2a;</mml:mo>
<mml:mn>100</mml:mn>
<mml:mo>%</mml:mo>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(11)</label>
</disp-formula>
<disp-formula id="e13">
<mml:math id="m31">
<mml:mrow>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mover accent="true">
<mml:mi>Y</mml:mi>
<mml:mo>&#x223c;</mml:mo>
</mml:mover>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mover accent="true">
<mml:mi>Y</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(12)</label>
</disp-formula>where <inline-formula id="inf20">
<mml:math id="m32">
<mml:mrow>
<mml:msub>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the VR value calculated by the R/T method. <inline-formula id="inf21">
<mml:math id="m33">
<mml:mrow>
<mml:msub>
<mml:mover accent="true">
<mml:mi>Y</mml:mi>
<mml:mo>&#x223c;</mml:mo>
</mml:mover>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the VR value calculated by the rapid method. <inline-formula id="inf22">
<mml:math id="m34">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>Y</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> represents the average VR of the dispersion curve calculated by the R/T method.</p>
<p>The data in <xref ref-type="table" rid="T2">Table 2</xref> indicate whether the <inline-formula id="inf23">
<mml:math id="m35">
<mml:mrow>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is more than .97, the MAE value is less than 10, and the MAPE is less than 3.5% in the training set or the validation set. These results suggest that the rapid method is highly correlated with the R/T method and has comparable calculation accuracy.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Evaluation standard.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Evaluation standard</th>
<th align="center">Training</th>
<th align="center">Validating</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">MAE</td>
<td align="center">8.30</td>
<td align="center">8.9387</td>
</tr>
<tr>
<td align="center">MSE</td>
<td align="center">160.89</td>
<td align="center">194.6913</td>
</tr>
<tr>
<td align="center">MAPE</td>
<td align="center">2.73%</td>
<td align="center">3.41%</td>
</tr>
<tr>
<td align="center">
<inline-formula id="inf24">
<mml:math id="m36">
<mml:mrow>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
<td align="center">.9803</td>
<td align="center">.9760</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>
<xref ref-type="fig" rid="F8">Figure 8</xref> shows the rapid method results and the ratio curves of 12 randomly selected dispersion curves. <xref ref-type="fig" rid="F8">Figures 8A&#x2013;D</xref> represent relatively smooth dispersion curves, and <xref ref-type="fig" rid="F8">Figures 8E&#x2013;H</xref> represent dispersion curves with local small-amplitude fluctuations. Their ratio curves between the predicted and true values are almost a line, with only slight fluctuations between 2.5 and 5&#xa0;Hz. The results show that the calculation accuracy of the rapid method is nearly the same as that of the R/T method on the dispersion curve, with smooth or small local fluctuations. <xref ref-type="fig" rid="F8">Figures 8I&#x2013;L</xref> represent dispersion curves with strong local fluctuations. There are two positions with local fluctuations. In contrast to the first local fluctuation position, the number of frequency points at the second location is relatively sparse, leading to some deviations in the local change at the high frequency.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>Phase velocity dispersion curve prediction and ratio curve.</p>
</caption>
<graphic xlink:href="feart-10-1084414-g009.tif"/>
</fig>
<p>The trend in <xref ref-type="fig" rid="F9">Figure 9</xref> shows that the time spent constructing the dispersion curve by the R/T method increases linearly with the increase in site layers, and the time spent by the rapid method changes little, if at all. When calculating the site model with 10, 50, and 100 layers, the R/T method takes 15.77&#xa0;s, 49.45&#xa0;s, and 88.18&#xa0;s, respectively. The rapid method takes less than .05&#xa0;s, and its computational cost only slightly increases with the increase in site layers. The rapid method and R/T method are used to calculate all the samples in the validation set and take 59&#xa0;s and 23,473.8&#xa0;s, respectively. The aforementioned analysis and time consumption ratio curve show that the rapid method improves the computational efficiency by at least 400 times while ensuring the same accuracy.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>Calculating time and time consumption ratio.</p>
</caption>
<graphic xlink:href="feart-10-1084414-g007.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="discussion" id="s5">
<title>5 Discussion</title>
<p>Compared with the conventional forward method, the rapid method has high computational efficiency and the same accuracy, but it also has some limitations.</p>
<sec id="s5-1">
<title>5.1 Influence of soft or hard interlayer</title>
<p>The conventional forward modeling method can directly calculate the site of the arbitrary velocity structure; however, the rapid method is affected by the size and distribution of the dataset for training the model. In the validation set, the mean ratio difference (MRD) of all 2000 site samples was calculated (Eq. <xref ref-type="disp-formula" rid="e14">14</xref>). Approximately 2% of the site samples had an MRD greater than 5%. <xref ref-type="fig" rid="F10">Figure 10</xref> shows the dispersion curves and velocity structures of the three site samples with the largest MRD. There is a large soft or hard interlayer in the shallow layer, and the overall velocity structure fluctuates back and forth. That situation makes a small part of the dispersion curves fluctuate locally at high frequencies. Both sites with a back-and-forth fluctuation velocity structure and high-frequency dispersion curve frequency points are very sparse in the dataset, which leads to insufficient DNN learning for those site characteristics. More research is required to balance the back-and-forth fluctuation site ratio and to maintain the main site characteristics in the dataset.<disp-formula id="e14">
<mml:math id="m37">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mrow>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mfenced open="|" close="|" separators="|">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mo>_</mml:mo>
<mml:mi>t</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mo>_</mml:mo>
<mml:mi>p</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
<mml:mo>/</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2a;</mml:mo>
<mml:mn>100</mml:mn>
<mml:mo>%</mml:mo>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(13)</label>
</disp-formula>where y_true and y_pred represent the VR calculated by the R/T method and the VR calculated by the rapid method, respectively.</p>
<fig id="F10" position="float">
<label>FIGURE 10</label>
<caption>
<p>Phase velocity dispersion curve prediction graph and velocity structure graph.</p>
</caption>
<graphic xlink:href="feart-10-1084414-g010.tif"/>
</fig>
</sec>
<sec id="s5-2">
<title>5.2 Selection of frequency points</title>
<p>The frequency points of the rapid method output are fixed. However, in practical engineering, the frequency points are generally equidistant, and the number of required frequency points is determined by the site complex. Therefore, we propose a linear interpolation method to calculate the VR at a specific frequency. Assuming that the VR of point x is needed, the rapid method can be used to calculate the left adjacent point x<sub>0</sub>, VR f (x<sub>0</sub>), right adjacent point x<sub>1</sub>, and VR f (x<sub>1</sub>). Moreover, it is substituted into Formula (14) to calculate the VR of this frequency point.<disp-formula id="e15">
<mml:math id="m38">
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2a;</mml:mo>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2a;</mml:mo>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(14)</label>
</disp-formula>
</p>
<p>As shown in <xref ref-type="fig" rid="F11">Figure 11</xref>, the solid green line and the red dots represent the .4&#xa0;Hz frequency step dispersion curve calculated by the interpolation and R/T methods, respectively. When using the rapid method to construct the dispersion curve with equal frequency steps, the interpolation process may ignore the local fluctuation of the high-frequency band because the high-frequency dispersion points are relatively sparse. Some deviations are applied to the population-based inversion algorithm. Considering that there are relatively few of these kinds of sites in the earth&#x2019;s sedimentary system, we will seek deep learning or data-driven methods to solve this problem in future studies, and the R/T method will be chosen when there is the rare need for high precision. For population-based inversion, our rapid method can be effectively applied to construct the dispersion curve of most sites. The R/T method may be further applied to sites with very thin soft/hard interlayers to improve the inversion accuracy.</p>
<fig id="F11" position="float">
<label>FIGURE 11</label>
<caption>
<p>Interpolated phase velocity dispersion curve graph and its velocity structure graph.</p>
</caption>
<graphic xlink:href="feart-10-1084414-g011.tif"/>
</fig>
</sec>
<sec id="s5-3">
<title>5.3 Influence of the number of datasets</title>
<p>
<xref ref-type="fig" rid="F12">Figure 12A</xref> shows the computational performance of the rapid method for nine datasets of different sizes. With the increase in samples, the box size gradually decreases, the 90% confidence interval gradually shrinks, and the distribution of the prediction results is gradually concentrated. <xref ref-type="fig" rid="F12">Figure 12B</xref> shows the variation in <inline-formula id="inf25">
<mml:math id="m39">
<mml:mrow>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> and MAE with the number of samples for the same nine datasets. As the number of samples in the dataset increases, <inline-formula id="inf26">
<mml:math id="m40">
<mml:mrow>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> rises quickly, and the MAE falls quickly. When the sample size of the dataset is more than 7,000, the change trend slows down significantly. This means that each DNN architecture needs only 7,000 samples, and the model is basically in a convergence state. Therefore, when rebuilding the rapid construction method of other frequency bands, a dataset size set at 7,000 samples is sufficient.</p>
<fig id="F12" position="float">
<label>FIGURE 12</label>
<caption>
<p>Effect of dataset size: <bold>(A)</bold> boxplot and 90% confidence interval; <bold>(B)</bold> R<sup>2</sup> and MAE with increasing sample size.</p>
</caption>
<graphic xlink:href="feart-10-1084414-g012.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="conclusion" id="s6">
<title>6 Conclusion</title>
<p>This paper proposes a deep learning-based method for rapidly constructing 1D dispersion curves. The method consists of two essential parts: building a reasonable random site dataset and building a high-precision, lightweight network architecture. For the dataset, we distill a general empirical relationship between Vs and the burial depth through Vs&#x2013;h empirical relationships in nine different regions of China. Based on this, a technique to generate a random layer site dataset that ensures the diversity and representativity of the dataset and improves the generalization and applicability of the rapid method is proposed. To build the network architecture, the Bayesian optimization algorithm is used to automatically find the optimal DNN architectures, which include three hidden layers, and &#x201c;gelu&#x201d; is applied as the activation function. This network architecture is well suited for learning the relationship between the site velocity structure and the dispersion curve. Finally, a validation set is used to verify the rapid method&#x2019;s accuracy, stability, generalization, and computational efficiency. Based on the analysis of the results, the following conclusions can be drawn:<list list-type="simple">
<list-item>
<p>1) In both the training and validation sets, <inline-formula id="inf27">
<mml:math id="m41">
<mml:mrow>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is more than .97, MAE is less than 10, and MAPE is less than 3.5%. The evaluation result from the aforementioned analysis indicates that the rapid method has comparable accuracy with the R/T method in calculating the dispersion curve.</p>
</list-item>
<list-item>
<p>2) It is difficult for the rapid method to identify the local fluctuations in the high-frequency band of the dispersion curve. In the dataset, 2% of the site samples have an MRD greater than 5%, and in most of those sites, Vs fluctuates back and forth, which causes a local fluctuation in the high-frequency dispersion curve. An implication of this is the possibility that the rapid method is poor at constructing dispersion curves where the site velocity structure fluctuates.</p>
</list-item>
<list-item>
<p>3) Once the training is completed, the rapid method improves the computational efficiency by at least 400 times while ensuring the same accuracy. There is almost no increase in computation expense with an increase in site layers. This method uses a trained neural network to replace the complex calculation processes of many matrix multiplications, transfer matrix iterations and recursive approximations in the matrix method. The computational cost of multiple iterative forward modeling is greatly reduced, especially when the population-based inversion algorithm searches for multilayer site structures.</p>
</list-item>
<list-item>
<p>4) The number of samples impacts the prediction performance of the rapid method. A dataset of 7,000 samples is sufficient to obtain a stable model.</p>
</list-item>
</list>
</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s7">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/<xref ref-type="sec" rid="s12">Supplementary Material</xref>; further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s8">
<title>Author contributions</title>
<p>DC handled most of the work of the article. LS provided the main idea of the article and was the main contributor to guide the revision of the article. KG determined the statistical Vs&#x2013;h empirical relationship and provided some suggestions for the article.</p>
</sec>
<sec id="s9">
<title>Funding</title>
<p>This work was supported by the National Natural Science Foundation of China (Grant No. 51978635).</p>
</sec>
<ack>
<p>The authors thank the editor and reviewers for their valuable comments on this paper and the Institute of Engineering Mechanics, China Earthquake Administration, for supporting this research. All programs have been created using the Python and Tensorflow computational frameworks.</p>
</ack>
<sec sec-type="COI-statement" id="s10">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s11">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors, and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s12">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/feart.2022.1084414/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/feart.2022.1084414/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="Table1.docx" id="SM1" mimetype="application/docx" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Abo-Zena</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>1979</year>). <article-title>Dispersion function computations for unlimited frequency values</article-title>. <source>Geophys. J. Int.</source> <volume>58</volume> (<issue>1</issue>), <fpage>91</fpage>&#x2013;<lpage>105</lpage>. <pub-id pub-id-type="doi">10.1111/j.1365-246X.1979.tb01011.x</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Aleardi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Stucchi</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A hybrid residual neural network&#x2013;Monte Carlo approach to invert surface wave dispersion data</article-title>. <source>Near Surf. Geophys.</source> <volume>19</volume> (<issue>4</issue>), <fpage>397</fpage>&#x2013;<lpage>414</lpage>. <pub-id pub-id-type="doi">10.1002/nsg.12163</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Alwosheel</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>van Cranenburgh</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Chorus</surname>
<given-names>C. G.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Is your dataset big enough? Sample size requirements when using artificial neural networks for discrete choice analysis</article-title>. <source>J. choice Model.</source> <volume>28</volume>, <fpage>167</fpage>&#x2013;<lpage>182</lpage>. <pub-id pub-id-type="doi">10.1016/j.jocm.2018.07.002</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Alyousuf</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Colombo</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Rovetta</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Sandoval-Curiel</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Near-surface velocity analysis for single-sensor data: An integrated workflow using surface waves, AI, and structure-regularized inversion</article-title>,&#x201d; in <source>SEG technical program expanded abstracts 2018 SEG technical program expanded abstracts</source> (<publisher-loc>Tulsa, Oklahoma</publisher-loc>: <publisher-name>Society of Exploration Geophysicists</publisher-name>), <fpage>2342</fpage>&#x2013;<lpage>2346</lpage>. <pub-id pub-id-type="doi">10.1190/segam2018-2994696.1</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Assi</surname>
<given-names>K. J.</given-names>
</name>
<name>
<surname>Nahiduzzaman</surname>
<given-names>K. M.</given-names>
</name>
<name>
<surname>Ratrout</surname>
<given-names>N. T.</given-names>
</name>
<name>
<surname>Aldosary</surname>
<given-names>A. S.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Mode choice behavior of high school goers: Evaluating logistic regression and MLP neural networks</article-title>. <source>Case Stud. Transp. policy</source> <volume>6</volume> (<issue>2</issue>), <fpage>225</fpage>&#x2013;<lpage>230</lpage>. <pub-id pub-id-type="doi">10.1016/j.cstp.2018.04.006</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Aung</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Leong</surname>
<given-names>E. C.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Discussion of &#x201c;near-field effects on array-based surface wave methods with active sources&#x201d; by S. Yoon and G. J. Rix</article-title>. <source>J. Geotechnical Geoenvironmental Eng.</source> <volume>136</volume>, <fpage>773</fpage>&#x2013;<lpage>775</lpage>. <pub-id pub-id-type="doi">10.1061/(ASCE)GT.1943-5606.0000131</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Bishop</surname>
<given-names>C. M.</given-names>
</name>
</person-group> (<year>1995</year>). <source>Neural networks for pattern recognition</source>. <comment>Clarendon Press</comment>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Oxford University Press</publisher-name>.</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Boore</surname>
<given-names>D. M.</given-names>
</name>
</person-group> (<year>1972</year>). <article-title>Finite difference methods for seismic wave propagation in heterogeneous materials</article-title>. <source>Methods Comput. Phys.</source> <volume>11</volume>, <fpage>1</fpage>&#x2013;<lpage>37</lpage>. <pub-id pub-id-type="doi">10.1016/B978-0-12-460811-5.50006-4</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chauhan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Dhingra</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Pattern recognition system using MLP neural networks</article-title>. <source>Pattern Recognit.</source> <volume>4</volume> (<issue>9</issue>), <fpage>990</fpage>&#x2013;<lpage>993</lpage>. <pub-id pub-id-type="doi">10.9790/3021-0205990993</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>1993</year>). <article-title>A systematic and efficient method of computing normal modes for multilayered half-space</article-title>. <source>Geophys. J. Int.</source> <volume>115</volume> (<issue>2</issue>), <fpage>391</fpage>&#x2013;<lpage>409</lpage>. <pub-id pub-id-type="doi">10.1111/j.1365-246X.1993.tb01194.x</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dai</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Ning</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Xi</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xing</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Deep learning for extracting dispersion curves</article-title>. <source>Surv. Geophys.</source> <volume>42</volume> (<issue>1</issue>), <fpage>69</fpage>&#x2013;<lpage>95</lpage>. <pub-id pub-id-type="doi">10.1007/s10712-020-09615-3</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Donoho</surname>
<given-names>D. L.</given-names>
</name>
</person-group> (<year>1995</year>). <article-title>De-noising by soft-thresholding</article-title>. <source>IEEE Trans. Inf. theory</source> <volume>41</volume> (<issue>3</issue>), <fpage>613</fpage>&#x2013;<lpage>627</lpage>. <pub-id pub-id-type="doi">10.1109/18.382009</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Faccioli</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Maggio</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Quarteroni</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Taghan</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>1996</year>). <article-title>Spectral-domain decomposition methods for the solution of acoustic and elastic wave equations</article-title>. <source>Geophysics</source> <volume>61</volume> (<issue>4</issue>), <fpage>1160</fpage>&#x2013;<lpage>1174</lpage>.</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fan</surname>
<given-names>Y. H.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>J. Q.</given-names>
</name>
<name>
<surname>Xiao</surname>
<given-names>B. X.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Fast vector-transfer algorithm for computation of Rayleigh wave dispersion curves</article-title>. <source>J. Hunan Univ.</source> <volume>29</volume> (<issue>5</issue>), <fpage>25</fpage>&#x2013;<lpage>30</lpage>. <pub-id pub-id-type="doi">10.1190/1.1444036</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Feurer</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Klein</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Eggensperger</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Springenberg</surname>
<given-names>J. T.</given-names>
</name>
<name>
<surname>Blum</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hutter</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2019</year>). <source>Auto-sklearn: Efficient and robust automated machine learning, part of the Springer series on challenges in machine learning book series (SSCML)</source>. <publisher-loc>Berlin, Germany</publisher-loc>: <publisher-name>Spinger</publisher-name>. <pub-id pub-id-type="doi">10.1007/978-3-030-05318-5_6</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Frazier</surname>
<given-names>P. I.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Bayesian optimization</article-title>. <source>Inf. TutORials Operations Res.</source> <volume>2018</volume>, <fpage>255</fpage>&#x2013;<lpage>278</lpage>. <pub-id pub-id-type="doi">10.1287/educ.2018.0188</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Pan</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Dong</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Retrieving S-wave velocity from surface wave multimode dispersion curves with DispINet</article-title>. <source>J. Appl. Geophys.</source> <volume>193</volume>, <fpage>104430</fpage>. <pub-id pub-id-type="doi">10.1016/j.jappgeo.2021.104430</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guo</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Ba</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Prestack seismic inversion with data-driven MRF-based regularization</article-title>. <source>IEEE Trans. Geoscience Remote Sens.</source> <volume>59</volume>, <fpage>7122</fpage>&#x2013;<lpage>7136</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2020.3019715</pub-id>
</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Haskell</surname>
<given-names>N. A.</given-names>
</name>
</person-group> (<year>1953</year>). <article-title>The dispersion of surface waves on multilayered media</article-title>. <source>Bull. Seismol. Soc. Am.</source> <volume>43</volume> (<issue>1</issue>), <fpage>17</fpage>&#x2013;<lpage>34</lpage>. <pub-id pub-id-type="doi">10.1785/BSSA0430010017</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>Y. F.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>C. F.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Normal mode computation by the generalized reflection transmission coefficient method in planar layered half space</article-title>. <source>Chin. J. Geophys.</source> <volume>49</volume> (<issue>4</issue>), <fpage>1074</fpage>&#x2013;<lpage>1081</lpage>. <pub-id pub-id-type="doi">10.3321/j.issn:0001-5733.2006.04.020</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hisada</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>1994</year>). <article-title>An efficient method for computing Green&#x27;s functions for a layered half-space with sources and receivers at close depths</article-title>. <source>Bull. Seismol. Soc. Am.</source> <volume>84</volume> (<issue>5</issue>), <fpage>1456</fpage>&#x2013;<lpage>1472</lpage>. <pub-id pub-id-type="doi">10.1785/BSSA0840051456</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hisada</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>1995</year>). <article-title>An efficient method for computing Green&#x27;s functions for a layered half-space with sources and receivers at close depths (Part 2)</article-title>. <source>Bull. Seismol. Soc. Am.</source> <volume>85</volume> (<issue>4</issue>), <fpage>1080</fpage>&#x2013;<lpage>1093</lpage>. <pub-id pub-id-type="doi">10.1785/BSSA0850041080</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Itano</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>de Sousa</surname>
<given-names>M. A. D. A.</given-names>
</name>
<name>
<surname>Del-Moral-Hernandez</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Extending MLP ANN hyper-parameters optimization by using genetic algorithm</article-title>,&#x201d; in <conf-name>2018 International joint conference on neural networks (IJCNN)</conf-name>, <conf-loc>Rio de Janeiro, Brazil</conf-loc>, <conf-date>08-13 July 2018</conf-date> (<publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1109/IJCNN.2018.8489520</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jo</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Jeong</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Min</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Park</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Kwon</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2022</year>). <article-title>Efficient deep-learning-based history matching for fluvial channel reservoirs</article-title>. <source>J. Petroleum Sci. Eng.</source> <volume>208</volume>, <fpage>109247</fpage>. <pub-id pub-id-type="doi">10.1016/j.petrol.2021.109247</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kennett</surname>
<given-names>B. L. N.</given-names>
</name>
<name>
<surname>Clarke</surname>
<given-names>T. J.</given-names>
</name>
</person-group> (<year>1983</year>). <article-title>Seismic waves in a stratified half-space&#x2014;IV: P&#x2014;SV wave decoupling and surface wave dispersion</article-title>. <source>Geophys. J. Int.</source> <volume>72</volume> (<issue>3</issue>), <fpage>633</fpage>&#x2013;<lpage>645</lpage>. <pub-id pub-id-type="doi">10.1111/j.1365-246X.1983.tb02824.x</pub-id>
</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kennett</surname>
<given-names>B. L. N.</given-names>
</name>
<name>
<surname>Kerry</surname>
<given-names>N. J.</given-names>
</name>
</person-group> (<year>1979</year>). <article-title>Seismic waves in a stratified half space</article-title>. <source>Geophys. J. Int.</source> <volume>57</volume> (<issue>3</issue>), <fpage>557</fpage>&#x2013;<lpage>583</lpage>. <pub-id pub-id-type="doi">10.1111/j.1365-246X.1979.tb06779.x</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kennett</surname>
<given-names>B. L. N.</given-names>
</name>
</person-group> (<year>1974</year>). <article-title>Reflections, rays, and reverberations</article-title>. <source>Bull. Seismol. Soc. Am.</source> <volume>64</volume> (<issue>6</issue>), <fpage>1685</fpage>&#x2013;<lpage>1696</lpage>. <pub-id pub-id-type="doi">10.1785/BSSA0640061685</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kingma</surname>
<given-names>D. P.</given-names>
</name>
<name>
<surname>Ba</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Adam: A method for stochastic optimization</article-title>. <comment>arXiv preprint arXiv:1412.6980</comment>. <pub-id pub-id-type="doi">10.48550/arXiv.1412.6980</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Komatitsch</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Vilotte</surname>
<given-names>J. P.</given-names>
</name>
</person-group> (<year>1998</year>). <article-title>The spectral element method: An efficient tool to simulate the seismic response of 2D and 3D geological structures</article-title>. <source>Bull. Seismol. Soc. Am.</source> <volume>88</volume> (<issue>2</issue>), <fpage>368</fpage>&#x2013;<lpage>392</lpage>. <pub-id pub-id-type="doi">10.1785/BSSA0880020368</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lei</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Shen</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Inversion of Rayleigh wave dispersion curves via adaptive GA and nested DLS</article-title>. <source>Geophys. J. Int.</source> <volume>218</volume> (<issue>1</issue>), <fpage>547</fpage>&#x2013;<lpage>559</lpage>. <pub-id pub-id-type="doi">10.1093/gji/ggz171</pub-id>
</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liang</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Image&#x2010;based post&#x2010;disaster inspection of reinforced concrete bridge systems using deep learning with Bayesian optimization</article-title>. <source>Computer&#x2010;Aided Civ. Infrastructure Eng.</source> <volume>34</volume> (<issue>5</issue>), <fpage>415</fpage>&#x2013;<lpage>430</lpage>. <pub-id pub-id-type="doi">10.1111/mice.12425</pub-id>
</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Du</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Rayleigh wave inversion using heat-bath simulated annealing algorithm</article-title>. <source>J. Appl. Geophys.</source> <volume>134</volume>, <fpage>267</fpage>&#x2013;<lpage>280</lpage>. <pub-id pub-id-type="doi">10.1016/j.jappgeo.2016.09.008</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Luco</surname>
<given-names>J. E.</given-names>
</name>
<name>
<surname>Apsel</surname>
<given-names>R. J.</given-names>
</name>
</person-group> (<year>1983</year>). <article-title>On the Green&#x27;s functions for a layered half-space. Part I</article-title>. <source>Bull. Seismol. Soc. Am.</source> <volume>73</volume> (<issue>4</issue>), <fpage>909</fpage>&#x2013;<lpage>929</lpage>. <pub-id pub-id-type="doi">10.1785/BSSA0730040909</pub-id>
</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Luo</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Ba</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Carcione</surname>
<given-names>J. M.</given-names>
</name>
</person-group> (<year>2022a</year>). <article-title>A hierarchical prestack seismic inversion scheme for VTI media based on the exact reflection coefficient</article-title>. <source>IEEE Trans. Geoscience Remote Sens.</source> <volume>60</volume>, <fpage>1</fpage>&#x2013;<lpage>16</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2021.3140133</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Luo</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2022b</year>). <article-title>Constructing shear velocity models from surface wave dispersion curves using deep learning</article-title>. <source>J. Appl. Geophys.</source> <volume>196</volume>, <fpage>104524</fpage>. <pub-id pub-id-type="doi">10.1016/j.jappgeo.2021.104524</pub-id>
</citation>
</ref>
<ref id="B36">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Manolis</surname>
<given-names>G. D.</given-names>
</name>
<name>
<surname>Beskos</surname>
<given-names>D. E.</given-names>
</name>
</person-group> (<year>1988</year>). <source>Boundary element methods in elastodynamics</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Taylor &#x26; Francis</publisher-name>. <pub-id pub-id-type="doi">10.1115/1.3173761</pub-id>
</citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nabian</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Meidani</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Deep learning for accelerated seismic reliability analysis of transportation networks</article-title>. <source>Computer-Aided Civ. Infrastructure Eng.</source> <volume>33</volume>, <fpage>443</fpage>&#x2013;<lpage>458</lpage>. <pub-id pub-id-type="doi">10.1111/mice.12359</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pan</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Yuan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Modified generalized R/T coefficient method for surface&#x2010;wave dispersion&#x2010;curve calculation in elastic and viscoelastic media</article-title>. <source>Bull. Seismol. Soc. Am.</source> <volume>112</volume>, <fpage>2280</fpage>&#x2013;<lpage>2296</lpage>. <pub-id pub-id-type="doi">10.1785/0120210294</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Picozzi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Albarello</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Combining genetic and linearized algorithms for a two-step joint inversion of Rayleigh wave dispersion and H/V spectral ratio curves</article-title>. <source>Geophys. J. Int.</source> <volume>169</volume> (<issue>1</issue>), <fpage>189</fpage>&#x2013;<lpage>200</lpage>. <pub-id pub-id-type="doi">10.1111/j.1365-246X.2006.03282.x</pub-id>
</citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Poormirzaee</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Fister</surname>
<given-names>I.</given-names>
<suffix>Jr</suffix>
</name>
</person-group> (<year>2021</year>). <article-title>Model-based inversion of Rayleigh wave dispersion curves via linear and nonlinear methods</article-title>. <source>Pure Appl. Geophys.</source> <volume>178</volume> (<issue>2</issue>), <fpage>341</fpage>&#x2013;<lpage>358</lpage>. <pub-id pub-id-type="doi">10.1007/s00024-021-02665-7</pub-id>
</citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Poormirzaee</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>S-wave velocity profiling from refraction microtremor Rayleigh wave dispersion curves via PSO inversion algorithm</article-title>. <source>Arabian J. Geosciences</source> <volume>9</volume> (<issue>16</issue>), <fpage>673</fpage>&#x2013;<lpage>710</lpage>. <pub-id pub-id-type="doi">10.1007/s12517-016-2701-6</pub-id>
</citation>
</ref>
<ref id="B42">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Prechelt</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Neural networks: Tricks of the trade</article-title>,&#x201d; in <source>Early stopping&#x2014;but when?</source> Editors <person-group person-group-type="editor">
<name>
<surname>Montavon</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Orr</surname>
<given-names>G. B.</given-names>
</name>
<name>
<surname>M&#xfc;ller</surname>
<given-names>K. R.</given-names>
</name>
</person-group> <edition>second edition</edition> (<publisher-loc>Berlin, Heidelberg</publisher-loc>: <publisher-name>Springer Berlin Heidelberg</publisher-name>), <fpage>53</fpage>&#x2013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-642-35289-8_5</pub-id>
</citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sameen</surname>
<given-names>M. I.</given-names>
</name>
<name>
<surname>Pradhan</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Application of convolutional neural networks featuring Bayesian optimization for landslide susceptibility assessment</article-title>. <source>Catena</source> <volume>186</volume>, <fpage>104249</fpage>. <pub-id pub-id-type="doi">10.1016/j.catena.2019.104249</pub-id>
</citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schwab</surname>
<given-names>F. A.</given-names>
</name>
<name>
<surname>Knopoff</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>1972</year>). <article-title>Fast surface wave and free mode computations</article-title>. <source>Methods Comput. Phys. Adv. Res. Appl.</source> <volume>11</volume>, <fpage>87</fpage>&#x2013;<lpage>180</lpage>. <pub-id pub-id-type="doi">10.1016/B978-0-12-460811-5.50008-8</pub-id>
</citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schwab</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Knopoff</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>1970</year>). <article-title>Surface-wave dispersion computations</article-title>. <source>Bull. Seismol. Soc. Am.</source> <volume>60</volume> (<issue>2</issue>), <fpage>321</fpage>&#x2013;<lpage>344</lpage>. <pub-id pub-id-type="doi">10.1785/BSSA0600020321</pub-id>
</citation>
</ref>
<ref id="B46">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Snoek</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Rippel</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Swersky</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kiros</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Satish</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Sundaram</surname>
<given-names>N.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Scalable bayesian optimization using deep neural networks <italic>Proceedings of the 32nd international Conference on machine learning</italic> (PMLR), 2171&#x2013;2180</article-title>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://proceedings.mlr.press/v37/snoek15.html">https://proceedings.mlr.press/v37/snoek15.html</ext-link> (Accessed October 29, 2022)</comment>.</citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Takeuchi</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Saito</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>1972</year>). <article-title>Seismic surface waves</article-title>. <source>Methods Comput. Phys.</source> <volume>11</volume>, <fpage>217</fpage>&#x2013;<lpage>295</lpage>. <pub-id pub-id-type="doi">10.1016/B978-0-12-460811-5.50010-6</pub-id>
</citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Thomson</surname>
<given-names>W. T.</given-names>
</name>
</person-group> (<year>1950</year>). <article-title>Transmission of elastic waves through a stratified solid medium</article-title>. <source>J. Appl. Phys.</source> <volume>21</volume> (<issue>2</issue>), <fpage>89</fpage>&#x2013;<lpage>93</lpage>. <pub-id pub-id-type="doi">10.1063/1.1699629</pub-id>
</citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tschannen</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Ghanim</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Ettrich</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Partial automation of the seismic to well tie with deep learning and Bayesian optimization</article-title>. <source>Comput. Geosciences</source> <volume>164</volume>, <fpage>105120</fpage>. <pub-id pub-id-type="doi">10.1016/j.cageo.2022.105120</pub-id>
</citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wamriew</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Charara</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Pissarenko</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Joint event location and velocity model update in real-time for downhole microseismic monitoring: A deep learning approach</article-title>. <source>Comput. Geosciences</source> <volume>158</volume>, <fpage>104965</fpage>. <pub-id pub-id-type="doi">10.1016/j.cageo.2021.104965</pub-id>
</citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Werbos</surname>
<given-names>P. J.</given-names>
</name>
</person-group> (<year>1988</year>). <article-title>Generalization of backpropagation with application to a recurrent gas market model</article-title>. <source>Neural Netw.</source> <volume>1</volume> (<issue>4</issue>), <fpage>339</fpage>&#x2013;<lpage>356</lpage>. <pub-id pub-id-type="doi">10.1016/0893-6080(88)90007-X</pub-id>
</citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wong</surname>
<given-names>T. T.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Performance evaluation of classification algorithms by k-fold and leave-one-out cross validation</article-title>. <source>Pattern Recognit.</source> <volume>48</volume> (<issue>9</issue>), <fpage>2839</fpage>&#x2013;<lpage>2846</lpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2015.03.009</pub-id>
</citation>
</ref>
<ref id="B53">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Jia</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Ross</surname>
<given-names>Z. E.</given-names>
</name>
<name>
<surname>Clayton</surname>
<given-names>R. W.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Extracting dispersion curves from ambient noise correlations using deep learning</article-title>. <source>IEEE Trans. Geoscience Remote Sens.</source> <volume>58</volume>, <fpage>8932</fpage>&#x2013;<lpage>8939</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2020.2992043</pub-id>
</citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Qiu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Armaghani</surname>
<given-names>D. J.</given-names>
</name>
<name>
<surname>Khandelwal</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Mohamad</surname>
<given-names>E. T.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Estimation of the TBM advance rate under hard rock conditions using XGBoost and Bayesian optimization</article-title>. <source>Undergr. Space</source> <volume>6</volume> (<issue>5</issue>), <fpage>506</fpage>&#x2013;<lpage>515</lpage>. <pub-id pub-id-type="doi">10.1016/j.undsp.2020.05.008</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>