<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Radiol.</journal-id>
<journal-title>Frontiers in Radiology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Radiol.</abbrev-journal-title>
<issn pub-type="epub">2673-8740</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fradi.2022.866974</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Radiology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Informative and Reliable Tract Segmentation for Preoperative Planning</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Lucena</surname> <given-names>Oeslle</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1652695/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Borges</surname> <given-names>Pedro</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Cardoso</surname> <given-names>Jorge</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Ashkan</surname> <given-names>Keyoumars</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/749688/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Sparks</surname> <given-names>Rachel</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Ourselin</surname> <given-names>Sebastien</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/222555/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>School of Biomedical Engineering and Imaging Sciences, King&#x00027;s College London</institution>, <addr-line>London</addr-line>, <country>United Kingdom</country></aff>
<aff id="aff2"><sup>2</sup><institution>Medical Physics and Biomedical Engineering, University College London</institution>, <addr-line>London</addr-line>, <country>United Kingdom</country></aff>
<aff id="aff3"><sup>3</sup><institution>King&#x00027;s College Hospital Foundation Trust</institution>, <addr-line>London</addr-line>, <country>United Kingdom</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Chuyang Ye, Beijing Institute of Technology, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Ye Wu, University of North Carolina at Chapel Hill, United States; Dong Hye Ye, Marquette University, United States; Hui Cui, La Trobe University, Australia</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Oeslle Lucena <email>oeslle.lucena&#x00040;kcl.ac.uk</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Artificial Intelligence in Radiology, a section of the journal Frontiers in Radiology</p></fn></author-notes>
<pub-date pub-type="epub">
<day>18</day>
<month>05</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>2</volume>
<elocation-id>866974</elocation-id>
<history>
<date date-type="received">
<day>31</day>
<month>01</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>31</day>
<month>03</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Lucena, Borges, Cardoso, Ashkan, Sparks and Ourselin.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Lucena, Borges, Cardoso, Ashkan, Sparks and Ourselin</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Identifying white matter (WM) tracts to locate eloquent areas for preoperative surgical planning is a challenging task. Manual WM tract annotations are often used but they are time-consuming, suffer from inter- and intra-rater variability, and noise intrinsic to diffusion MRI may make manual interpretation difficult. As a result, in clinical practice direct electrical stimulation is necessary to precisely locate WM tracts during surgery. A measure of WM tract segmentation unreliability could be important to guide surgical planning and operations. In this study, we use deep learning to perform reliable tract segmentation in combination with uncertainty quantification to measure segmentation unreliability. We use a 3D U-Net to segment white matter tracts. We then estimate model and data uncertainty using test time dropout and test time augmentation, respectively. We use a volume-based calibration approach to compute representative predicted probabilities from the estimated uncertainties. In our findings, we obtain a Dice of &#x02248;0.82 which is comparable to the state-of-the-art for multi-label segmentation and Hausdorff distance &#x0003C;10<italic>mm</italic>. We demonstrate a high positive correlation between volume variance and segmentation errors, which indicates a good measure of reliability for tract segmentation ad uncertainty estimation. Finally, we show that calibrated predicted volumes are more likely to encompass the ground truth segmentation volume than uncalibrated predicted volumes. This study is a step toward more informed and reliable WM tract segmentation for clinical decision-making.</p></abstract>
<kwd-group>
<kwd>diffusion MRI</kwd>
<kwd>tract segmentation</kwd>
<kwd>deep learning</kwd>
<kwd>uncertainty quantification</kwd>
<kwd>calibration</kwd>
<kwd>tractography</kwd>
</kwd-group>
<contract-sponsor id="cn001">EPSRC Centre for Doctoral Training in Medical Imaging<named-content content-type="fundref-id">10.13039/501100013915</named-content></contract-sponsor>
<counts>
<fig-count count="7"/>
<table-count count="5"/>
<equation-count count="3"/>
<ref-count count="65"/>
<page-count count="15"/>
<word-count count="9262"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Segmentation of white matter (WM) tracts is important in several tasks including understanding brain organization, preoperative neurosurgical planning to identify eloquent areas, and identification of surgical approaches to reducing post-operative damage (<xref ref-type="bibr" rid="B1">1</xref>, <xref ref-type="bibr" rid="B2">2</xref>). Clinically, manual tract annotations can help to plan a surgical approach but direct electrical stimulation is often used in complex cases as the ground truth to determine the precise location of eloquent areas during surgery (<xref ref-type="bibr" rid="B3">3</xref>). Manual tract annotations are time-consuming, often relying on fine-tuning tractography which depends on diffusion MRI (dMRI) acquisition parameters (<xref ref-type="bibr" rid="B4">4</xref>, <xref ref-type="bibr" rid="B5">5</xref>), and suffer from inter- and intra-rater variability due to the complexity of the WM tracts (<xref ref-type="bibr" rid="B6">6</xref>&#x02013;<xref ref-type="bibr" rid="B8">8</xref>). Automatic segmentation approaches have emerged to produce faster and more reproducible tract annotations. Nonetheless, to the best of our knowledge, none of the automatic approaches has investigated model reliability, such as uncertainty awareness, for tract segmentation.</p>
<p>Automatic tract segmentation methods can be divided into (1) region-of-interest-based (ROI) (or connectivity-based), (2) clustering-based, and (3) direct segmentation (<xref ref-type="bibr" rid="B8">8</xref>&#x02013;<xref ref-type="bibr" rid="B10">10</xref>). ROI-based approaches focus on filtering tract fibers based on known anatomical regions before or after computing whole-brain tractography (<xref ref-type="bibr" rid="B11">11</xref>, <xref ref-type="bibr" rid="B12">12</xref>). ROI-based methods require either brain parcellation or registration of subject-specific images to an atlas. Clustering-based methods focus on computing similarity metrics (i.e., distance) to classify WM fiber tracts (<xref ref-type="bibr" rid="B13">13</xref>&#x02013;<xref ref-type="bibr" rid="B15">15</xref>). These approaches are computationally expensive due to the need to perform registration, whole-brain parcellation, or other preprocessing steps. Additionally, these methods demonstrate poor reproducibility in tracts that have high anatomical variability across subjects (<xref ref-type="bibr" rid="B16">16</xref>) which may introduce an unacceptable level of risk to the patient for preoperative neurosurgical planning.</p>
<p>Direct methods output tract masks or fiber tracts directly from input data without the intermediate steps of ROI-based or clustering-based methods (<xref ref-type="bibr" rid="B10">10</xref>). Direct methods can be divided into voxel-based or fiber-based classification approaches. Voxel-based methods classify voxels as being inside or outside a specific WM tract from volumetric input data while fiber-based methods classify whether or not a particular fiber belongs to a specific tract. Deep learning-based (DL) methods are currently the state-of-the-art for direct WM tract segmentation (<xref ref-type="bibr" rid="B8">8</xref>, <xref ref-type="bibr" rid="B10">10</xref>, <xref ref-type="bibr" rid="B17">17</xref>, <xref ref-type="bibr" rid="B18">18</xref>). Once trained, DL models can quickly perform inference (<xref ref-type="bibr" rid="B19">19</xref>).</p>
<p>TractSeg (<xref ref-type="bibr" rid="B10">10</xref>), a voxel-based approach, uses 2D U-Nets (<xref ref-type="bibr" rid="B20">20</xref>), in a tri-planar approach, to segment 72 tracts. TractSeg uses as input the 3 major peak directions obtained from fiber orientation distributions (FODs) computed from constrained spherical deconvolution (CSD) (<xref ref-type="bibr" rid="B21">21</xref>). Neuro4Neuro (<xref ref-type="bibr" rid="B17">17</xref>), another voxel-based method, uses a 3D U-Net (<xref ref-type="bibr" rid="B22">22</xref>) to segment 25 tracts from an in-house dataset. As input for the 3D convolutional neural network (CNN), Neuro4Neuro uses diffusion tensor imaging (DTI) (<xref ref-type="bibr" rid="B23">23</xref>). DeepWMA (<xref ref-type="bibr" rid="B18">18</xref>), a fiber-based approach, uses a 2D CNN to classify fibers as belonging to one of 54 possible WM tracts. As input, DeepWMA uses a 2D multi-channel fiber feature descriptor computed for fibers obtained from whole-brain tractography. Similarly to DeepWMA, Classifyber (<xref ref-type="bibr" rid="B8">8</xref>) uses a set of features (i.e., spatial position, connectivity, etc) to describe a fiber of a tract and classify fibers into specific WM tracts. Classifyber uses a logistic regression (LR) model for classification.</p>
<p>Deep learning approaches tend to be overconfident in their segmentations which can lead to mistaken conclusions. For instance, underestimating the likelihood of a voxel being a false positive, missing a pathologic finding, or false negative leads to damaging an eloquent region (<xref ref-type="bibr" rid="B1">1</xref>, <xref ref-type="bibr" rid="B24">24</xref>). Therefore, it is important to ensure model uncertainty is reflective of the ground truth data. Uncertainty estimation is also important as it enables DL results to be more transparent to the end-user giving clinicians more confidence in segmentation results.</p>
<p>Uncertainty quantification (UQ) can be used as a metric of reliability for DL approaches. UQ in DL has been investigated for a variety of medical imaging applications, such as physics-informed uncertainty-aware Brain MRI segmentation (<xref ref-type="bibr" rid="B25">25</xref>), modality synthesis (<xref ref-type="bibr" rid="B26">26</xref>), dMRI super resolution (<xref ref-type="bibr" rid="B27">27</xref>), brain parcellation (<xref ref-type="bibr" rid="B28">28</xref>), electrode bending prediction (<xref ref-type="bibr" rid="B29">29</xref>, <xref ref-type="bibr" rid="B30">30</xref>), and tumor segmentation (<xref ref-type="bibr" rid="B24">24</xref>, <xref ref-type="bibr" rid="B31">31</xref>, <xref ref-type="bibr" rid="B32">32</xref>).</p>
<p>Uncertainty quantification is often divided into two types: epistemic, noise caused by variation in the model&#x00027;s parameters, or aleatoric, the noise inherent in the data (<xref ref-type="bibr" rid="B33">33</xref>). For tract segmentation, we expect the uncertainty to be both data and model-dependent. Thus, it is important to know where the model&#x00027;s parameters&#x00027; variability causes uncertainty and where noise in the data causes uncertainty.</p>
<p>Epistemic uncertainty can be computed through Bayesian inference networks (BNNs) (<xref ref-type="bibr" rid="B33">33</xref>). BNNs offer a mathematically grounded method where they compute distribution functions for the trained parameters instead of regular scalars as in regular neural networks. However, they are hard to implement, and their training stage is computationally expensive (<xref ref-type="bibr" rid="B33">33</xref>). Bayesian approximation using dropout layers at the inference stage has been proposed to overcome training limitations of BNNs by doing multiple forward inferences (<xref ref-type="bibr" rid="B34">34</xref>) and has been successfully applied to medical imaging tasks (<xref ref-type="bibr" rid="B24">24</xref>, <xref ref-type="bibr" rid="B27">27</xref>, <xref ref-type="bibr" rid="B29">29</xref>). Following the Bayesian inference approximation, other methods such as Markov chain Monte Carlo (MCMC) (<xref ref-type="bibr" rid="B35">35</xref>) and Monte Carlo Batch Normalization (MCBN) (<xref ref-type="bibr" rid="B36">36</xref>) have been proposed where batch normalization at the inference stage approximates the outputs of a BNN.</p>
<p>Aleatoric uncertainty <italic>via</italic> learned loss attenuation, where a network is designed to have two branches one for the final prediction and one for uncertainty, has been proposed by Kendall and Gal (<xref ref-type="bibr" rid="B33">33</xref>). While this has been successfully applied to medical imaging tasks (<xref ref-type="bibr" rid="B27">27</xref>, <xref ref-type="bibr" rid="B31">31</xref>, <xref ref-type="bibr" rid="B37">37</xref>), the addition of a second branch makes the network challenging to train and prone to instability. Another method to compute aleatoric uncertainty is to augment input data at the inference stage and compute uncertainty over several rounds of inference. This approach is easy to implement once a CNN is trained and does not require modifying network architecture or retraining. Test time augmentation has been shown robust in medical imaging (<xref ref-type="bibr" rid="B32">32</xref>, <xref ref-type="bibr" rid="B38">38</xref>).</p>
<p>An accurate segmentation model is important to achieve the best possible results and enable uncertainty to be generally low so that it highlights regions that are difficult to segment (either due to data or model limitations). However, the probability associated with the predicted class label does not always reflect its ground truth likelihood (<xref ref-type="bibr" rid="B39">39</xref>). Calibration makes predicted probabilities more aligned to the ground truth accuracy, meaning that output predictions reflect a measurable property in the annotations of a validation dataset. Calibration has been widely used as a post-processing step, e.g., in classification (<xref ref-type="bibr" rid="B40">40</xref>&#x02013;<xref ref-type="bibr" rid="B42">42</xref>) and segmentation tasks (<xref ref-type="bibr" rid="B24">24</xref>, <xref ref-type="bibr" rid="B43">43</xref>).</p>
<p>In this study, we aim to provide uncertainty awareness for tract segmentation with accurate and reliable predicted probabilities so that clinicians can use it as a safety tool in preoperative neurosurgical planning. We present a 3D CNN that takes as input raw dMRI intensities transformed into the spherical harmonics (SH) space to align data across subjects. We design a system to output calibrated epistemic and aleatoric uncertainties that are reflective of measured ground truth volumes. We demonstrate that our approach has comparable performance to the state-of-the-art tract segmentation approaches while providing an estimation of model and data uncertainty. The significance of this study is that it provides a method to augment information so clinicians can make more informed clinical choices.</p>
</sec>
<sec sec-type="materials and methods" id="s2">
<title>2. Materials and Methods</title>
<sec>
<title>2.1. Pipeline Overview</title>
<p>We project dMRI signal intensities into the SH space (Section 2.3) to align data across different acquisitions without fitting a model. Next, we train a 3D CNN to segment WM tracts from the SH coefficients (Section 2.3). Given a trained model, we calculate epistemic uncertainty (Section 2.7.1) and aleatoric uncertainty (Section 2.7.2). Finally, we perform volume-based calibration to make predicted probabilities and uncertainty measurements more representative of the ground truth volume (Section 2.8).</p>
</sec>
<sec>
<title>2.2. Dataset</title>
<p>We use dMRI from 105 subjects provided by the Human Connectome Project (HCP) (<xref ref-type="bibr" rid="B44">44</xref>). HCP dMRI were acquired on a 3T scanner with the following parameters: the spatial size of 145 &#x000D7; 174 &#x000D7; 145 with 1.25 mm isotropic resolution, 90 gradient directions for each b = &#x02208; {1,000, 2,000, 3,000 s/mm<sup>2</sup>} and 18 images at b = 0 s/mm<sup>2</sup>. Data is corrected following the protocols described in Sotiropoulos et al. (<xref ref-type="bibr" rid="B44">44</xref>) prior to download. For each one of the 105 HCP subjects, a set of 72 annotated tracts in the 3D spatial coordinate space, corrected by a human rater is provided by Wasserthal et al. (<xref ref-type="bibr" rid="B10">10</xref>) and available for download<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref>. For each tract, a binary mask was generated similar to the approach of TractSeg (<xref ref-type="bibr" rid="B10">10</xref>). We set a voxel to the foreground if one or more fibers are present within the voxel.</p>
</sec>
<sec>
<title>2.3. Data Preprocessing</title>
<p>A single-shell (b = 2,000 s/mm<sup>2</sup>) was selected, and its dMRI signal intensities are transformed into SH coefficients without any model fitting. SH coefficients are then normalized by the b-zero shell using the algorithm provided in MRtrix (<xref ref-type="bibr" rid="B45">45</xref>). Then, we clamped all SH voxels outside the 5th and 99th intervals to remove outliers due to noise. In this study, we used <italic>l</italic><sub><italic>max</italic></sub> &#x0003D; 4 to compute SH coefficients as it has previously been demonstrated to provide comparable performance in CNN-based CSD model coefficient regression as <italic>l</italic><sub><italic>max</italic></sub> &#x0003D; 8 (<xref ref-type="bibr" rid="B46">46</xref>).</p>
</sec>
<sec>
<title>2.4. CNN Architecture</title>
<p>The CNN architecture used is the 3D U-Net (<xref ref-type="bibr" rid="B22">22</xref>) implementation provided with the nnU-Net framework presented in Isensee et al. (<xref ref-type="bibr" rid="B47">47</xref>). The nnU-Net implementation has four downsampling blocks in the encoder pathway and four upsampling blocks in the decoder pathway. Each downsampling block is comprised of 2 &#x000D7; (Convolution, Dropout, InstanceNorm, LeakyRelu) &#x0002B; pooling layer. The upsampling block has a similar structure but the pooling layer is replaced by an upsampling layer.</p>
</sec>
<sec>
<title>2.5. Data Augmentation</title>
<p>Classical techniques for on-the-fly augmentation including axis flipping, scaling, and rotation have been successfully applied to training DL models for small 3D medical imaging datasets (<xref ref-type="bibr" rid="B10">10</xref>, <xref ref-type="bibr" rid="B48">48</xref>, <xref ref-type="bibr" rid="B49">49</xref>). However, traditional medical image processing tools apply these augmentations to the 3D spatial domain which are inappropriate to apply to the SH domain. Therefore, to account for the SH coefficient properties, we apply the same random 3D rotation to both 3D spatial location and SH coefficients in order to ensure location and orientation are preserved during data augmentation as in Nath et al. (<xref ref-type="bibr" rid="B50">50</xref>).</p>
</sec>
<sec>
<title>2.6. CNN Training</title>
<p>For a given training dataset, let &#x1D54F; &#x0003D; [<italic><bold>X</bold></italic><sub><bold>1</bold></sub>, &#x02026;, <italic><bold>X</bold></italic><sub><bold>&#x003C4;</bold></sub>] be the input images mapped to SH coefficients of order <italic>l</italic><sub><italic>max</italic></sub> &#x0003D; 4 and &#x1D550; &#x0003D; [<italic><bold>Y</bold></italic><sub><bold>1</bold></sub>, &#x02026;, <italic><bold>Y</bold></italic><bold><sub>&#x003C4;</sub></bold>] is the corresponding ground truth tract masks, where &#x003C4; is the number of subjects. For a given pair of image <italic><bold>X</bold></italic><bold><sub>&#x003C4;</sub></bold> &#x0003D; [<italic><bold>x</bold></italic><bold><sub>1</sub></bold>, &#x02026;, <italic><bold>x</bold></italic><bold><sub><italic>J</italic></sub></bold>], and mask <italic><bold>Y</bold></italic><bold><sub>&#x003C4;</sub></bold> &#x0003D; [<italic><bold>y</bold></italic><bold><sub>1</sub></bold>, &#x02026;, <italic><bold>y</bold></italic><bold><sub><italic>J</italic></sub>]</bold>, <italic>J</italic> is the number of voxels and <italic><bold>x</bold></italic><bold><sub><italic>j</italic></sub></bold> &#x0003D; [<italic>x</italic><sub><italic>j</italic>1</sub>, &#x02026;, <italic>x</italic><sub><italic>jM</italic></sub>], where <italic>M</italic> is the number of SH coefficients and <italic><bold>y</bold></italic><bold><sub><italic>j</italic></sub></bold> &#x0003D; [<italic>y</italic><sub><italic>j</italic>1</sub>, &#x02026;, <italic>y</italic><sub><italic>jN</italic></sub>], where <italic>N</italic> is the number of classes (tracts to be predicted). The training stage consists of optimizing the CNN model <italic>f</italic><sub>&#x003B8;</sub>(&#x000B7;) to minimize a mapping as follows:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mo class="qopname">arg</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">min</mml:mo></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>&#x003B8;</mml:mi></mml:mstyle></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>&#x1D54F;</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x1D550;</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic><bold>&#x003B8;</bold></italic> are the learned weights. At the inference stage, for an image <italic><bold>X</bold></italic><bold><sub>&#x003C4;</sub></bold>, we compute the predicted probabilities as <inline-formula><mml:math id="M2"><mml:mstyle mathvariant="bold-italic"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>Y</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mstyle><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:msub><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<sec>
<title>2.6.1. Loss Function</title>
<p>We used weighted binary cross-entropy (wBCE) loss during our CNN training stage. We calculate the distribution of each class over all subjects as <italic><bold>C</bold></italic> &#x0003D; [<italic>c</italic><sub>1</sub>, &#x02026;, <italic>c</italic><sub><italic>n</italic></sub>] where <inline-formula><mml:math id="M3"><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>J</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the number of positive labels for a given ground truth tract mask for a class <italic>n</italic> in the training set. A class weight <inline-formula><mml:math id="M4"><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>C</mml:mi></mml:mstyle></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></inline-formula> is used to preferentially optimize wBCE for classes with small numbers of positive labels. For a given predicted probability <inline-formula><mml:math id="M5"><mml:mstyle mathvariant="bold-italic"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mrow><mml:mi>Y</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003C4;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mstyle></mml:math></inline-formula> and ground truth tract masks <italic><bold>Y</bold></italic><bold><sub>&#x003C4;</sub></bold>, we compute wBCE as:</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M6"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mi>w</mml:mi><mml:mi>B</mml:mi><mml:mi>C</mml:mi><mml:mi>E</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mover accent='true'><mml:mrow><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>Y</mml:mi></mml:mstyle><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x003C4;</mml:mi></mml:mstyle></mml:msub></mml:mrow><mml:mo stretchy='true'>&#x0005E;</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>Y</mml:mi></mml:mstyle><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>&#x003C4;</mml:mi></mml:mstyle></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:mstyle><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>J</mml:mi></mml:munderover><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>y</mml:mi></mml:mstyle><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>j</mml:mi><mml:mi>n</mml:mi></mml:mstyle></mml:mrow></mml:msub><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mover accent='true'><mml:mi>y</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover><mml:mrow><mml:mi>j</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mstyle></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>+</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>y</mml:mi></mml:mstyle><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>j</mml:mi><mml:mi>n</mml:mi></mml:mstyle></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mover accent='true'><mml:mi>y</mml:mi><mml:mo>&#x0005E;</mml:mo></mml:mover></mml:mstyle><mml:mrow><mml:mstyle mathvariant='bold' mathsize='normal'><mml:mi>j</mml:mi><mml:mi>n</mml:mi></mml:mstyle></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
</sec>
<sec>
<title>2.6.2. Training Setup</title>
<p>The 3D U-Net is initialized using the He uniform function (<xref ref-type="bibr" rid="B51">51</xref>) and is trained for 400 epochs, with a weight decay of 1<italic>E</italic> &#x02212; 6, and a dropout of <italic>r</italic> &#x0003D; 0.25 in the encoder branch, based on experimentally chosen convergence. The learning rate is initialized to 1<italic>E</italic> &#x02212; 3 and is reduced by 1/2 every 50 epochs. For each iteration in an epoch, a subject from the training set is randomly selected and 3D rotations (Section 2.5) are applied to augment the data in the range of [&#x02212;20, 20]. From this image, 50 patches of size 64 &#x000D7; 64 &#x000D7; 64 &#x000D7; 15 are randomly sampled from within a binary mask corresponding to the intracranial space, where 15 is the number of SH coefficients. The number of patches was experimentally selected to achieve optimal convergence on the validation set while having all patches loaded on the available graphics processing units (GPUs). An epoch finishes when all subjects from the training set have been selected once. For every new subject, a new set of random rotations and patches are computed.</p>
</sec>
</sec>
<sec>
<title>2.7. Uncertainty Quantification</title>
<p>Uncertainty can be divided into two types: epistemic (model&#x00027;s parameters&#x00027; variability) or aleatoric (noise inherent in the data) (<xref ref-type="bibr" rid="B33">33</xref>). A well-established method for estimating epistemic uncertainty in deep learning is test time dropout (TTD) (<xref ref-type="bibr" rid="B34">34</xref>). In TTD, dropout layers are enabled at the inference stage to output multiple stochastic predictions. For aleatoric uncertainty, test time augmentation (TTA) is implemented by performing multiple data augmentations to the input data at the inference stage to output multiple stochastic predictions (<xref ref-type="bibr" rid="B32">32</xref>).</p>
<sec>
<title>2.7.1. Epistemic Uncertainty Modeling</title>
<p>We modeled epistemic uncertainty using TTD as described in Gal et al. (<xref ref-type="bibr" rid="B34">34</xref>). TTD estimates a tractable parametrized distribution <italic>q</italic><sup>&#x0002A;</sup>(<italic><bold>&#x003B8;</bold></italic>) which minimizes the Kullback-Leibler divergence of the true model posterior <italic>p</italic>(<italic><bold>&#x003B8;</bold></italic>|&#x1D54F;, &#x1D550;) (<xref ref-type="bibr" rid="B33">33</xref>). However, <italic>p</italic>(<italic><bold>&#x003B8;</bold></italic>|&#x1D54F;, &#x1D550;) is often not directly computable. In practice, without any further change in the model during the training, dropout layers that switch off neurons activation at a rate of <italic>r</italic>, sampled from a Bernoulli distribution, can be used at the inference stage to approximate probability distribution for the model weights <italic><bold>&#x003B8;</bold></italic>. This approximation is done by computing <italic>T</italic> forward passes where random model weights are set to 0 for each iteration. The final prediction is calculated by averaging the predicted probabilities from the <italic>T</italic> forward passes. The epistemic uncertainty is computed as the SD of the predicted probabilities from the <italic>T</italic> forward passes. In this study, we define a total of <italic>T</italic> &#x0003D; 20 forward passes and a dropout rate <italic>r</italic> &#x0003D; 0.25.</p>
</sec>
<sec>
<title>2.7.2. Aleatoric Uncertainty Modeling</title>
<p>We modeled aleatoric uncertainty using TTA. This technique combines the predicted probabilities of multiple augmentation transforms at the inference stage to generate a final output to take into account noise inherent to the input data. TTA is common practice in classification problems in computer vision (<xref ref-type="bibr" rid="B52">52</xref>) and has also been applied in medical imaging for segmentation (<xref ref-type="bibr" rid="B32">32</xref>, <xref ref-type="bibr" rid="B53">53</xref>). For TTA, the same data augmentation techniques as presented in Section 2.5 were used to augment the input data. Similarly to TTD, we define <italic>T</italic> &#x0003D; 20 forward passes, each with a random data augmentation.</p>
</sec>
<sec>
<title>2.7.3. Aleatoric and Epistemic Uncertainty Modeling</title>
<p>Similarly to Wang et al. (<xref ref-type="bibr" rid="B32">32</xref>), we compute a Hybrid approach to compute both epistemic and aleatoric uncertainty using TTD and TTA, respectively. We keep <italic>T</italic> &#x0003D; 20 forward passes, where for each pass, we have a random data augmentation and a dropout rate <italic>r</italic> &#x0003D; 0.25. This gives a total of 20 predicted probabilities for each subject.</p>
</sec>
</sec>
<sec>
<title>2.8. Calibration</title>
<p>We perform post-processing volumetric calibration as presented in Eaton-Rosen et al. (<xref ref-type="bibr" rid="B24">24</xref>). For a given subject <italic><bold>x</bold></italic><bold><sub><italic>j</italic></sub></bold>, after running <italic>T</italic> stochastic forward passes (TTD, TTA, or Hybrid), we have <italic>T</italic> predicted probabilities per voxel and per class <italic><bold>&#x01EF9;</bold></italic><sub><italic>jn</italic></sub> &#x0003D; [&#x01EF9;<sub><italic>jn</italic>1</sub>, &#x02026;, &#x01EF9;<sub><italic>jnT</italic></sub>] for <italic>n</italic> &#x02208; [1, &#x02026;, <italic>N</italic>] classes. We compute predicted probabilities quantiles &#x003C9;<sub><italic>kj</italic></sub> for each voxel given the <italic>T</italic> output predicted probabilities <italic><bold>&#x01EF9;</bold></italic><sub><italic>jn</italic></sub> for <inline-formula><mml:math id="M8"><mml:mi>k</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mo>&#x02026;</mml:mo><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:mfrac></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>. Then, we compute the volume <inline-formula><mml:math id="M9"><mml:msub><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>J</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mrow><mml:mi>&#x003C9;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> for each quantile <italic>k</italic>th. Next, the cumulative distribution function (CDF) <italic>F</italic>(<italic>v</italic>) &#x0003D; <italic>P</italic>(<italic>V</italic><sub><italic>k</italic></sub> &#x0003C; <italic>v</italic>) is computed where <italic>V</italic><sub><italic>k</italic></sub> is the kth quantile volume.</p>
<p>The calibration step is performed by fitting <italic>F</italic>(<italic>v</italic>) to a cumulative uniform distribution using a 1D linear interpolation. The scaling parameters for the linear interpolation are computed using a one-shot approach. This interpolation realigns <italic>F</italic>(<italic>v</italic>) so that the correct proportion of ground truth volumes appears in a given confidence interval (<xref ref-type="bibr" rid="B24">24</xref>). This calibration is performed per tract since each tract varies in shape, size, and uncertainty. We calculate all scaling parameters on the validation set, to prevent contamination with the test set, and subsequently use these trained parameters to calibrate predicted probabilities on the test set.</p>
</sec>
<sec>
<title>2.9. Evaluation Metrics</title>
<p>We evaluate the quality of predicted tract segmentation using the following metrics: sensitivity, specificity, Dice, Hausdorff distance, and average surface-to-surface distance (ASSD). Sensitivity is true positives over all predicted positive voxels, specificity is true negatives over all predictive negative voxels and Dice is the intersection of the predicted and ground truth masks over two times the union. The sensitivity, specificity, and Dice metrics are overlap metrics (larger numbers are best, the maximum value is 1.0). The Hausdorff distance measures the maximum of the directed distances between the boundaries of the predicted and ground truth segmentations, while ASSD is the average of all distances from points on the boundary of the predicted segmentation to the boundary of the ground truth segmentation and the boundary of the ground truth segmentation to boundary of the predicted segmentation (<xref ref-type="bibr" rid="B54">54</xref>). ASSD and Hausdorff distance measures often indicate if outliers are present in the predicted segmentation (<xref ref-type="bibr" rid="B55">55</xref>). The Hausdorff distance and ASSD are distance-based metrics (smaller numbers are best, the minimum value is 0.0).</p>
</sec>
<sec>
<title>2.10. Experiments</title>
<p>We assess the following tract segmentation approaches, deterministic (U-Net) and stochastic (TTD, TTA, and Hybrid), in the following scenarios: (1) how well the deterministic and stochastic approaches perform tract segmentation (<italic>Segmentation performance and comparison to state-of-the-art</italic>), (2) how well the deterministic and Hybrid approaches perform on clinical quality data (<italic>Segmentation performance on clinical quality data</italic>), (3) how well uncertainty maps computed by the stochastic approaches correlate with tract segmentation error (<italic>Correlation between uncertainty and segmentation error</italic>), and (4) how well volume-based calibration adjusts predicted probabilities from TTD, TTA, and Hybrid approaches (<italic>Calibration impact on predicted tract volume</italic>). The details of each experiment are described below.</p>
<sec>
<title>2.10.1. Segmentation Performance and Comparison to State-of-the-Art</title>
<p>We assess how well deterministic and stochastic tract segmentation approaches perform in terms of the evaluation metrics described in Section 2.9. In this experiment, 5-fold cross-validation is conducted where 4 folds are used for training (10% of training data was used for validation) and 1 fold for testing. We also compare the results of our approaches against state-of-the-art approaches, including TractSeg (<xref ref-type="bibr" rid="B10">10</xref>) (CNN-based method), Classifyber (<xref ref-type="bibr" rid="B8">8</xref>) (machine learning-based method), and RecoBundles (<xref ref-type="bibr" rid="B56">56</xref>) (a clustering-based method for segmentation) in terms of average Dice.</p>
</sec>
<sec>
<title>2.10.2. Segmentation Performance on Clinical Quality Data</title>
<p>We assess the robustness of the deterministic and hybrid approaches perform on a clinical quality dataset. From the original HCP data, we first select the b = 1,000 s/mm<sup>2</sup> shell to mimic a shell commonly acquired in a clinical protocol. Then, similarly to Lucena et al. (<xref ref-type="bibr" rid="B46">46</xref>), for each dataset, we first reorder the set of gradient directions such that if a scan is truncated the acquired gradient directions will still be close to optimally distributed on the half-sphere (<xref ref-type="bibr" rid="B45">45</xref>), and we then synthetically generate a clinical quality dMRI scan by truncating the number of gradient directions for b = 0 and b = 1,000 s/mm<sup>2</sup> to 45 gradient directions. Finally, we apply the method described in this article (Section 2).</p>
</sec>
<sec>
<title>2.10.3. Correlation Between Uncertainty and Segmentation Error</title>
<p>We assess how well uncertainty quantification correlates with tract segmentation errors. We use structure-wise uncertainty measured by the volume variation coefficient (VVC) and correlation this to segmentation errors as measured by 1 - average Dice as presented by Wang et al. (<xref ref-type="bibr" rid="B32">32</xref>). We compute Dice for each of <italic>T</italic> forward passes and then compute the average of the <italic>T</italic> Dice scores (output predict probabilities images are thresholded at &#x02265;0.5 to obtain a binary segmentation). We compute tract volume for each forward pass, <italic>V</italic> &#x0003D; [<italic>v</italic><sub>1</sub>, ...<italic>v</italic><sub><italic>T</italic></sub>] where <italic>v</italic><sub><italic>t</italic></sub> is the total sum over all voxels on the binary image and <italic>t</italic> &#x02208; [0, &#x02026;, <italic>T</italic>]. VVC is then computed as <italic>VVC</italic> &#x0003D; &#x003C3;<sub><italic>V</italic></sub>/&#x003BC;<sub><italic>V</italic></sub> where &#x003BC;<sub><italic>V</italic></sub> and &#x003C3;<sub><italic>V</italic></sub> are the mean and SD for all volumes in <italic>V</italic>, respectively. We compute the strength of the VVC and 1 - Dice correlation using Spearman&#x00027;s rank correlation coefficient (<xref ref-type="bibr" rid="B57">57</xref>). Spearmans&#x00027; correlation assesses monotonic relationships (whether linear or not). If there are no repeated data values, a perfect Spearmans&#x00027; correlation of 1 or &#x02212;1 occurs when each of the variables is a perfect monotone function of the other (<xref ref-type="bibr" rid="B57">57</xref>).</p>
</sec>
<sec>
<title>2.10.4. Calibration Impact on Predicted Tract Volume</title>
<p>We assess how well volume-calibrated stochastic methods (TTA, TTD, and Hybrid) correspond to ground truth volumes. In this experiment, we evaluate the correlation of predicted volumes at different quantiles obtained over <italic>T</italic> forward passes, with and without calibration, to the ground truth volumes for individual tract structures.</p>
</sec>
</sec>
<sec>
<title>2.11. Implementation</title>
<p>All experiments were performed on a workstation equipped with an Intel CPU (Xeon&#x000AE;W-2123, 8 &#x000D7; 3.60 GHz; Intel), 32 GB of memory, and an NVIDIA GPU (GeForce Titan V) with 12 GB of on-board memory. All code was implemented in Python 3.6. PyTorch 1.6.0 (<xref ref-type="bibr" rid="B58">58</xref>) and PyTorch lightning (<xref ref-type="bibr" rid="B59">59</xref>) were used for network training. MONAI 0.5.2 and TorchIO 0.18.15 (<xref ref-type="bibr" rid="B60">60</xref>) were used for data loading and sampling. Data augmentation was performed using SHtools 4.6.2 (<xref ref-type="bibr" rid="B61">61</xref>). All code used for training the models is available online as an open-source project<xref ref-type="fn" rid="fn0002"><sup>2</sup></xref>.</p>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>3. Results</title>
<p>Overall quantitative analyses are reported for all 72 tracts (<xref ref-type="table" rid="T1">Tables 1</xref>, <bold>4</bold>) while other detailed results for specific qualitative and quantitative are presented for a small number of select tracts (<xref ref-type="table" rid="T2">Table 2</xref> and <xref ref-type="fig" rid="F1">Figures 1</xref>&#x02013;<bold>7</bold>). A complete list of all tracts can be found online<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref>. For these cases, we chose to report the representative tracts: corticospinal tract (CST), inferior longitudinal fascicle (ILF), and uncinate fascicle tract (UF) for the left side of the brain. CST is a large, well-represented tract with a straight shape that fans out close to the cortex. ILF is a complex longitudinal tract that starts from the anterior side and goes to the posterior side of the brain. Finally, the UF is a complex tract that has a large &#x0201C;C&#x0201D; shaped curvature.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Dice, sensitivity, specificity, ASSD, and Hausdorff distance evaluation for deterministic (U-Net) and stochastic (TTD, TTA, and Hybrid) approaches.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Network<xref ref-type="table-fn" rid="TN1"><sup>a</sup></xref></bold></th>
<th valign="top" align="center"><bold>Dice</bold></th>
<th valign="top" align="center"><bold>Sensitivity</bold></th>
<th valign="top" align="center"><bold>Specificity</bold></th>
<th valign="top" align="center"><bold>ASSD (mm)</bold></th>
<th valign="top" align="center"><bold>Hausdorff distance (mm)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">TractSeg<xref ref-type="table-fn" rid="TN1"><sup>a</sup></xref></td>
<td valign="top" align="center"><bold>0.84</bold></td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left">U-Net</td>
<td valign="top" align="center">0.83 (0.06)</td>
<td valign="top" align="center">0.81 (0.09)</td>
<td valign="top" align="center"><bold>0.85 (0.07)</bold></td>
<td valign="top" align="center">0.69 (0.63)</td>
<td valign="top" align="center">17.32 (21.66)</td>
</tr>
<tr>
<td valign="top" align="left">TTD</td>
<td valign="top" align="center">0.82 (0.06)</td>
<td valign="top" align="center">0.85 (0.08)</td>
<td valign="top" align="center">0.80 (0.08)</td>
<td valign="top" align="center">0.65(0.28)</td>
<td valign="top" align="center">10.57 (9.47)</td>
</tr>
<tr>
<td valign="top" align="left">TTA</td>
<td valign="top" align="center">0.82 (0.07)</td>
<td valign="top" align="center">0.85 (0.08)</td>
<td valign="top" align="center">0.80 (0.09)</td>
<td valign="top" align="center"><bold>0.63 (0.30)</bold></td>
<td valign="top" align="center"><bold>9.24 (3.73)</bold></td>
</tr>
<tr>
<td valign="top" align="left">Hybrid</td>
<td valign="top" align="center">0.82 (0.07)</td>
<td valign="top" align="center"><bold>0.86 (0.08)</bold></td>
<td valign="top" align="center">0.78 (0.09)</td>
<td valign="top" align="center">0.66 (0.33)</td>
<td valign="top" align="center">9.46 (3.74)</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="TN1"><label>a</label><p><italic>TractSeg results are taken from Wasserthal et al. (<xref ref-type="bibr" rid="B10">10</xref>). The best value, the minimum value for ASSD and Hausdorff distance, and the maximum value for Dice, sensitivity, and specificity, are indicated by bold text</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Comparison with state-of-the-art approaches for the following 5 tracts from both left and right sides of the brain: arcuate fascicle (AF), corticospinal tract (CST), inferior fronto-occipital fascicle (IFO), inferior longitudinal fascicle (ILF), and uncinate fascicle (UF).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Tract</bold></th>
<th valign="top" align="center"><bold>RecoBundles<xref ref-type="table-fn" rid="TN2"><sup><bold>a</bold></sup></xref></bold></th>
<th valign="top" align="center"><bold>TractSeg<xref ref-type="table-fn" rid="TN3"><sup><bold>b</bold></sup></xref></bold></th>
<th valign="top" align="center"><bold>Classifyber<xref ref-type="table-fn" rid="TN4"><sup><bold>c</bold></sup></xref></bold></th>
<th valign="top" align="center"><bold>U-Net</bold></th>
<th valign="top" align="center"><bold>TTD</bold></th>
<th valign="top" align="center"><bold>TTA</bold></th>
<th valign="top" align="center"><bold>Hybrid</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Right CST</td>
<td valign="top" align="center">0.62</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">0.84 (0.03)</td>
<td valign="top" align="center">0.84 (0.03)</td>
<td valign="top" align="center">0.84 (0.03)</td>
<td valign="top" align="center">0.84 (0.03)</td>
</tr>
<tr>
<td valign="top" align="left">Left CST</td>
<td valign="top" align="center">0.62</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.85 (0.03)</td>
<td valign="top" align="center">0.85 (0.02)</td>
<td valign="top" align="center">0.85 (0.02)</td>
<td valign="top" align="center">0.84 (0.02)</td>
</tr>
<tr>
<td valign="top" align="left">Right UF</td>
<td valign="top" align="center">0.57</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.77 (0.04)</td>
<td valign="top" align="center">0.78 (0.03)</td>
<td valign="top" align="center">0.78 (0.03)</td>
<td valign="top" align="center">0.78 (0.03)</td>
</tr>
<tr>
<td valign="top" align="left">Left UF</td>
<td valign="top" align="center">0.55</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.75 (0.07)</td>
<td valign="top" align="center">0.75 (0.06)</td>
<td valign="top" align="center">0.76 (0.06)</td>
<td valign="top" align="center">0.75 (0.06)</td>
</tr>
<tr>
<td valign="top" align="left">Right AF</td>
<td valign="top" align="center">0.53</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.83 (0.03)</td>
<td valign="top" align="center">0.83 (0.02)</td>
<td valign="top" align="center">0.84 (0.02)</td>
<td valign="top" align="center">0.83 (0.02)</td>
</tr>
<tr>
<td valign="top" align="left">Left AF</td>
<td valign="top" align="center">0.71</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.84 (0.03)</td>
<td valign="top" align="center">0.84 (0.02)</td>
<td valign="top" align="center">0.85 (0.02)</td>
<td valign="top" align="center">0.84 (0.02)</td>
</tr>
<tr>
<td valign="top" align="left">Right ILF</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.80 (0.03)</td>
<td valign="top" align="center">0.79 (0.02)</td>
<td valign="top" align="center">0.81 (0.02)</td>
<td valign="top" align="center">0.80 (0.02)</td>
</tr>
<tr>
<td valign="top" align="left">Left ILF</td>
<td valign="top" align="center">0.57</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.80 (0.03)</td>
<td valign="top" align="center">0.79 (0.02)</td>
<td valign="top" align="center">0.80 (0.03)</td>
<td valign="top" align="center">0.79 (0.02)</td>
</tr>
<tr>
<td valign="top" align="left">Right IFO</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.80 (0.04)</td>
<td valign="top" align="center">0.80 (0.03)</td>
<td valign="top" align="center">0.80 (0.04)</td>
<td valign="top" align="center">0.79 (0.03)</td>
</tr>
<tr>
<td valign="top" align="left">Left IFO</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.78 (0.04)</td>
<td valign="top" align="center">0.78 (0.03)</td>
<td valign="top" align="center">0.78 (0.03)</td>
<td valign="top" align="center">0.78 (0.03)</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="TN2"><label>a</label><p><italic>Garyfallidis et al. (<xref ref-type="bibr" rid="B56">56</xref>)</italic>.</p></fn>
<fn id="TN3">
<label>b</label>
<p><italic>Wasserthal et al. (<xref ref-type="bibr" rid="B10">10</xref>)</italic>.</p></fn>
<fn id="TN4">
<label>c</label>
<p><italic>Bert&#x000F2; et al. (<xref ref-type="bibr" rid="B8">8</xref>)</italic>.</p></fn>
<p><italic>Results for TractSeg, ReconBundles, and Classifyber are reported in Bert&#x000F2; et al. (<xref ref-type="bibr" rid="B8">8</xref>)</italic>.</p>
</table-wrap-foot>
</table-wrap>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Tract segmentation comparisons between deterministic and stochastic approaches for the left CST, ILF, and UF tracts. Red contours show ground truth segmentations.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fradi-02-866974-g0001.tif"/>
</fig>
<sec>
<title>3.1. Segmentation Performance and Comparison to State-of-the-Art</title>
<p><xref ref-type="table" rid="T1">Table 1</xref> reports the mean (standard deviation) for the metrics described in Section 2.9 for both deterministic (U-Net) and stochastic (TTA, TTD, Hybrid) approaches. Similar performance in terms of Dice is found between U-Net and TTD, TTA, and Hybrid (average Dice &#x02248;0.82). Stochastic approaches are more sensitive than the deterministic approach but less specific, which can be explained due to the presence of fewer false negative predictions (segmenting a voxel that belongs to a specific tract as background) by the TTD, TTA, and Hybrid approaches.</p>
<p>Pronounced improvements in Hausdorff distance are found for stochastic approaches compared to the deterministic approach (for TTA, it is a &#x02248;50% improvement, 9 mm difference). This demonstrates that stochastic approaches are less likely to make large mistakes when segmenting a tract. TTD and TTA are comparable in terms of Dice performance to TractSeg which currently has the best achieving results for multi-label tract segmentation.</p>
<p><xref ref-type="table" rid="T2">Table 2</xref> compares the Dice performance between our approaches and four state-of-the-art tract segmentation methods. Both deterministic and stochastic approaches have comparable performance to TractSeg as expected from the results in <xref ref-type="table" rid="T1">Table 1</xref>. RecoBundles, an ROI-based method, has the lowest Dice performance while Classifyber, an LR fiber-based classification approach has the highest Dice performance. However, differences in Dice are relatively small (1&#x02013;7%) between methods. ReconBundles is known to demonstrate poor reproducibility in tracts with high anatomical variability across subjects (<xref ref-type="bibr" rid="B16">16</xref>) while Classifyber relies on LR where each tract is treated as a separated classification task and does not face the challenges of multi-label classification.</p>
<p><xref ref-type="fig" rid="F1">Figure 1</xref> shows qualitative segmentations for CST, ILF, and UF on the left side of the brain. Both deterministic and stochastic approaches provide tract segmentation that has similar shapes and sizes compared to the ground truth tract masks. For the UF, there is larger anatomical variability and fewer &#x0201C;spurious&#x0201D; regions in the inner part of the tract resulting in a cleaner &#x0201C;C&#x0201D; shape, when compared to the ground truth. This can be explained due to CNN&#x00027;s learning an average pattern across different subjects during the training stage leading to smoother results.</p>
</sec>
<sec>
<title>3.2. Segmentation Performance on Clinical Quality Data</title>
<p><xref ref-type="table" rid="T3">Table 3</xref> reports the mean (SD) for the metrics described in Section 2.9 for both U-Net and Hybrid approaches evaluated on the clinical data. As expected, due to the lower quality of the clinical data and different acquisition parameters, we observe a drop in performance on clinical data (45 gradient directions, b = 1,000 s/mm<sup>2</sup>) for both U-Net and Hybrid approaches when compared to the original data (90 gradient directions, b = 2,000 s/mm<sup>2</sup>). For the Hybrid approach, we observe greater uncertainty in the boundary regions resulting in under segmentation of the tract (<xref ref-type="fig" rid="F2">Figure 2</xref>). This trend is observed across all tracts, with a drop in performance of 4%. These results highlight the importance of computing uncertainty for lower quality data in preoperative planning.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Dice, sensitivity, specificity, ASSD, and Hausdorff distance evaluation for deterministic (U-Net) and stochastic Hybrid approaches for clinical quality data.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Network</bold></th>
<th valign="top" align="center"><bold>Dice</bold></th>
<th valign="top" align="center"><bold>Sensitivity</bold></th>
<th valign="top" align="center"><bold>Specificity</bold></th>
<th valign="top" align="center"><bold>ASSD (mm)</bold></th>
<th valign="top" align="center"><bold>Hausdorff distance (mm)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">U-Net</td>
<td valign="top" align="center">0.82 (0.06)</td>
<td valign="top" align="center">0.83 (0.09)</td>
<td valign="top" align="center">0.81 (0.07)</td>
<td valign="top" align="center">0.68 (0.55)</td>
<td valign="top" align="center">16.18 (20.84)</td>
</tr>
<tr>
<td valign="top" align="left">Hybrid</td>
<td valign="top" align="center">0.78 (0.08)</td>
<td valign="top" align="center">0.89 (0.08)</td>
<td valign="top" align="center">0.72 (0.11)</td>
<td valign="top" align="center">0.81 (0.44)</td>
<td valign="top" align="center">10.20 (3.90)</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Difference between Hybrid the SDs of the predicted probabilities (Uncertainties) computed on original and clinical data (top). A selection of Dice distribution for 13 select tracts is computed across all 105 subjects (bottom).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fradi-02-866974-g0002.tif"/>
</fig>
</sec>
<sec>
<title>3.3. Correlation Between Uncertainty and Segmentation Error</title>
<p><xref ref-type="fig" rid="F3">Figure 3</xref> plots VVC (tract uncertainty) vs. 1-Dice (segmentation error). A high correlation between VCC and 1-Dice is an indication of segmentation reliability, whereas uncertainty is indicative of poor segmentation performance (i.e., higher error). Therefore, a positive correlation should be expected if the uncertainty computed is a good measure of segmentation reliability.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Structure-wise uncertainty as measured by volume variation coefficient (VVC) vs. 1 - Dice for TTD, TTA, and Hybrid approaches in the left CST, ILF, and UF tracts. Spearman&#x00027;s correlation coefficient indicates the strength of the correlation between VVC and 1 - Dice.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fradi-02-866974-g0003.tif"/>
</fig>
<p>For the tracts segmented by the TTD approach, the correlation between structure-wise uncertainty and segmentation error is absent (ILF) or weak (CST, UF). This indicates that model uncertainty is an inadequate measure of segmentation reliability. The TTA approach had stronger correlations between structure-wise uncertainty and segmented error. This demonstrates that aleatoric uncertainty provides a good measure of segmentation reliability in this task. For the tracts segmented by the Hybrid approach, the strongest correlations between structure-wise uncertainty and segmentation error are observed, suggesting that including both model and data uncertainty is more beneficial to estimating the reliability of the segmentation than either measure individually. For these specific tracts, for all stochastic approaches, UF and ILF exhibit higher slopes compared to CST which can be explained due to these structures having more complex tract anatomy.</p>
<p><xref ref-type="fig" rid="F4">Figures 4</xref>&#x02013;<bold>6</bold> show 2D reconstructions of the residuals (<italic><bold>y</bold></italic> &#x02212; <italic><bold>&#x00177;</bold></italic>) for U-Net, TTD, TTA, and Hybrid approaches and uncertainty maps output by the stochastic approaches. For all tracts, residuals tend to be at boundary voxels (regions more likely to be mistaken for other tracts) and these regions are also associated with higher uncertainty (TTD, TTA, and Hybrid only).</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Uncertainty and residuals maps for the left CST tract. Red arrows point to one area of high uncertainty within the ground truth. This region may not represent real CST anatomy. Uncertainty maps are not applicable (N/A) for the deterministic U-Net approach.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fradi-02-866974-g0004.tif"/>
</fig>
<p>For the CST (<xref ref-type="fig" rid="F4">Figure 4</xref>), stochastic approaches have larger residuals in areas that may not be part of the tract (small protuberance on the right side of the CST with high uncertainty). TTD demonstrates high uncertainty for areas with high residuals values whereas TTA and Hybrid approaches identify uncertainty in more dispersed regions, although these approaches still have high uncertainty near boundary regions. This may occur due to variability in the tract shape caused by multiple augmentations.</p>
<p>Similar to CST (<xref ref-type="fig" rid="F5">Figure 5</xref>), for the ILF, TTD uncertainty is highest at boundary voxels while TTA and Hybrid also output high uncertainty for regions inside the tract. For this case, the Hybrid approach outputs high uncertainty in many areas inside the tract which may be due to the complex tract structure. For the UF, a tract with a very high curvature that varies between subjects, TTA residuals output high values inside the tract similarly to the Hybrid approach (<xref ref-type="fig" rid="F6">Figure 6</xref>).</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Uncertainty and residual maps for the left ILF tract. Uncertainty maps are not applicable (N/A) for the deterministic U-Net approach.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fradi-02-866974-g0005.tif"/>
</fig>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Uncertainty and residual maps for the left UF tract. Uncertainty maps are not applicable (N/A) for the deterministic U-Net approach.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fradi-02-866974-g0006.tif"/>
</fig>
</sec>
<sec>
<title>3.4. Calibration Impact on Predicted Tract Volume</title>
<p>We evaluated how well volume-based calibration for the stochastic approaches impacts the reliability of the predicted probabilities. <xref ref-type="table" rid="T4">Table 4</xref> reports the mean (SD) for all metrics described in Section 2.9 for uncalibrated and calibrated approaches. As expected, no pronounced difference is found within the segmentation metrics between the uncalibrated and calibrated approaches.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Segmentation metrics for 21 random subjects in the test set.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Network</bold></th>
<th valign="top" align="center"><bold>Dice</bold></th>
<th valign="top" align="center"><bold>Sensitivity</bold></th>
<th valign="top" align="center"><bold>Specificity</bold></th>
<th valign="top" align="center"><bold>ASSD (mm)</bold></th>
<th valign="top" align="center"><bold>Hausdorff distance (mm)</bold></th>
<th valign="top" align="center"><bold>Calibration</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">TTD</td>
<td valign="top" align="center">0.81 (0.07)</td>
<td valign="top" align="center">0.84 (0.08)</td>
<td valign="top" align="center"><bold>0.81 (0.09)</bold></td>
<td valign="top" align="center">0.67 (0.31)</td>
<td valign="top" align="center">10.75 (9.82)</td>
<td valign="top" align="center">NO</td>
</tr>
<tr>
<td valign="top" align="left">TTD</td>
<td valign="top" align="center">0.81 (0.07)</td>
<td valign="top" align="center">0.84 (0.08)</td>
<td valign="top" align="center"><bold>0.81 (0.09)</bold></td>
<td valign="top" align="center">0.67 (0.32)</td>
<td valign="top" align="center">10.78 (10.04)</td>
<td valign="top" align="center">YES</td>
</tr>
<tr>
<td valign="top" align="left">TTA</td>
<td valign="top" align="center">0.82 (0.08)</td>
<td valign="top" align="center">0.84 (0.09)</td>
<td valign="top" align="center">0.80 (0.10)</td>
<td valign="top" align="center">0.67 (0.38)</td>
<td valign="top" align="center">9.49(4.01)</td>
<td valign="top" align="center">NO</td>
</tr>
<tr>
<td valign="top" align="left">TTA</td>
<td valign="top" align="center"><bold>0.83 (0.05)</bold></td>
<td valign="top" align="center">0.84 (0.08)</td>
<td valign="top" align="center">0.81 (0.10)</td>
<td valign="top" align="center"><bold>0.67 (0.22)</bold></td>
<td valign="top" align="center"><bold>9.48 (4.04)</bold></td>
<td valign="top" align="center">YES</td>
</tr>
<tr>
<td valign="top" align="left">Hybrid</td>
<td valign="top" align="center">0.81 (0.08)</td>
<td valign="top" align="center"><bold>0.86 (0.08)</bold></td>
<td valign="top" align="center">0.79 (0.11)</td>
<td valign="top" align="center">0.69 (0.4)</td>
<td valign="top" align="center">9.65 (3.86)</td>
<td valign="top" align="center">NO</td>
</tr>
<tr>
<td valign="top" align="left">Hybrid</td>
<td valign="top" align="center">0.81 (0.08)</td>
<td valign="top" align="center"><bold>0.86 (0.08)</bold></td>
<td valign="top" align="center">0.78 (0.11)</td>
<td valign="top" align="center">0.70 (0.42)</td>
<td valign="top" align="center">9.70 (3.90)</td>
<td valign="top" align="center">YES</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>In this table, TTD, TTD, and Hybrid results are computed from uncalibrated and calibrated predicted probabilities. The best value, minimum value for ASSD and Hausdorff distance, and maximum value for Dice, sensitivity, and specificity, are indicated by bold text</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>Volume-based calibration makes the distribution of predicted volumes more uniform to consequently make predicted probabilities more representative of the ground truth volume. Calibrated predicted volumes (orange dots) tend to encompass the ground truth segmentation volume (black bars) compared to uncalibrated predicted volumes (blue dots) (<xref ref-type="fig" rid="F7">Figure 7</xref>).</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Uncalibrated and calibrated predicted volumes for TTD, TTA, and Hybrid approaches. Volume deviation is computed as the predicted volume subtracted by the ground truth volume. For each subject, the ground truth (black horizontal line) is at zero, and the predicted volume deviation is more visually apparent.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fradi-02-866974-g0007.tif"/>
</fig>
<p>For all 105 subjects, we compute the minimum volume difference between the ground truth volume and the predicted volume for the Hybrid approach before and after calibration (<xref ref-type="table" rid="T5">Table 5</xref>). The average minimum volume difference between ground truth volume and calibrated predicted volumes is lower than the average minimum volume difference between ground truth volumes and uncalibrated predicted volumes. These results indicate that calibration makes the distribution of predicted volumes more representative of the ground truth volumes for the dataset.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>The minimum volume difference between ground truth volume and the predicted volume for the Hybrid method before and after calibration.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Tracts</bold></th>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Minimum volume difference</bold></th>
</tr>
<tr>
<th/>
<th valign="top" align="center"><bold>Calibrated</bold></th>
<th valign="top" align="center"><bold>Uncalibrated</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Left CST</td>
<td valign="top" align="center">347.83 (444.96)</td>
<td valign="top" align="center">90.82 (50.87)</td>
</tr>
<tr>
<td valign="top" align="left">Left ILF</td>
<td valign="top" align="center">755.93 (570.75)</td>
<td valign="top" align="center">137.33 (162.31)</td>
</tr>
<tr>
<td valign="top" align="left">Left UF</td>
<td valign="top" align="center">579.18 (607.17)</td>
<td valign="top" align="center">100.22 (84.91)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>4. Discussion</title>
<p>We evaluated techniques for uncertainty quantification to provide more accurate and reliable predicted probabilities segmentation outputs applied to DL-based tract segmentation. We show quantitatively (<xref ref-type="table" rid="T1">Tables 1</xref>, <xref ref-type="table" rid="T2">2</xref> and <xref ref-type="fig" rid="F1">Figure 1</xref>) that stochastic approaches with uncertainty awareness have comparable performance to state-of-the-art methods. Additionally, these uncertainty measures have a positive correlation with tract segmentation errors, which indicates uncertainty is a good measure of reliability for tract segmentation (<xref ref-type="fig" rid="F4">Figures 4</xref>&#x02013;<xref ref-type="fig" rid="F6">6</xref>). Finally, we demonstrate that calibrated predicted probabilities are more representative of the ground truth volume compared to uncalibrated predicted probabilities (<xref ref-type="fig" rid="F7">Figure 7</xref>).</p>
<p>Similar studies have proposed DL-based tract segmentation. TractSeg (<xref ref-type="bibr" rid="B10">10</xref>) is a voxel-based approach introduced for multi-label segmentation using 2D U-Nets to segment 72 tracts. As input to the network, TractSeg uses the 3 major peak directions computed from FODs based on CSD (<xref ref-type="bibr" rid="B21">21</xref>). TractSeg has an average Dice of 0.84 on 105 subjects from the HCP dataset which is currently the best multi-label tract segmentation performance achieved. Although the authors use a tri-planar approach, the use of 2D CNNs cannot leverage spatial context between slices (<xref ref-type="bibr" rid="B62">62</xref>). Additionally, selecting only the 3 major peaks of the FODs computed using CSD may discard important information contained within the dMRI data.</p>
<p>Neuro4Neuro is another voxel-based approach (<xref ref-type="bibr" rid="B17">17</xref>). A 3D U-Net is used to segment 25 tracts with an average Dice of 0.75. As input for the 3D CNN, this approach uses DTI. Although the authors used a large cohort to train their models (&#x0003E;1,000 scans), they segment a small number of tracts. This approach obtained lower Dice for the task compared to TractSeg. One explanation is that DTI only models single fiber populations and cannot resolve complex fiber configurations such as fiber crossings (<xref ref-type="bibr" rid="B63">63</xref>), resulting in &#x0201C;poor&#x0201D; segmentation for complex structures (i.e., inferior longitudinal fasciculus) (<xref ref-type="bibr" rid="B17">17</xref>). Direct comparisons between Neuro4Neuro and our work were not possible since the code was not publicly available, and their test data was an in-house dataset.</p>
<p>DeepWMA (<xref ref-type="bibr" rid="B18">18</xref>) is a fiber-based approach that uses a 2D multi-channel fiber feature descriptor to describe fibers obtained from whole-brain tractography. DeepWMA uses a 2D CNN to classify individual fibers into one of 54 possible WM tracts. This approach generalizes well for scans of independently acquired populations. This method reports a performance comparable to TractSeg, average Dice 0.83, for 34 tracts on the HCP dataset (<xref ref-type="bibr" rid="B18">18</xref>). However, for a new given patient, DeepWMA requires preprocessing whole-brain tractography which is a time-consuming and computationally expensive step.</p>
<p>Classifyber is another fiber-based approach (<xref ref-type="bibr" rid="B8">8</xref>). Classifyber uses an LR classification model to predict whether individual streamlines belong to a tract of interest. Similar to DeepWMA, Classifyber also has a descriptor to represent a streamline based on a set of features (i.e., spatial position, connectivity, etc). Although the method provides high Dice (&#x02265; 0.80 per tract), Classifyber relies on LR where each tract is treated as a separated classification task and does not address the challenges of multi-label classification. As with DeepWMA, tractography is a preprocessing step that is time-consuming and computationally expensive.</p>
<p>In this study, we used TTD and TTA to model epistemic and aleatoric uncertainty, respectively. We evaluated the combination of both in a Hybrid approach. The aim of this approach is to provide additional information about model and data reliability to help inform clinicians&#x00027; decision making. The TTA approach has the best performance in terms of Dice for all stochastic approaches, however, the Hybrid approach provides a stronger correlation between structure-wise uncertainty (VVC) and segmentation error (1-Dice), indicating it is a good measure of segmentation reliability (<xref ref-type="fig" rid="F3">Figure 3</xref>). These results are observed more strongly in tracts with complex anatomy that are difficult to segment such as ILF (<xref ref-type="fig" rid="F5">Figures 5</xref>, <xref ref-type="fig" rid="F6">6</xref>).</p>
<p>We used single-rater ground truth annotations for this study. However, multi-rater annotations could improve the quality of segmentation results (less bias toward a single rater) but would introduce inter-rater variance. In this context, we would expect an increase in uncertainty for areas with high variability among the raters and high certainty in areas with high concordance among the raters.</p>
<p>Calibration was performed per tract using a volume-based approach. Calibrated volume estimates are more likely to encompass the ground truth volume than the uncalibrated volume estimates (<xref ref-type="fig" rid="F7">Figure 7</xref>). Calibration can help to reduce false negatives by allowing the user to select a volume from a larger quantile to err on the side of caution, e.g., during surgery/planning, stochastic approaches can include the segmentation regions with a low likelihood of belonging to the tract to ensure no potential damage will occur even if it is a low probability event. However, calibration is sensitive to the size of the training set (<xref ref-type="bibr" rid="B24">24</xref>) and the quality of the ground truth, meaning that tracts with complex anatomy and inter-subject variability, such as the UF tract, may be difficult to calibrate. Additionally, volume as a metric for calibration may not be sufficient for tract segmentation, and other metrics such as a topology-based assessment to describe tract segmentation coverage should be investigated.</p>
<p>There are two key limitations in this study. First, we validated our methods on research data acquired using a single protocol. For clinical data with different acquisition protocols, DL-based methods can output &#x0201C;poor&#x0201D; segmentation and high epistemic uncertainty due to domain shift (<xref ref-type="bibr" rid="B64">64</xref>). While domain adaption (<xref ref-type="bibr" rid="B65">65</xref>) can overcome low model performance and high uncertainty, we have not investigated this in the current study. Second, we did not validate subjects with pathologies that would distort WM tissue connectivity, which can result in unusual tract shape and location that might lower the Dice of our proposed method. One future avenue of research is to evaluate our approach to subjects with pathologies that distort normal anatomy, such as brain tumors, in a clinical setting.</p>
</sec>
<sec sec-type="conclusions" id="s5">
<title>5. Conclusion</title>
<p>In this study, we presented uncertainty awareness for tract segmentation with accurate and reliable predicted probabilities so that clinicians can use it as a safety tool in preoperative neurosurgical planning. Our stochastic approaches, TTD, TTA, and Hybrid, achieved performance comparable to the state-of-the-art methods while outputting measures of uncertainty. We demonstrated a strong positive correlation between segmentation error and structure-wise uncertainty for our stochastic approaches indicating that our output uncertainties are a good measure of reliability for tract segmentation. We confirmed the importance of volume-based calibration in tract segmentation showing an improved ability to measure tract volumes in complex structures compared to uncalibrated approaches. However, other metrics that describe tracts topology could improve calibration results but require further investigation. We focused our analysis on healthy subjects from the HCP dataset. Future validation is required to demonstrate our approach generalizes to datasets acquired at clinical sites and on patients with brain pathologies that distort normal anatomies, such as edemas or tumors.</p>
</sec>
<sec sec-type="data-availability" id="s6">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article material. The Human Connectome Project (HCP) data is a publicly available dataset (<ext-link ext-link-type="uri" xlink:href="https://db.humanconnectome.org">https://db.humanconnectome.org</ext-link>) where we used the 105 subjects. Further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>OL and RS: conceptualization. OL: software, data curation, validation, formal analysis, investigation, writing&#x02014;original draft, and visualization. PB and OL: methodology. PB, JC, SO, KA, and RS: writing&#x02014;review and editing. RS: resources. SO: funding acquisition. KA, RS, and SO: supervision and project administration. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>This research was funded by the National Institute for Health Research (NIHR) Biomedical Research Centre based at Guy&#x00027;s and St Thomas&#x00027; NHS Foundation Trust and King&#x00027;s College London and the NIHR Clinical Research Facility. OL was funded by EPSRC Research Council (EPSRC DTP EP/R513064/1). PB was funded by the Wellcome Flagship Programme (WT213038/Z/18/Z) and Wellcome EPSRC CME (WT203148/Z/16/Z).</p>
</sec>
<sec id="s9">
<title>Author Disclaimer</title>
<p>The views expressed are those of the author(s) and not necessarily those of the NHS, the NIHR or the Department of Health.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ack><p>We thank NVIDIA for providing the Titan V GPU used in this work.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<label>1.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Essayed</surname> <given-names>WI</given-names></name> <name><surname>Zhang</surname> <given-names>F</given-names></name> <name><surname>Unadkat</surname> <given-names>P</given-names></name> <name><surname>Cosgrove</surname> <given-names>GR</given-names></name> <name><surname>Golby</surname> <given-names>AJ</given-names></name> <name><surname>O&#x00027;Donnell</surname> <given-names>LJ</given-names></name></person-group>. <article-title>White matter tractography for neurosurgical planning: a topography-based review of the current state of the art</article-title>. <source>Neuroimage Clin</source>. (<year>2017</year>) <volume>15</volume>:<fpage>659</fpage>&#x02013;<lpage>72</lpage>. <pub-id pub-id-type="doi">10.1016/j.nicl.2017.06.011</pub-id><pub-id pub-id-type="pmid">28664037</pub-id></citation></ref>
<ref id="B2">
<label>2.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mancini</surname> <given-names>M</given-names></name> <name><surname>Vos</surname> <given-names>SB</given-names></name> <name><surname>Vakharia</surname> <given-names>VN</given-names></name> <name><surname>O&#x00027;Keeffe</surname> <given-names>AG</given-names></name> <name><surname>Trimmel</surname> <given-names>K</given-names></name> <name><surname>Barkhof</surname> <given-names>F</given-names></name> <etal/></person-group>. <article-title>Automated fiber tract reconstruction for surgery planning: extensive validation in language-related white matter tracts</article-title>. <source>Neuroimage Clin</source>. (<year>2019</year>) <volume>23</volume>:<fpage>101883</fpage>. <pub-id pub-id-type="doi">10.1016/j.nicl.2019.101883</pub-id><pub-id pub-id-type="pmid">31163386</pub-id></citation></ref>
<ref id="B3">
<label>3.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Calabrese</surname> <given-names>E</given-names></name></person-group>. <article-title>Diffusion tractography in deep brain stimulation surgery: a review</article-title>. <source>Front Neuroanat</source>. (<year>2016</year>) <volume>10</volume>:<fpage>45</fpage>. <pub-id pub-id-type="doi">10.3389/fnana.2016.00045</pub-id><pub-id pub-id-type="pmid">27199677</pub-id></citation></ref>
<ref id="B4">
<label>4.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schilling</surname> <given-names>KG</given-names></name> <name><surname>Tax</surname> <given-names>CM</given-names></name> <name><surname>Rheault</surname> <given-names>F</given-names></name> <name><surname>Hansen</surname> <given-names>C</given-names></name> <name><surname>Yang</surname> <given-names>Q</given-names></name> <name><surname>Yeh</surname> <given-names>FC</given-names></name> <etal/></person-group>. <article-title>Fiber tractography bundle segmentation depends on scanner effects, vendor effects, acquisition resolution, diffusion sampling scheme, diffusion sensitization, and bundle segmentation workflow</article-title>. <source>Neuroimage</source>. (<year>2021</year>) <volume>242</volume>:<fpage>118451</fpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2021.118451</pub-id><pub-id pub-id-type="pmid">34358660</pub-id></citation></ref>
<ref id="B5">
<label>5.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schilling</surname> <given-names>KG</given-names></name> <name><surname>Daducci</surname> <given-names>A</given-names></name> <name><surname>Maier-Hein</surname> <given-names>K</given-names></name> <name><surname>Poupon</surname> <given-names>C</given-names></name> <name><surname>Houde</surname> <given-names>JC</given-names></name> <name><surname>Nath</surname> <given-names>V</given-names></name> <etal/></person-group>. <article-title>Challenges in diffusion MRI tractography-Lessons learned from international benchmark competitions</article-title>. <source>Magn Reson Imaging</source>. (<year>2018</year>) <pub-id pub-id-type="doi">10.1016/j.mri.2018.11.014</pub-id><pub-id pub-id-type="pmid">30503948</pub-id></citation></ref>
<ref id="B6">
<label>6.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Andreisek</surname> <given-names>G</given-names></name> <name><surname>White</surname> <given-names>LM</given-names></name> <name><surname>Kassner</surname> <given-names>A</given-names></name> <name><surname>Sussman</surname> <given-names>MS</given-names></name></person-group>. <article-title>Evaluation of diffusion tensor imaging and fiber tractography of the median nerve: preliminary results on intrasubject variability and precision of measurements</article-title>. <source>Am J Roentgenol</source>. (<year>2010</year>) <volume>194</volume>:<fpage>W65</fpage>&#x02013;<lpage>72</lpage>. <pub-id pub-id-type="doi">10.2214/AJR.09.2517</pub-id><pub-id pub-id-type="pmid">20028893</pub-id></citation></ref>
<ref id="B7">
<label>7.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>De Schotten</surname> <given-names>MT</given-names></name> <name><surname>Bizzi</surname> <given-names>A</given-names></name> <name><surname>Dell&#x00027;Acqua</surname> <given-names>F</given-names></name> <name><surname>Allin</surname> <given-names>M</given-names></name> <name><surname>Walshe</surname> <given-names>M</given-names></name> <name><surname>Murray</surname> <given-names>R</given-names></name> <etal/></person-group>. <article-title>Atlasing location, asymmetry and inter-subject variability of white matter tracts in the human brain with MR diffusion tractography</article-title>. <source>Neuroimage</source>. (<year>2011</year>) <volume>54</volume>:<fpage>49</fpage>&#x02013;<lpage>59</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2010.07.055</pub-id><pub-id pub-id-type="pmid">20682348</pub-id></citation></ref>
<ref id="B8">
<label>8.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bert&#x000F2;</surname> <given-names>G</given-names></name> <name><surname>Bullock</surname> <given-names>D</given-names></name> <name><surname>Astolfi</surname> <given-names>P</given-names></name> <name><surname>Hayashi</surname> <given-names>S</given-names></name> <name><surname>Zigiotto</surname> <given-names>L</given-names></name> <name><surname>Annicchiarico</surname> <given-names>L</given-names></name> <etal/></person-group>. <article-title>Classifyber, a robust streamline-based linear classifier for white matter bundle segmentation</article-title>. <source>Neuroimage</source>. (<year>2021</year>) <volume>224</volume>:<fpage>117402</fpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2020.117402</pub-id><pub-id pub-id-type="pmid">32979520</pub-id></citation></ref>
<ref id="B9">
<label>9.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sydnor</surname> <given-names>VJ</given-names></name> <name><surname>Rivas-Grajales</surname> <given-names>AM</given-names></name> <name><surname>Lyall</surname> <given-names>AE</given-names></name> <name><surname>Zhang</surname> <given-names>F</given-names></name> <name><surname>Bouix</surname> <given-names>S</given-names></name> <name><surname>Karmacharya</surname> <given-names>S</given-names></name> <etal/></person-group>. <article-title>A comparison of three fiber tract delineation methods and their impact on white matter analysis</article-title>. <source>Neuroimage</source>. (<year>2018</year>) <volume>178</volume>:<fpage>318</fpage>&#x02013;<lpage>331</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2018.05.044</pub-id><pub-id pub-id-type="pmid">29787865</pub-id></citation></ref>
<ref id="B10">
<label>10.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wasserthal</surname> <given-names>J</given-names></name> <name><surname>Neher</surname> <given-names>P</given-names></name> <name><surname>Maier-Hein</surname> <given-names>KH</given-names></name></person-group>. <article-title>Tractseg-fast and accurate white matter tract segmentation</article-title>. <source>Neuroimage</source>. (<year>2018</year>) <volume>183</volume>:<fpage>239</fpage>&#x02013;<lpage>53</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2018.07.070</pub-id><pub-id pub-id-type="pmid">30086412</pub-id></citation></ref>
<ref id="B11">
<label>11.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yeh</surname> <given-names>CH</given-names></name> <name><surname>Smith</surname> <given-names>RE</given-names></name> <name><surname>Dhollander</surname> <given-names>T</given-names></name> <name><surname>Calamante</surname> <given-names>F</given-names></name> <name><surname>Connelly</surname> <given-names>A</given-names></name></person-group>. <article-title>Connectomes from streamlines tractography: assigning streamlines to brain parcellations is not trivial but highly consequential</article-title>. <source>Neuroimage</source>. (<year>2019</year>) <volume>199</volume>:<fpage>160</fpage>&#x02013;<lpage>71</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2019.05.005</pub-id><pub-id pub-id-type="pmid">31082471</pub-id></citation></ref>
<ref id="B12">
<label>12.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wassermann</surname> <given-names>D</given-names></name> <name><surname>Makris</surname> <given-names>N</given-names></name> <name><surname>Rathi</surname> <given-names>Y</given-names></name> <name><surname>Shenton</surname> <given-names>M</given-names></name> <name><surname>Kikinis</surname> <given-names>R</given-names></name> <name><surname>Kubicki</surname> <given-names>M</given-names></name> <etal/></person-group>. <article-title>The white matter query language: a novel approach for describing human white matter anatomy</article-title>. <source>Brain Struct Funct</source>. (<year>2016</year>) <volume>221</volume>:<fpage>4705</fpage>&#x02013;<lpage>21</lpage>. <pub-id pub-id-type="doi">10.1007/s00429-015-1179-4</pub-id><pub-id pub-id-type="pmid">26754839</pub-id></citation></ref>
<ref id="B13">
<label>13.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Siless</surname> <given-names>V</given-names></name> <name><surname>Chang</surname> <given-names>K</given-names></name> <name><surname>Fischl</surname> <given-names>B</given-names></name> <name><surname>Yendiki</surname> <given-names>A</given-names></name></person-group>. <article-title>AnatomiCuts: hierarchical clustering of tractography streamlines based on anatomical similarity</article-title>. <source>Neuroimage</source>. (<year>2018</year>) <volume>166</volume>:<fpage>32</fpage>&#x02013;<lpage>45</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2017.10.058</pub-id><pub-id pub-id-type="pmid">29100937</pub-id></citation></ref>
<ref id="B14">
<label>14.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Garyfallidis</surname> <given-names>E</given-names></name> <name><surname>C&#x000F4;t&#x000E9;</surname> <given-names>MA</given-names></name> <name><surname>Rheault</surname> <given-names>F</given-names></name> <name><surname>Sidhu</surname> <given-names>J</given-names></name> <name><surname>Hau</surname> <given-names>J</given-names></name> <name><surname>Petit</surname> <given-names>L</given-names></name> <etal/></person-group>. <article-title>Recognition of white matter bundles using local and global streamline-based registration and clustering</article-title>. <source>Neuroimage</source>. (<year>2018</year>) <volume>170</volume>:<fpage>283</fpage>&#x02013;<lpage>95</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2017.07.015</pub-id><pub-id pub-id-type="pmid">28712994</pub-id></citation></ref>
<ref id="B15">
<label>15.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>O&#x00027;Donnell</surname> <given-names>LJ</given-names></name> <name><surname>Westin</surname> <given-names>CF</given-names></name></person-group>. <article-title>Automatic tractography segmentation using a high-dimensional white matter atlas</article-title>. <source>IEEE Trans Med Imaging</source>. (<year>2007</year>) <volume>26</volume>:<fpage>1562</fpage>&#x02013;<lpage>75</lpage>. <pub-id pub-id-type="doi">10.1109/TMI.2007.906785</pub-id><pub-id pub-id-type="pmid">18041271</pub-id></citation></ref>
<ref id="B16">
<label>16.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wakana</surname> <given-names>S</given-names></name> <name><surname>Caprihan</surname> <given-names>A</given-names></name> <name><surname>Panzenboeck</surname> <given-names>MM</given-names></name> <name><surname>Fallon</surname> <given-names>JH</given-names></name> <name><surname>Perry</surname> <given-names>M</given-names></name> <name><surname>Gollub</surname> <given-names>RL</given-names></name> <etal/></person-group>. <article-title>Reproducibility of quantitative tractography methods applied to cerebral white matter</article-title>. <source>Neuroimage</source>. (<year>2007</year>) <volume>36</volume>:<fpage>630</fpage>&#x02013;<lpage>44</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2007.02.049</pub-id><pub-id pub-id-type="pmid">17481925</pub-id></citation></ref>
<ref id="B17">
<label>17.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>B</given-names></name> <name><surname>De Groot</surname> <given-names>M</given-names></name> <name><surname>Steketee</surname> <given-names>RM</given-names></name> <name><surname>Meijboom</surname> <given-names>R</given-names></name> <name><surname>Smits</surname> <given-names>M</given-names></name> <name><surname>Vernooij</surname> <given-names>MW</given-names></name> <etal/></person-group>. <article-title>Neuro4Neuro: a neural network approach for neural tract segmentation using large-scale population-based diffusion imaging</article-title>. <source>Neuroimage</source>. (<year>2020</year>) <volume>218</volume>:<fpage>116993</fpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2020.116993</pub-id><pub-id pub-id-type="pmid">32492510</pub-id></citation></ref>
<ref id="B18">
<label>18.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>F</given-names></name> <name><surname>Karayumak</surname> <given-names>SC</given-names></name> <name><surname>Hoffmann</surname> <given-names>N</given-names></name> <name><surname>Rathi</surname> <given-names>Y</given-names></name> <name><surname>Golby</surname> <given-names>AJ</given-names></name> <name><surname>O&#x00027;Donnell</surname> <given-names>LJ</given-names></name></person-group>. <article-title>Deep white matter analysis (DeepWMA): fast and consistent tractography segmentation</article-title>. <source>Med Image Anal</source>. (<year>2020</year>) <volume>65</volume>:<fpage>101761</fpage>. <pub-id pub-id-type="doi">10.1016/j.media.2020.101761</pub-id><pub-id pub-id-type="pmid">32622304</pub-id></citation></ref>
<ref id="B19">
<label>19.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeCun</surname> <given-names>Y</given-names></name> <name><surname>Bengio</surname> <given-names>Y</given-names></name> <name><surname>Hinton</surname> <given-names>G</given-names></name></person-group>. <article-title>Deep learning</article-title>. <source>Nature</source>. (<year>2015</year>) <volume>521</volume>:<fpage>436</fpage>&#x02013;<lpage>444</lpage>. <pub-id pub-id-type="doi">10.1038/nature14539</pub-id><pub-id pub-id-type="pmid">26017442</pub-id></citation></ref>
<ref id="B20">
<label>20.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ronneberger</surname> <given-names>O</given-names></name> <name><surname>Fischer</surname> <given-names>P</given-names></name> <name><surname>Brox</surname> <given-names>T</given-names></name></person-group>. <article-title>U-net: convolutional networks for biomedical image segmentation</article-title>. In: <source>International Conference on Medical Image Computing and Computer-Assisted Intervention</source>. <publisher-loc>Springer</publisher-loc> (<year>2015</year>). p. <fpage>234</fpage>&#x02013;<lpage>41</lpage>.</citation>
</ref>
<ref id="B21">
<label>21.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jeurissen</surname> <given-names>B</given-names></name> <name><surname>Tournier</surname> <given-names>JD</given-names></name> <name><surname>Dhollander</surname> <given-names>T</given-names></name> <name><surname>Connelly</surname> <given-names>A</given-names></name> <name><surname>Sijbers</surname> <given-names>J</given-names></name></person-group>. <article-title>Multi-tissue constrained spherical deconvolution for improved analysis of multi-shell diffusion MRI data</article-title>. <source>Neuroimage</source>. (<year>2014</year>) <volume>103</volume>:<fpage>411</fpage>&#x02013;<lpage>26</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2014.07.061</pub-id><pub-id pub-id-type="pmid">25109526</pub-id></citation></ref>
<ref id="B22">
<label>22.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>&#x000C7;i&#x000E7;ek</surname> <given-names>&#x000D6;</given-names></name> <name><surname>Abdulkadir</surname> <given-names>A</given-names></name> <name><surname>Lienkamp</surname> <given-names>SS</given-names></name> <name><surname>Brox</surname> <given-names>T</given-names></name> <name><surname>Ronneberger</surname> <given-names>O</given-names></name></person-group>. <article-title>3D U-Net: learning dense volumetric segmentation from sparse annotation</article-title>. In: <source>International Conference on Medical Image Computing and Computer-Assisted Intervention</source>. <publisher-loc>Springer</publisher-loc> (<year>2016</year>). p. <fpage>424</fpage>&#x02013;<lpage>32</lpage>.</citation>
</ref>
<ref id="B23">
<label>23.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Basser</surname> <given-names>PJ</given-names></name> <name><surname>Mattiello</surname> <given-names>J</given-names></name> <name><surname>LeBihan</surname> <given-names>D</given-names></name></person-group>. <article-title>MR diffusion tensor spectroscopy and imaging</article-title>. <source>Biophys J</source>. (<year>1994</year>) <volume>66</volume>:<fpage>259</fpage>&#x02013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.1016/S0006-3495(94)80775-1</pub-id><pub-id pub-id-type="pmid">8130344</pub-id></citation></ref>
<ref id="B24">
<label>24.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Eaton-Rosen</surname> <given-names>Z</given-names></name> <name><surname>Bragman</surname> <given-names>F</given-names></name> <name><surname>Bisdas</surname> <given-names>S</given-names></name> <name><surname>Ourselin</surname> <given-names>S</given-names></name> <name><surname>Cardoso</surname> <given-names>MJ</given-names></name></person-group>. <article-title>Towards safe deep learning: accurately quantifying biomarker uncertainty in neural network predictions</article-title>. In: <source>International Conference on Medical Image Computing and Computer-Assisted Intervention</source>. <publisher-loc>Springer</publisher-loc> (<year>2018</year>). p. <fpage>691</fpage>&#x02013;<lpage>9</lpage>.</citation>
</ref>
<ref id="B25">
<label>25.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Borges</surname> <given-names>P</given-names></name> <name><surname>Shaw</surname> <given-names>R</given-names></name> <name><surname>Varsavsky</surname> <given-names>T</given-names></name> <name><surname>Klaser</surname> <given-names>K</given-names></name> <name><surname>Thomas</surname> <given-names>D</given-names></name> <name><surname>Drobnjak</surname> <given-names>I</given-names></name> <etal/></person-group>. <article-title>Acquisition-invariant brain mri segmentation with informative uncertainties</article-title>. <source>arXiv [Preprint]</source>. (<year>2021</year>). arXiv: 2111.04094478. <pub-id pub-id-type="doi">10.48550/ARXIV.2111.04094</pub-id></citation>
</ref>
<ref id="B26">
<label>26.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kl&#x000E4;ser</surname> <given-names>K</given-names></name> <name><surname>Varsavsky</surname> <given-names>T</given-names></name> <name><surname>Markiewicz</surname> <given-names>P</given-names></name> <name><surname>Vercauteren</surname> <given-names>T</given-names></name> <name><surname>Atkinson</surname> <given-names>D</given-names></name> <name><surname>Thielemans</surname> <given-names>K</given-names></name> <etal/></person-group>. <article-title>Improved MR to CT synthesis for PET/MR attenuation correction using Imitation Learning</article-title>. In: <source>International Workshop on Simulation and Synthesis in Medical Imaging</source>. <publisher-loc>Springer</publisher-loc> (<year>2019</year>). p. <fpage>13</fpage>&#x02013;<lpage>21</lpage>.</citation>
</ref>
<ref id="B27">
<label>27.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tanno</surname> <given-names>R</given-names></name> <name><surname>Worrall</surname> <given-names>DE</given-names></name> <name><surname>Kaden</surname> <given-names>E</given-names></name> <name><surname>Ghosh</surname> <given-names>A</given-names></name> <name><surname>Grussu</surname> <given-names>F</given-names></name> <name><surname>Bizzi</surname> <given-names>A</given-names></name> <etal/></person-group>. <article-title>Uncertainty modelling in deep learning for safer neuroimage enhancement: demonstration in diffusion MRI</article-title>. <source>Neuroimage</source>. (<year>2021</year>) <volume>225</volume>:<fpage>117366</fpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2020.117366</pub-id><pub-id pub-id-type="pmid">33039617</pub-id></citation></ref>
<ref id="B28">
<label>28.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Graham</surname> <given-names>MS</given-names></name> <name><surname>Sudre</surname> <given-names>CH</given-names></name> <name><surname>Varsavsky</surname> <given-names>T</given-names></name> <name><surname>Tudosiu</surname> <given-names>PD</given-names></name> <name><surname>Nachev</surname> <given-names>P</given-names></name> <name><surname>Ourselin</surname> <given-names>S</given-names></name> <etal/></person-group>. <article-title>Hierarchical brain parcellation with uncertainty</article-title>. In: <source>Uncertainty for Safe Utilization of Machine Learning in Medical Imaging, and Graphs in Biomedical Image Analysis</source>. <publisher-loc>Springer</publisher-loc> (<year>2020</year>). p. <fpage>23</fpage>&#x02013;<lpage>31</lpage>.</citation>
</ref>
<ref id="B29">
<label>29.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Granados</surname> <given-names>A</given-names></name> <name><surname>Lucena</surname> <given-names>O</given-names></name> <name><surname>Vakharia</surname> <given-names>V</given-names></name> <name><surname>Miserocchi</surname> <given-names>A</given-names></name> <name><surname>McEvoy</surname> <given-names>AW</given-names></name> <name><surname>Vos</surname> <given-names>SB</given-names></name> <etal/></person-group>. <article-title>Towards uncertainty quantification for electrode bending prediction in stereotactic neurosurgery</article-title>. In: <source>2020 IEEE 17th International Symposium on Biomedical Imaging (ISBI)</source>. <publisher-loc>Iowa City, IA</publisher-loc>: <publisher-name>IEEE</publisher-name> (<year>2020</year>). p. <fpage>674</fpage>&#x02013;<lpage>7</lpage>.</citation>
</ref>
<ref id="B30">
<label>30.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Granados</surname> <given-names>A</given-names></name> <name><surname>Han</surname> <given-names>Y</given-names></name> <name><surname>Lucena</surname> <given-names>O</given-names></name> <name><surname>Vakharia</surname> <given-names>V</given-names></name> <name><surname>Rodionov</surname> <given-names>R</given-names></name> <name><surname>Vos</surname> <given-names>SB</given-names></name> <etal/></person-group>. <article-title>Patient-specific prediction of SEEG electrode bending for stereotactic neurosurgical planning</article-title>. <source>Int J Comput Assist Radiol Surg</source>. (<year>2021</year>) <volume>16</volume>:<fpage>789</fpage>&#x02013;<lpage>798</lpage>. <pub-id pub-id-type="doi">10.1007/s11548-021-02347-8</pub-id><pub-id pub-id-type="pmid">33761063</pub-id></citation></ref>
<ref id="B31">
<label>31.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jungo</surname> <given-names>A</given-names></name> <name><surname>Balsiger</surname> <given-names>F</given-names></name> <name><surname>Reyes</surname> <given-names>M</given-names></name></person-group>. <article-title>Analyzing the quality and challenges of uncertainty estimations for brain tumor segmentation</article-title>. <source>Front Neurosci</source>. (<year>2020</year>) <volume>14</volume>:<fpage>282</fpage>. <pub-id pub-id-type="doi">10.3389/fnins.2020.00282</pub-id><pub-id pub-id-type="pmid">32322186</pub-id></citation></ref>
<ref id="B32">
<label>32.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>G</given-names></name> <name><surname>Li</surname> <given-names>W</given-names></name> <name><surname>Aertsen</surname> <given-names>M</given-names></name> <name><surname>Deprest</surname> <given-names>J</given-names></name> <name><surname>Ourselin</surname> <given-names>S</given-names></name> <name><surname>Vercauteren</surname> <given-names>T</given-names></name></person-group>. <article-title>Aleatoric uncertainty estimation with test-time augmentation for medical image segmentation with convolutional neural networks</article-title>. <source>Neurocomputing</source>. (<year>2019</year>) <volume>338</volume>:<fpage>34</fpage>&#x02013;<lpage>45</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2019.01.103</pub-id><pub-id pub-id-type="pmid">31595105</pub-id></citation></ref>
<ref id="B33">
<label>33.</label>
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Kendal</surname> <given-names>A</given-names></name> <name><surname>Gal</surname> <given-names>Y</given-names></name></person-group>. <article-title>What uncertainties do we need in bayesian deep learning for computer vision?</article-title> <source>Adv Neural Inf Process Syst.</source> (<year>2017</year>) <fpage>30</fpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/pdf/1703.04977.pdf">https://arxiv.org/pdf/1703.04977.pdf</ext-link> <pub-id pub-id-type="pmid">32405271</pub-id></citation></ref>
<ref id="B34">
<label>34.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gal</surname> <given-names>Y</given-names></name> <name><surname>Ghahramani</surname> <given-names>Z</given-names></name></person-group>. <article-title>Dropout as a bayesian approximation: representing model uncertainty in deep learning</article-title>. In: <source>International Conference on Machine Learning</source>. <publisher-loc>PMLR</publisher-loc> (<year>2016</year>). p. <fpage>1050</fpage>&#x02013;<lpage>9</lpage>.</citation>
</ref>
<ref id="B35">
<label>35.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Neal</surname> <given-names>RM</given-names></name></person-group>. <source>Bayesian Learning for Neural Networks. Vol. 118</source>. Springer Science &#x00026; Business Media (<year>2012</year>).</citation>
</ref>
<ref id="B36">
<label>36.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Teye</surname> <given-names>M</given-names></name> <name><surname>Azizpour</surname> <given-names>H</given-names></name> <name><surname>Smith</surname> <given-names>K</given-names></name></person-group>. <article-title>Bayesian uncertainty estimation for batch normalized deep networks</article-title>. In: <source>International Conference on Machine Learning</source>. <publisher-loc>PMLR</publisher-loc> (<year>2018</year>). p. <fpage>4907</fpage>&#x02013;<lpage>16</lpage>.</citation>
</ref>
<ref id="B37">
<label>37.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Klaser</surname> <given-names>K</given-names></name> <name><surname>Borges</surname> <given-names>P</given-names></name> <name><surname>Shaw</surname> <given-names>R</given-names></name> <name><surname>Ranzini</surname> <given-names>M</given-names></name> <name><surname>Modat</surname> <given-names>M</given-names></name> <name><surname>Atkinson</surname> <given-names>D</given-names></name> <etal/></person-group>. <article-title>A multi-channel uncertainty-aware multi-resolution network for MR to CT synthesis</article-title>. <source>Appl Sci</source>. (<year>2021</year>) <volume>11</volume>:<fpage>1667</fpage>. <pub-id pub-id-type="doi">10.3390/app11041667</pub-id><pub-id pub-id-type="pmid">33763236</pub-id></citation></ref>
<ref id="B38">
<label>38.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ayhan</surname> <given-names>MS</given-names></name> <name><surname>Berens</surname> <given-names>P</given-names></name></person-group>. <article-title>Test-time data augmentation for estimation of heteroscedastic aleatoric uncertainty in deep neural networks</article-title>. In: <source>Proceedings of Medical Imaging with Deep Learning.</source> (<year>2018</year>).</citation>
</ref>
<ref id="B39">
<label>39.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>C</given-names></name> <name><surname>Pleiss</surname> <given-names>G</given-names></name> <name><surname>Sun</surname> <given-names>Y</given-names></name> <name><surname>Weinberger</surname> <given-names>KQ</given-names></name></person-group>. <article-title>On calibration of modern neural networks</article-title>. In: <source>International Conference on Machine Learning</source>. <publisher-loc>PMLR</publisher-loc> (<year>2017</year>). p. <fpage>1321</fpage>&#x02013;<lpage>30</lpage>.</citation>
</ref>
<ref id="B40">
<label>40.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zadrozny</surname> <given-names>B</given-names></name> <name><surname>Elkan</surname> <given-names>C</given-names></name></person-group>. <article-title>Obtaining calibrated probability estimates from decision trees and naive bayesian classifiers</article-title>. In: <source>ICML, vol. 1</source>. <publisher-loc>Citeseer</publisher-loc> (<year>2001</year>). p. <fpage>609</fpage>&#x02013;<lpage>16</lpage>.</citation>
</ref>
<ref id="B41">
<label>41.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Naeini</surname> <given-names>MP</given-names></name> <name><surname>Cooper</surname> <given-names>G</given-names></name> <name><surname>Hauskrecht</surname> <given-names>M</given-names></name></person-group>. <article-title>Obtaining well calibrated probabilities using bayesian binning</article-title>. In: <source>Twenty-Ninth AAAI Conference on Artificial Intelligence.</source> (<year>2015</year>). p. <fpage>2901</fpage>&#x02013;<lpage>7</lpage>. <pub-id pub-id-type="pmid">25927013</pub-id></citation></ref>
<ref id="B42">
<label>42.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Platt</surname> <given-names>J</given-names></name> <etal/></person-group>. <article-title>Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods</article-title>. <source>Adv Large Margin Classifiers</source>. (<year>1999</year>) <volume>10</volume>:<fpage>61</fpage>&#x02013;<lpage>74</lpage>.</citation>
</ref>
<ref id="B43">
<label>43.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mehrtash</surname> <given-names>A</given-names></name> <name><surname>Wells</surname> <given-names>WM</given-names></name> <name><surname>Tempany</surname> <given-names>CM</given-names></name> <name><surname>Abolmaesumi</surname> <given-names>P</given-names></name> <name><surname>Kapur</surname> <given-names>T</given-names></name></person-group>. <article-title>Confidence calibration and predictive uncertainty estimation for deep medical image segmentation</article-title>. <source>IEEE Trans Med Imaging</source>. (<year>2020</year>) <volume>39</volume>:<fpage>3868</fpage>&#x02013;<lpage>78</lpage>. <pub-id pub-id-type="doi">10.1109/TMI.2020.3006437</pub-id><pub-id pub-id-type="pmid">32746129</pub-id></citation></ref>
<ref id="B44">
<label>44.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sotiropoulos</surname> <given-names>SN</given-names></name> <name><surname>Jbabdi</surname> <given-names>S</given-names></name> <name><surname>Xu</surname> <given-names>J</given-names></name> <name><surname>Andersson</surname> <given-names>JL</given-names></name> <name><surname>Moeller</surname> <given-names>S</given-names></name> <name><surname>Auerbach</surname> <given-names>EJ</given-names></name> <etal/></person-group>. <article-title>Advances in diffusion MRI acquisition and processing in the human connectome project</article-title>. <source>Neuroimage</source>. (<year>2013</year>) <volume>80</volume>:<fpage>125</fpage>&#x02013;<lpage>43</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2013.05.057</pub-id><pub-id pub-id-type="pmid">23702418</pub-id></citation></ref>
<ref id="B45">
<label>45.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tournier</surname> <given-names>J-D</given-names></name> <name><surname>Smith</surname> <given-names>R</given-names></name> <name><surname>Raffelt</surname> <given-names>D</given-names></name> <name><surname>Tabbara</surname> <given-names>R</given-names></name> <name><surname>Dhollander</surname> <given-names>T</given-names></name> <name><surname>Pietsch</surname> <given-names>M</given-names></name> <etal/></person-group>. <article-title>Mrtrix3: A fast, flexible and open software framework for medical image processing and visualisation</article-title>. <source>Neuroimage</source>. (<year>2019</year>) <volume>202</volume>:<fpage>116137</fpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2019.116137</pub-id><pub-id pub-id-type="pmid">31473352</pub-id></citation></ref>
<ref id="B46">
<label>46.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lucena</surname> <given-names>O</given-names></name> <name><surname>Vos</surname> <given-names>SB</given-names></name> <name><surname>Vakharia</surname> <given-names>V</given-names></name> <name><surname>Duncan</surname> <given-names>J</given-names></name> <name><surname>Ashkan</surname> <given-names>K</given-names></name> <name><surname>Sparks</surname> <given-names>R</given-names></name> <etal/></person-group>. <article-title>Enhancing the estimation of fiber orientation distributions using convolutional neural networks</article-title>. <source>Comput Biol Med</source>. (<year>2021</year>) <volume>135</volume>:<fpage>104643</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104643</pub-id><pub-id pub-id-type="pmid">34280774</pub-id></citation></ref>
<ref id="B47">
<label>47.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Isensee</surname> <given-names>F</given-names></name> <name><surname>Jaeger</surname> <given-names>PF</given-names></name> <name><surname>Kohl</surname> <given-names>SA</given-names></name> <name><surname>Petersen</surname> <given-names>J</given-names></name> <name><surname>Maier-Hein</surname> <given-names>KH</given-names></name></person-group>. <article-title>nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation</article-title>. <source>Nat Methods</source>. (<year>2021</year>) <volume>18</volume>:<fpage>203</fpage>&#x02013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1038/s41592-020-01008-z</pub-id><pub-id pub-id-type="pmid">33288961</pub-id></citation></ref>
<ref id="B48">
<label>48.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>W</given-names></name> <name><surname>Wang</surname> <given-names>G</given-names></name> <name><surname>Fidon</surname> <given-names>L</given-names></name> <name><surname>Ourselin</surname> <given-names>S</given-names></name> <name><surname>Cardoso</surname> <given-names>MJ</given-names></name> <name><surname>Vercauteren</surname> <given-names>T</given-names></name></person-group>. <article-title>On the compactness, efficiency, and representation of 3D convolutional networks: brain parcellation as a pretext task</article-title>. In: <source>International Conference on Information Processing in Medical Imaging</source>. <publisher-loc>Springer</publisher-loc> (<year>2017</year>). p. <fpage>348</fpage>&#x02013;<lpage>60</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-59050-9_28</pub-id></citation>
</ref>
<ref id="B49">
<label>49.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gibson</surname> <given-names>E</given-names></name> <name><surname>Li</surname> <given-names>W</given-names></name> <name><surname>Sudre</surname> <given-names>C</given-names></name> <name><surname>Fidon</surname> <given-names>L</given-names></name> <name><surname>Shakir</surname> <given-names>DI</given-names></name> <name><surname>Wang</surname> <given-names>G</given-names></name> <etal/></person-group>. <article-title>NiftyNet: a deep-learning platform for medical imaging</article-title>. <source>Comput Methods Programs Biomed</source>. (<year>2018</year>) <volume>158</volume>:<fpage>113</fpage>&#x02013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.1016/j.cmpb.2018.01.025</pub-id><pub-id pub-id-type="pmid">29544777</pub-id></citation></ref>
<ref id="B50">
<label>50.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nath</surname> <given-names>V</given-names></name> <name><surname>Schilling</surname> <given-names>KG</given-names></name> <name><surname>Parvathaneni</surname> <given-names>P</given-names></name> <name><surname>Hansen</surname> <given-names>CB</given-names></name> <name><surname>Hainline</surname> <given-names>AE</given-names></name> <name><surname>Huo</surname> <given-names>Y</given-names></name> <etal/></person-group>. <article-title>Deep learning reveals untapped information for local white-matter fiber reconstruction in diffusion-weighted MRI</article-title>. <source>Magn Reson Imaging</source>. (<year>2019</year>) <volume>62</volume>:<fpage>220</fpage>&#x02013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1016/j.mri.2019.07.012</pub-id><pub-id pub-id-type="pmid">31323317</pub-id></citation></ref>
<ref id="B51">
<label>51.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K</given-names></name> <name><surname>Zhang</surname> <given-names>X</given-names></name> <name><surname>Ren</surname> <given-names>S</given-names></name> <name><surname>Sun</surname> <given-names>J</given-names></name></person-group>. <article-title>Delving deep into rectifiers: surpassing human-level performance on imagenet classification</article-title>. In: <source>Proceedings of the IEEE International Conference on Computer Vision</source>. <publisher-loc>Santiago</publisher-loc>: <publisher-name>IEEE</publisher-name> (<year>2015</year>). p. <fpage>1026</fpage>&#x02013;<lpage>34</lpage>.</citation>
</ref>
<ref id="B52">
<label>52.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shanmugam</surname> <given-names>D</given-names></name> <name><surname>Blalock</surname> <given-names>D</given-names></name> <name><surname>Balakrishnan</surname> <given-names>G</given-names></name> <name><surname>Guttag</surname> <given-names>J</given-names></name></person-group>. <article-title>Better aggregation in test-time augmentation</article-title>. In: <source>Proceedings of the IEEE/CVF International Conference on Computer Vision.</source> (<year>2021</year>). p. <fpage>1214</fpage>&#x02013;<lpage>23</lpage>.</citation>
</ref>
<ref id="B53">
<label>53.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Amiri</surname> <given-names>M</given-names></name> <name><surname>Brooks</surname> <given-names>R</given-names></name> <name><surname>Behboodi</surname> <given-names>B</given-names></name> <name><surname>Rivaz</surname> <given-names>H</given-names></name></person-group>. <article-title>Two-stage ultrasound image segmentation using U-Net and test time augmentation</article-title>. <source>Int J Comput Assist Radiol Surg</source>. (<year>2020</year>) <volume>15</volume>:<fpage>981</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1007/s11548-020-02158-3</pub-id><pub-id pub-id-type="pmid">32350786</pub-id></citation></ref>
<ref id="B54">
<label>54.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yeghiazaryan</surname> <given-names>V</given-names></name> <name><surname>Voiculescu</surname> <given-names>ID</given-names></name></person-group>. <article-title>Family of boundary overlap metrics for the evaluation of medical image segmentation</article-title>. <source>J Med Imaging</source>. (<year>2018</year>) <volume>5</volume>:<fpage>015006</fpage>. <pub-id pub-id-type="doi">10.1117/1.JMI.5.1.015006</pub-id><pub-id pub-id-type="pmid">29487883</pub-id></citation></ref>
<ref id="B55">
<label>55.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lucena</surname> <given-names>O</given-names></name> <name><surname>Souza</surname> <given-names>R</given-names></name> <name><surname>Rittner</surname> <given-names>L</given-names></name> <name><surname>Frayne</surname> <given-names>R</given-names></name> <name><surname>Lotufo</surname> <given-names>R</given-names></name></person-group>. <article-title>Convolutional neural networks for skull-stripping in brain MR imaging using silver standard masks</article-title>. <source>Artif Intell Med</source>. (<year>2019</year>) <volume>98</volume>:<fpage>48</fpage>&#x02013;<lpage>58</lpage>. <pub-id pub-id-type="doi">10.1016/j.artmed.2019.06.008</pub-id><pub-id pub-id-type="pmid">31521252</pub-id></citation></ref>
<ref id="B56">
<label>56.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Garyfallidis</surname> <given-names>E</given-names></name> <name><surname>Ocegueda</surname> <given-names>O</given-names></name> <name><surname>Wassermann</surname> <given-names>D</given-names></name> <name><surname>Descoteaux</surname> <given-names>M</given-names></name></person-group>. <article-title>Robust and efficient linear registration of white-matter fascicles in the space of streamlines</article-title>. <source>Neuroimage</source>. (<year>2015</year>) <volume>117</volume>:<fpage>124</fpage>&#x02013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1016/j.neuroimage.2015.05.016</pub-id><pub-id pub-id-type="pmid">25987367</pub-id></citation></ref>
<ref id="B57">
<label>57.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sedgwick</surname> <given-names>P</given-names></name></person-group>. <article-title>Spearman&#x00027;s rank correlation coefficient</article-title>. <source>Bmj</source>. (<year>2014</year>) <volume>349</volume>:<fpage>7327</fpage>. <pub-id pub-id-type="doi">10.1136/bmj.g7327</pub-id><pub-id pub-id-type="pmid">25432873</pub-id></citation></ref>
<ref id="B58">
<label>58.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Paszke</surname> <given-names>A</given-names></name> <name><surname>Gross</surname> <given-names>S</given-names></name> <name><surname>Massa</surname> <given-names>F</given-names></name> <name><surname>Lerer</surname> <given-names>A</given-names></name> <name><surname>Bradbury</surname> <given-names>J</given-names></name> <name><surname>Chanan</surname> <given-names>G</given-names></name> <etal/></person-group>. <article-title>PyTorch: an imperative style, high-performance deep learning library</article-title>. In: <source>Advances in Neural Information Processing Systems.</source> (<year>2019</year>). p. <fpage>8024</fpage>&#x02013;<lpage>35</lpage>.</citation>
</ref>
<ref id="B59">
<label>59.</label>
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Falcon</surname> <given-names>W.</given-names></name></person-group> <source>PyTorch Lightning</source>. (<year>2019</year>). GitHub Note: Available online at: <ext-link ext-link-type="uri" xlink:href="https://githubcom/PyTorchLightning/pytorch-lightning">https://githubcom/PyTorchLightning/pytorch-lightning</ext-link></citation>
</ref>
<ref id="B60">
<label>60.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>P&#x000E9;rez-Garc&#x000ED;a</surname> <given-names>F</given-names></name> <name><surname>Sparks</surname> <given-names>R</given-names></name> <name><surname>Ourselin</surname> <given-names>S</given-names></name></person-group>. <article-title>TorchIO: a Python library for efficient loading, preprocessing, augmentation and patch-based sampling of medical images in deep learning</article-title>. <source>Comput Methods Programs Biomed</source>. (<year>2021</year>) 208: 106236. <pub-id pub-id-type="doi">10.1016/j.cmpb.2021.106236</pub-id><pub-id pub-id-type="pmid">34311413</pub-id></citation></ref>
<ref id="B61">
<label>61.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wieczorek</surname> <given-names>MA</given-names></name> <name><surname>Meschede</surname> <given-names>M</given-names></name></person-group>. <article-title>Shtools: tools for working with spherical harmonics</article-title>. <source>Geochem Geophys Geosyst</source>. (<year>2018</year>) <volume>19</volume>:<fpage>2574</fpage>&#x02013;<lpage>92</lpage>. <pub-id pub-id-type="doi">10.1029/2018GC007529</pub-id></citation>
</ref>
<ref id="B62">
<label>62.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mzoughi</surname> <given-names>H</given-names></name> <name><surname>Njeh</surname> <given-names>I</given-names></name> <name><surname>Wali</surname> <given-names>A</given-names></name> <name><surname>Slima</surname> <given-names>MB</given-names></name> <name><surname>BenHamida</surname> <given-names>A</given-names></name> <name><surname>Mhiri</surname> <given-names>C</given-names></name> <etal/></person-group>. <article-title>Deep multi-scale 3D convolutional neural network (CNN) for MRI gliomas brain tumor classification</article-title>. <source>J Digit Imaging</source>. (<year>2020</year>) <volume>33</volume>:<fpage>903</fpage>&#x02013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1007/s10278-020-00347-9</pub-id><pub-id pub-id-type="pmid">32440926</pub-id></citation></ref>
<ref id="B63">
<label>63.</label>
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alexander</surname> <given-names>D</given-names></name> <name><surname>Barker</surname> <given-names>G</given-names></name> <name><surname>Arridge</surname> <given-names>S</given-names></name></person-group>. <article-title>Detection and modeling of non-Gaussian apparent diffusion coefficient profiles in human brain data</article-title>. <source>Magn Reson Med</source>. (<year>2002</year>) <volume>48</volume>:<fpage>331</fpage>&#x02013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1002/mrm.10209</pub-id><pub-id pub-id-type="pmid">12210942</pub-id></citation></ref>
<ref id="B64">
<label>64.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ghafoorian</surname> <given-names>M</given-names></name> <name><surname>Mehrtash</surname> <given-names>A</given-names></name> <name><surname>Kapur</surname> <given-names>T</given-names></name> <name><surname>Karssemeijer</surname> <given-names>N</given-names></name> <name><surname>Marchiori</surname> <given-names>E</given-names></name> <name><surname>Pesteie</surname> <given-names>M</given-names></name> <etal/></person-group>. <article-title>Transfer learning for domain adaptation in mri: Application in brain lesion segmentation</article-title>. In: <source>International Conference on Medical Image Computing and Computer-Assisted Intervention</source>. <publisher-loc>Springer</publisher-loc> (<year>2017</year>). p. <fpage>516</fpage>&#x02013;<lpage>24</lpage>.</citation>
</ref>
<ref id="B65">
<label>65.</label>
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chiou</surname> <given-names>E</given-names></name> <name><surname>Giganti</surname> <given-names>F</given-names></name> <name><surname>Punwani</surname> <given-names>S</given-names></name> <name><surname>Kokkinos</surname> <given-names>I</given-names></name> <name><surname>Panagiotaki</surname> <given-names>E</given-names></name></person-group>. <article-title>Harnessing uncertainty in domain adaptation for mri prostate lesion segmentation</article-title>. In: <source>International Conference on Medical Image Computing and Computer-Assisted Intervention</source>. <publisher-loc>Springer</publisher-loc> (<year>2020</year>). p. <fpage>510</fpage>&#x02013;<lpage>20</lpage>.</citation>
</ref>
</ref-list>
<fn-group>
<fn id="fn0001"><p><sup>1</sup><ext-link ext-link-type="uri" xlink:href="https://zenodo.org/record/1285152&#x00023;.YDeqj-qnxH4">https://zenodo.org/record/1285152&#x00023;.YDeqj-qnxH4</ext-link></p></fn>
<fn id="fn0002"><p><sup>2</sup><ext-link ext-link-type="uri" xlink:href="https://github.com/OeslleLucena/TractSegmentation">https://github.com/OeslleLucena/TractSegmentation</ext-link></p></fn>
</fn-group>
</back>
</article>