<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Mater.</journal-id>
<journal-title>Frontiers in Materials</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Mater.</abbrev-journal-title>
<issn pub-type="epub">2296-8016</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">761229</article-id>
<article-id pub-id-type="doi">10.3389/fmats.2021.761229</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Materials</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A Modular U-Net for Automated Segmentation of X-Ray Tomography Images in Composite Materials</article-title>
<alt-title alt-title-type="left-running-head">Bertoldo et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Modular U-Net for Composite Materials Segmentation</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Bertoldo</surname>
<given-names>Jo&#xe3;o P. C.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1448115/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Decenci&#xe8;re&#x2009;</surname>
<given-names>Etienne</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1547161/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Ryckelynck&#x2009;</surname>
<given-names>David</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1515817/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Proudhon</surname>
<given-names>Henry</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1490404/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>Centre des Mat&#xe9;riaux, MINES ParisTech, PSL University, <addr-line>Paris</addr-line>, <country>France</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>Centre de Morphologie Math&#xe9;matique, MINES ParisTech, PSL University, <addr-line>Paris</addr-line>, <country>France</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/125847/overview">Zhenyu Li</ext-link>, University of Science and Technology of China, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/600892/overview">Orkun Furat</ext-link>, University of Ulm, Germany</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/100865/overview">Xiaojun Wu</ext-link>, University of Science and Technology of China, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Jo&#xe3;o P. C. Bertoldo, <email>joao.bertoldo@mines-paristech.fr</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Computational Materials Science, a section of the journal Frontiers in Materials</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>25</day>
<month>11</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>8</volume>
<elocation-id>761229</elocation-id>
<history>
<date date-type="received">
<day>19</day>
<month>08</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>22</day>
<month>10</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Bertoldo, Decenci&#xe8;re&#x2009;, Ryckelynck&#x2009; and Proudhon.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Bertoldo, Decenci&#xe8;re&#x2009;, Ryckelynck&#x2009; and Proudhon</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>X-Ray Computed Tomography (XCT) techniques have evolved to a point that high-resolution data can be acquired so fast that classic segmentation methods are prohibitively cumbersome, demanding automated data pipelines capable of dealing with non-trivial 3D images. Meanwhile, deep learning has demonstrated success in many image processing tasks, including materials science applications, showing a promising alternative for a human-free segmentation pipeline. However, the rapidly increasing number of available architectures can be a serious drag to the wide adoption of this type of models by the end user. In this paper a modular interpretation of U-Net (Modular U-Net) is proposed with a parametrized architecture that can be easily tuned to optimize it. As an example, the model is trained to segment 3D tomography images of a three-phased glass fiber-reinforced Polyamide 66. We compare 2D and 3D versions of our model, finding that the former is slightly better than the latter. We observe that human-comparable results can be achievied even with only 13 annotated slices and using a shallow U-Net yields better results than a deeper one. As a consequence, neural networks show indeed a promising venue to automate XCT data processing pipelines needing no human, adhoc intervention.</p>
</abstract>
<kwd-group>
<kwd>deep learning</kwd>
<kwd>U-net</kwd>
<kwd>modular network architecture</kwd>
<kwd>semantic segmentation</kwd>
<kwd>3D X-ray computed tomography (XCT)</kwd>
<kwd>composite material</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>X-ray Computed Tomography (XCT), a characterization technique used by material scientists for non-invasive analysis, has tremendously progressed over the last 10&#xa0;years with improvements in both spatial resolution and throughput (<xref ref-type="bibr" rid="B35">Withers et&#x20;al., 2021</xref>; <xref ref-type="bibr" rid="B18">Maire and Withers, 2014</xref>). Progress with synchrotron sources, including the recent European Synchrotron Radiation Facility (ESRF) upgrade (<xref ref-type="bibr" rid="B21">Pacchioni, 2019</xref>), made it possible to look inside a specimen without destroying it in a matter of seconds (<xref ref-type="bibr" rid="B31">Shuai et&#x20;al., 2016</xref>)&#x2014;sometimes even faster (<xref ref-type="bibr" rid="B17">Maire et&#x20;al., 2016</xref>).</p>
<p>This results in a wealth of 3D tomography images (stack of 2D images) that need to be analyzed and, in some applications, it is desirable to segment them (i.e.,&#x20;transform the gray-scaled voxels into semantic categorical values). A segmented image is crucial for quantitative analyses; for instance, measuring the distribution of precipitate length and orientation (<xref ref-type="bibr" rid="B29">Shashank Kaira et&#x20;al., 2018</xref>), or phase characteristics, which can be useful for more downstream applications like estimating thermo-mechanical properties (<xref ref-type="bibr" rid="B34">Strohmann et&#x20;al., 2019</xref>).</p>
<p>XCT images typically have billions of voxels, weighting several gigabytes, and remain complex to inspect manually even using dedicated costly software (e.g., Avizo, VGStudioMax). Using thresholding techniques on the gray level image is an easy, useful method to segment phases in tomographies, but it fails in complex cases, in particular when acquisition artifacts (e.g., rings, beam hardening, phantom gradients) are present. Algorithms based on mathematical morphology like the watershed segmentation (<xref ref-type="bibr" rid="B8">Beucher, 1994</xref>) help tackling more complex scenarios, but they need human parametrization, which often requires expertise in the application domain. Thus, scaling quantitative analyses is expensive, creating a bottleneck to process 3D XCT&#x2014;or even 4D (3D with time steps).</p>
<p>Deep learning approaches offer a viable solution to attack this issue because neural networks can generalize patterns learned from annotated data. A neural network is a statistical model originated from perceptrons (<xref ref-type="bibr" rid="B26">Rosenblatt, 1958</xref>) capable of approximating a generic function. Convolutional neural networks (CNNs) (<xref ref-type="bibr" rid="B5">Bengio and Lecun, 1997</xref>), a variation adapted to spatially-structured data (time series, images, volumes), made great advances in computer vision tasks. Since the emergence of popular deep learning frameworks, more problem-specific architectures have been proposed, such as fully-convolutional neural networks (<xref ref-type="bibr" rid="B30">Shelhamer et&#x20;al., 2017</xref>), a convolution-only type of model used to map image pixel values to another domain (e.g., classification or regression).</p>
<p>
<xref ref-type="bibr" rid="B29">Shashank Kaira et&#x20;al. (2018)</xref> trained a model to segment three phases in 3D nanotomographies of an Al-Cu alloy, showing that even a simple convolutional neural network can reproduce patterns of a human-made segmentation. <xref ref-type="bibr" rid="B32">Stan et&#x20;al. (2020)</xref> optimized a SegNet (<xref ref-type="bibr" rid="B4">Badrinarayanan et&#x20;al., 2017</xref>) to segment dendrites of different alloys, including a 4D XCT. <xref ref-type="bibr" rid="B34">Strohmann et&#x20;al. (2019)</xref> identified Aluminides and Si phases in XCT using a U-Net, an architecture that, along with its many flavors (<xref ref-type="bibr" rid="B25">Ronneberger et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B11">&#xc7;i&#xe7;ek et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B24">Qin et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B38">Zhou et&#x20;al., 2018</xref>), has shown success in a variety of applications (<xref ref-type="bibr" rid="B20">Oktay et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B33">Stoller et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B37">Zhang et&#x20;al., 2018</xref>). Finally, <xref ref-type="bibr" rid="B13">Furat et&#x20;al. (2019)</xref> combined U-Nets with classic segmentation algorithms (e.g., marker-based watershed) to segment grain boundaries in successive XCTs of an Al-Cu specimen as it is submitted to Ostwald ripening&#x20;steps.</p>
<p>If great proof of concepts have been made in the field, the variety of architectures available and the ability to optimize them is a real challenge for the end user. We present a modular U-Net architecture that has the potential for wider adoption of CNN-based segmentation tasks as it makes it easier for the end user to tune its architecture based on objective measures of the performances. Our architecture, the Modular U-Net (<xref ref-type="fig" rid="F1">Figure&#x20;1</xref> and <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>), is proposed as a generalized representation of the U-Net, explicitly factorizing the U-like structure from its composing blocks. Describing U-Nets in this way makes it easier to analyze its components as hyperparameters and provides elements of vocabulary to communicate their details more easily.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Modular U-Net: a generalization of the U-Net architecture.</p>
</caption>
<graphic xlink:href="fmats-08-761229-g001.tif"/>
</fig>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Examples of Modular U-Net blocks. Left: our Convolutional Block. Middle/right: rigid/learnable Down-sampling Block and Up-sampling Block.</p>
</caption>
<graphic xlink:href="fmats-08-761229-g002.tif"/>
</fig>
<p>In this paper, an annotated 3D XCT of glass fiber-reinforced Polyamide 66 (<xref ref-type="fig" rid="F3">Figure&#x20;3</xref>) is presented as an example of segmentation problem in Materials Science that can be automated with a deep learning approach. Like <xref ref-type="bibr" rid="B13">Furat et&#x20;al. (2019)</xref>, we compare three variants on the composite material dataset focusing on the dimensionality of the convolutions (2D or 3D), obtaining qualitatively human-like segmentation (<xref ref-type="fig" rid="F4">Figures 4</xref>&#x2013;<xref ref-type="fig" rid="F6">6</xref>) with all of them although 2D-convolutions yield better results (<xref ref-type="fig" rid="F8">Figure&#x20;8</xref>). We find that (for the considered material) the U-Net architecture can be shallow without loss of performance, but batch normalization is necessary for the optimization (<xref ref-type="fig" rid="F9">Figure&#x20;9</xref>). Finally, a learning curve (<xref ref-type="fig" rid="F10">Figure&#x20;10</xref>) shows that only ten annotated 2D slices are necessary to train our&#x20;model.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Glass fiber-reinforced Polyamide 66 tomography. Left: A 1300 &#xd7; 1040 slice on the XY plane of the volume Train-Val-Test. In the upper left corner, a histogram of the gray level values in the image (linear scale in black, log scale in gray). Right: Zoom. Ring artifacts&#x2014;from the acquisition process&#x2014;can be as dark as porosities, making it harder to segment such regions.</p>
</caption>
<graphic xlink:href="fmats-08-761229-g003.tif"/>
</fig>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Segmentation generated by the 2D model on the test set. Phase color code: blue represents voxels correctly classified as fiber (hatched), yellow as porosity (contours), and red represents misclassification. Notice that some regions in blue are in fact mistakenly segmented as fiber, although it is considered as correct because the annotations contain the same mistake. <bold>(B)</bold> is a zoom of the yellow-countoured area in <bold>(A)</bold>.</p>
</caption>
<graphic xlink:href="fmats-08-761229-g004.tif"/>
</fig>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Segmentation generated by the 2D model on the volume Crack. Image rendered with standard tools in Avizo for the sake of visualization: two orthogonal planes inside the specimen, fibers rendered in 3D at the bottom, and the fracture rendered as a surface. Phase color code: blue represents the fiber and yellow represents the porosity.</p>
</caption>
<graphic xlink:href="fmats-08-761229-g005.tif"/>
</fig>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Segmentation of the volume Crack. A crop from the vertical plane in a slice passing through the fracture. The fiber phase is hatched in blue and the porosity phase is contoured in yellow.</p>
</caption>
<graphic xlink:href="fmats-08-761229-g006.tif"/>
</fig>
<p>Our results confirm that NNs can be not only a quality-wise satisfactory but also a viable solution in practice for XCT segmentation as it requires little annotated data and shallow models (therefore faster to train). Furthermore, the Modular U-Net, where the architecture becomes a set of hyperparameters, allows to optimize the model more easily and may ease the adoption of CNN-based models for automated X-ray tomography segmentation&#x20;tasks.</p>
</sec>
<sec id="s2">
<title>2 Materials and Methods</title>
<sec id="s2-1">
<title>2.1 Data</title>
<p>The data used in this work is composed of synchrotron X-ray tomography volumes recorded using 2&#xa0;mm &#xd7; 2&#xa0;mm cross section composite specimens of Polyamide 66 reinforced by glass fibers, a material commonly used for structural pieces in different applications. XCT allows to assess the relation between the microstructure and the mechanical properties such as the resistance to the formation of damage under load. For this matter, segmenting such images automatically would, for instance, make it possible to visualize the propagation of cracks inside a specimen.</p>
<p>Two volumetric images (of different specimen) were used in this study. First, a volume of 2048<sup>3</sup> voxels, referred to as <italic>Train-Val-Test</italic> (<xref ref-type="fig" rid="F3">Figure&#x20;3</xref>, <xref ref-type="fig" rid="F4">4</xref>), was cropped to get rid of the specimen&#x2019;s borders, and its ground truth segmentation was created semi-manually with ImageJ (<xref ref-type="bibr" rid="B28">Schneider et&#x20;al., 2012</xref>) using Fiji (<xref ref-type="bibr" rid="B27">Schindelin et&#x20;al., 2012</xref>) and Avizo. Second, another specimen containing a fracture was segmented in order to qualitatively evaluate our models (<xref ref-type="fig" rid="F5">Figures 5</xref>, <xref ref-type="fig" rid="F6">6</xref>)&#x2014;it is further referred to as &#x201c;<italic>Crack</italic>&#x201d;&#x2014;on a representative application of interest. All the data used in this work is publicly available; <xref ref-type="table" rid="T1">Table&#x20;1</xref> describes the files published on Zeonodo (<xref ref-type="bibr" rid="B6">Bertoldo et al., 2021a</xref>).</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Default hyperparameters. Parameters not mentioned are the default in TensorFlow 2.2.0.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Parameter</th>
<th align="center">2D</th>
<th align="center">2.5D</th>
<th align="center">3D</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">U-depth</td>
<td align="center">3</td>
<td align="center">3</td>
<td align="center">3</td>
</tr>
<tr>
<td align="left">Convolution kernel</td>
<td align="center">3&#x20;&#xd7; 3</td>
<td align="center">3&#x20;&#xd7; 3</td>
<td align="center">3&#x20;&#xd7; 3&#x20;&#xd7; 3</td>
</tr>
<tr>
<td align="left">Batch size (nb. of crops)</td>
<td align="center">10</td>
<td align="center">10</td>
<td align="center">10</td>
</tr>
<tr>
<td align="left">Epoch size (nb. of batches)</td>
<td align="center">10</td>
<td align="center">10</td>
<td align="center">10</td>
</tr>
<tr>
<td align="left">Crop shape</td>
<td align="center">160&#x20;&#xd7; 160</td>
<td align="center">160&#x20;&#xd7; 160&#x20;&#xd7; 5</td>
<td align="center">32&#x20;&#xd7; 32&#x20;&#xd7; 32</td>
</tr>
<tr>
<td align="left">Dropout</td>
<td align="center">10%</td>
<td align="center">10%</td>
<td align="center">10%</td>
</tr>
<tr>
<td align="left">Gaussian noise (zero mean) standard deviation</td>
<td align="center">0.03</td>
<td align="center">0.03</td>
<td align="center">0.03</td>
</tr>
<tr>
<td align="left">f0</td>
<td align="center">16</td>
<td align="center">16</td>
<td align="center">16</td>
</tr>
<tr>
<td align="left">Up/Down-sampling stride or Max pooling size</td>
<td align="center">2&#x20;&#xd7; 2</td>
<td align="center">2&#x20;&#xd7; 2</td>
<td align="center">2&#x20;&#xd7; 2&#x20;&#xd7; 2</td>
</tr>
<tr>
<td align="left">Batch normalization momentum</td>
<td align="center">0.5</td>
<td align="center">0.5</td>
<td align="center">0.5</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s2-1-1">
<title>2.1.1 Acquisition</title>
<p>X-ray tomography scans were recorded on the Psich&#xe9; beamline at the Synchrotron SOLEIL using a parallel pink beam. The incident beam spectrum was characterized by a peak intensity at 25&#xa0;keV, defined by the silver absorption edge, with a full width at half maximum bandwidth of approximately 1.8&#xa0;keV. The total flux at the sample position was about 2.8e12 photons/s/mm<sup>2</sup>. The detector placed after the sample was constituted by a LuAG scintillator, a 5&#xd7; magnifying optics, and a Hamamatsu CMOS 2048&#x20;&#xd7; 2048 pixels detector (effective pixel size of 1.3&#xa0;&#xb5;m). 1,500 radiographs were collected over a 180&#xb0; rotation and an exposure of 50&#xa0;ms (full scan duration of 2&#xa0;min). The sets of radiographs were then processed using PyHST2 reconstruction software (<xref ref-type="bibr" rid="B19">Mirone et&#x20;al., 2014</xref>) with the Paganin filter (<xref ref-type="bibr" rid="B22">Paganin et&#x20;al., 2002</xref>) activated to enhance the contrast between the phases.</p>
</sec>
<sec id="s2-1-2">
<title>2.1.2 Phases</title>
<p>The three phases in the material are visible in <xref ref-type="fig" rid="F3">Figures 3</xref>, <xref ref-type="fig" rid="F4">4</xref>: the polymer matrix (gray), the glass fibers (white, hatched in blue), and damage in form of pores (dark gray and black, contoured in yellow). Although the pores (absence of material) do not constitute a real phase, they can be thought of as such for the sake of this study, and will be referred as porosity. Furthermore, a segmentation can be seen as a voxel-wise classification so the phases here are also referred to as &#x201c;classes.&#x201d;</p>
</sec>
<sec id="s2-1-3">
<title>2.1.3 Ground Truth</title>
<p>The data was annotated in two steps: first, the fiber and the porosity phases were independently segmented using Seeded Region Growing (<xref ref-type="bibr" rid="B2">Adams and Bischof, 1994</xref>); then, ring artifacts that leaked to the porosity class were manually corrected. A detailed description of the procedure is presented in the <xref ref-type="sec" rid="s11">Supplementary Material</xref>.</p>
</sec>
<sec id="s2-1-4">
<title>2.1.4 Data Split</title>
<p>The ground truth tomography slices (of the Train-Val-Test volume) were sequentially split&#x2014;their order was preserved to train the 3D models (<xref ref-type="sec" rid="s2-2-2">Section 2.2.2</xref>)&#x2014;into three sets: train (1300 slices), validation (128 slices), and test (300 slices). A margin of 86 slices between these sets was adopted to avoid information leakage. The &#x201c;train&#x201d; range was used to train the models, the &#x201c;validation&#x201d;) (or &#x201c;val&#x201d;) was used to select the best model (during the optimization), and the &#x201c;test&#x201d; was used to evaluate the models (<xref ref-type="sec" rid="s3">Section&#x20;3</xref>).</p>
</sec>
<sec id="s2-1-5">
<title>2.1.5 Class Imbalance</title>
<p>Due to the material&#x2019;s nature, the classes (phases) in this dataset are intrinsically imbalanced. The matrix, the fiber, and the porosity represent, respectively, 82.3, 17.2, and 0.5% of the voxels.</p>
</sec>
</sec>
<sec id="s2-2">
<title>2.2 Neural Network</title>
<p>Let <inline-formula id="inf1">
<mml:math id="m1">
<mml:mspace width="0.22em"/>
<mml:mi>x</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">X</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>0,1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>h</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mspace width="0.22em"/>
</mml:math>
</inline-formula> be a normalized gray 3D image. Its segmentation <inline-formula id="inf2">
<mml:math id="m2">
<mml:mspace width="0.22em"/>
<mml:mi>y</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">Y</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mi>C</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>h</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mspace width="0.22em"/>
</mml:math>
</inline-formula>, where [[<italic>C</italic>]] &#x3d; {1, 2, <italic>&#x2026;</italic> , <italic>C</italic>}, contains a class value in each voxel, which may represent any categorical information, such as the phase of the material. In this setting, a segmentation algorithm is a function <inline-formula id="inf3">
<mml:math id="m3">
<mml:mspace width="0.22em"/>
<mml:mi>f</mml:mi>
<mml:mo>:</mml:mo>
<mml:mi mathvariant="script">X</mml:mi>
<mml:mo>&#x2192;</mml:mo>
<mml:mi mathvariant="script">Y</mml:mi>
<mml:mspace width="0.22em"/>
</mml:math>
</inline-formula>. In this section we present our approach (that is, the <italic>f</italic>) used to segment the data described in the previous section.</p>
<p>In this section, a generic U-Net architecture, which we coined Modular U-Net (<xref ref-type="fig" rid="F1">Figure&#x20;1</xref>), is proposed. We explain how it is composed and describe the modules used in this work. Three variants of the Modular U-Net were considered; their differences&#x2014;based on the input, convolution, and output nature (2D or 3D)&#x2014;are then analyzed (summary in <xref ref-type="table" rid="T2">Table&#x20;2</xref>). Finally, our training setup is described, specifying the loss function, optimizer, learning rate, data augmentation, software, and hardware&#x20;used.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Modular U-Net variations: input, convolutional layer, and output nature (2D or 3D).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Model</th>
<th align="center">Input (data)</th>
<th align="center">Convolution</th>
<th align="center">Output (segm.)</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">2D</td>
<td align="char" char=".">2D</td>
<td align="char" char=".">2D</td>
<td align="char" char=".">2D</td>
</tr>
<tr>
<td align="left">2.5D</td>
<td align="char" char=".">3D</td>
<td align="char" char=".">2D</td>
<td align="char" char=".">2D</td>
</tr>
<tr>
<td align="left">3D</td>
<td align="char" char=".">3D</td>
<td align="char" char=".">3D</td>
<td align="char" char=".">3D</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s2-2-1">
<title>2.2.1 Modular U-Net</title>
<p>Since <xref ref-type="bibr" rid="B25">Ronneberger et&#x20;al. (2015)</xref> proposed U-Net, variations of it emerged in the literature [e.g., <xref ref-type="bibr" rid="B11">&#xc7;i&#xe7;ek et&#x20;al. (2016)</xref>; <xref ref-type="bibr" rid="B24">Qin et&#x20;al. (2020)</xref>]. Here we propose a generalized version, preserving its overall structure. The <italic>Modular U-Net</italic> is based on three blocks (<xref ref-type="fig" rid="F2">Figure&#x20;2</xref>): the Convolutional Block (ConvBlock), the Down-sampling Block (DownBlock), and the Up-sampling Block (UpBlock).</p>
<p>The left/right side of the architecture corresponds to an encoder/decoder: a repetition of pairs of ConvBlock and DownBlock/UpBlock modules. The two sides are connected by concatenations between their respective parts at the same U-level&#x2014;which corresponds to the inner tensors&#x2019; resolutions (higher U-level means lower resolution). The U-depth, a hyperparameter, is the number of U-levels in a model, corresponding to the number of DownBlock (and equivalently UpBlock) modules.</p>
<p>The ConvBlock is a combination of operations that outputs a tensor with the same spatial dimensions of its input, though the number of channels may differ&#x2014;in our models it always doubles. The assumption of equally-sized input/output is optional, but we admit it for the sake of simplicity because it makes the model easier to be used with an arbitrarily shaped volume. The number of channels after the ConvBlocks is 2<sup>U&#x2212;level</sup> &#xd7; <italic>f</italic>
<sub>0</sub>, where <italic>f</italic>
<sub>0</sub> is the number of filters in the first convolution.</p>
<p>The DownBlock/UpBlock applies a down/up-sampling operation&#x2014;which can be a learnable convolution/transposed convolution or a &#x201c;rigid&#x201d; max pooling/up-sampling (see <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>)&#x2014;changing the tensor&#x2019;s shape by dividing/multipling by two every spatial dimension: width, length, and depth in the 3D case. In other words, a tensor with shape (<italic>w</italic>, <italic>h</italic>, <italic>d</italic>, <italic>c</italic>)&#x2014;where <italic>w</italic>, <italic>h</italic>, and <italic>d</italic> are the dimensions of the volume (respectively, the width, height, and depth), and <italic>c</italic> is the number of channels&#x2014;becomes <inline-formula id="inf4">
<mml:math id="m4">
<mml:mspace width="0.22em"/>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>d</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mspace width="0.22em"/>
</mml:math>
</inline-formula> after a DownBlock and (2<italic>w</italic>, 2<italic>h</italic>, 2<italic>d</italic>, <italic>c</italic>) after an UpBlock.</p>
<p>In <xref ref-type="bibr" rid="B25">Ronneberger et&#x20;al. (2015)</xref>, for instance, the ConvBlock is a sequence of two 3&#x20;&#xd7; 3 convolutions with ReLU activation, the DownBlock is a max pooling, and the UpBlock is an up-sampling layer. In <xref ref-type="bibr" rid="B11">&#xc7;i&#xe7;ek et&#x20;al. (2016)</xref>, the ConvBlock is a 3D convolutional layer, and in <xref ref-type="bibr" rid="B24">Qin et&#x20;al. (2020)</xref> it is a nested U-Net.</p>
<p>The ConvBlock used here (<xref ref-type="fig" rid="F2">Figure&#x20;2</xref>) is a sequence of two 3&#x20;&#xd7; 3 (x3 in the 3D case) convolutions with ReLU activation, a residual connection with another convolution, batch normalization before each activation, and dropout at the end. The DownBlock is a 3 &#xd7; 3 convolution with 2 &#xd7; 2 stride, and the UpBlock is a 3 &#xd7; 3 transposed convolution with 2 &#xd7; 2 stride.</p>
</sec>
<sec id="s2-2-2">
<title>2.2.2 Variations: 2D, 2.5D, and 3D</title>
<p>Since our dataset contains intrinsically 3D structures, we compared the performances of this architecture using 2D and 3D convolutions. The 2D model processes individual tomography slices independently, while the 3D one processes several at once (i.e. a volume). We also compared an intermediate version, which we coined 2.5D; it processes one tomography slice at a time with 2D convolutions taking five consecutive slices at the input (the processed slice plus two above and two below), as if the extra slices were channels of a 2D&#x20;image.</p>
<p>The 2D and 2.5D models can only process each slice of a plane individually; the XY plane (slices in the <italic>z</italic>-axis, as in <xref ref-type="fig" rid="F3">Figures 3</xref>, <xref ref-type="fig" rid="F4">4</xref>) was chosen because the visual characteristics in this direction are mostly invariant&#x2014;we observed a high correlation between adjacent slices.</p>
<p>These three variants have different receptive fields (therefore access to different information) and mix up the input entries differently (<xref ref-type="table" rid="T2">Table&#x20;2</xref> summarizes these differences). The 2.5D model has access to extra information relative to the 2D model; its first convolution layer can correlate local information from adjacent tomography slices, although the rest of the network is identical in terms of structure. Meanwhile, the 3D model uses information from adjacent slices in every network layer, enabling more complex correlations but also increasing the model&#x2019;s variance.</p>
</sec>
<sec id="s2-2-3">
<title>2.2.3 Training</title>
<p>The next paragraphs describe how our models were trained; <xref ref-type="table" rid="T1">Table&#x20;1</xref> summarizes the hyperparameters of our neural network.</p>
<sec id="s2-2-3-1">
<title>2.2.3.1 Loss Function</title>
<p>Our models were trained using a custom loss inspired on the Jaccard index, also known as Intersection over Union (IoU).</p>
<p>DEF. 1. <italic>Let</italic> <italic>A</italic> and <italic>B</italic> be two sets, and &#x7c;&#x22c5;&#x7c; denote a set&#x2019;s cardinality. The Jaccard index <italic>J</italic>&#x20;&#x2208; [0, 1] is<disp-formula id="e1">
<mml:math id="m5">
<mml:mi>J</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>B</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>A</mml:mi>
<mml:mo>&#x2229;</mml:mo>
<mml:mi>B</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>A</mml:mi>
<mml:mo>&#x222a;</mml:mo>
<mml:mi>B</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>A</mml:mi>
<mml:mo>&#x2229;</mml:mo>
<mml:mi>B</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>A</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>&#x2b;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>B</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>A</mml:mi>
<mml:mo>&#x2229;</mml:mo>
<mml:mi>B</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:mtext>.</mml:mtext>
</mml:math>
<label>(1)</label>
</disp-formula>
</p>
<p>Similar to <xref ref-type="bibr" rid="B12">Duque-Arias et&#x20;al. (2021)</xref>, we adapt the right hand side of <xref ref-type="disp-formula" rid="e1">Eq. 1</xref> to define the multi-class Jaccard<sup>2</sup> loss for a batch of images.</p>
<p>For the sake of simplicity, the spatial dimensions (see <xref ref-type="sec" rid="s2-2">Section 2.2</xref>) are omitted, and a single index <italic>n</italic> refers to position of individual <italic>voxels in a batch</italic>. In other words, instead of referring to a batch of normalized 3D images in [0, 1]<sup>
<italic>w</italic>&#xd7;<italic>h</italic>&#xd7;<italic>d</italic>&#xd7;<italic>b</italic>
</sup>, where <italic>b</italic> is the batch size, we refer to an unraveled batch of voxels in [0, 1]<sup>
<italic>N</italic>
</sup>, where <italic>N</italic>&#x20;&#x3d; <italic>w</italic>&#x20;&#xd7; <italic>h</italic>&#x20;&#xd7; <italic>d</italic>&#x20;&#xd7; <italic>b</italic>. Nevertheless, the 3D structure remains implicitly unchanged&#x2014;notice that the definition below remains the same for 2D images.</p>
<p>DEF. 2. Let <italic>y</italic>&#x20;&#x2208; {0,1}<sup>
<italic>N</italic>&#xd7;<italic>C</italic>
</sup> be a batch of ground truth voxels, where <italic>N</italic> is the number of voxels, each belonging to one out of <italic>C</italic> classes, such that<disp-formula id="equ1">
<mml:math id="m6">
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="cases">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>if&#x2009;the&#x2009;voxel</mml:mtext>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mtext>belongs&#x2009;to&#x2009;the&#x2009;class</mml:mtext>
<mml:mspace width="0.28em"/>
<mml:mi>c</mml:mi>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mn>0</mml:mn>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mtext>otherwise</mml:mtext>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
<mml:mtext>,</mml:mtext>
</mml:math>
</disp-formula>where <italic>y</italic>
<sub>
<italic>n</italic>
</sub> is the voxel at position <italic>n</italic>&#x20;&#x2208; [[<italic>N</italic>]], and <italic>y</italic>
<sub>
<italic>n</italic>
</sub>,<sub>
<italic>c</italic>
</sub> is its cth entry (<italic>c</italic>&#x20;&#x2208; [[<italic>C</italic>]])<italic>.</italic>
</p>
<p>A model&#x2019;s last activation map, a per-voxel softmax, is a tensor <inline-formula id="inf5">
<mml:math id="m7">
<mml:mspace width="0.22em"/>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>0,1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mspace width="0.22em"/>
</mml:math>
</inline-formula>, where each row is a probability vector <inline-formula id="inf6">
<mml:math id="m8">
<mml:mspace width="0.22em"/>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>0,1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mspace width="0.22em"/>
</mml:math>
</inline-formula>, and the component <inline-formula id="inf7">
<mml:math id="m9">
<mml:mspace width="0.22em"/>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> corresponds to the probability assigned to the class&#x20;<italic>c</italic>.</p>
<p>The Jaccard<sup>2</sup> loss of the batch <inline-formula id="inf8">
<mml:math id="m10">
<mml:mspace width="0.22em"/>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mspace width="0.22em"/>
</mml:math>
</inline-formula> is defined as<disp-formula id="e2">
<mml:math id="m11">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mi>J</mml:mi>
<mml:mn>2</mml:mn>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right"/>
<mml:mtd columnalign="left">
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2a;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2a;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfrac>
<mml:mtext>,</mml:mtext>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(2)</label>
</disp-formula>where <inline-formula id="inf9">
<mml:math id="m12">
<mml:mspace width="0.22em"/>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2a;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mspace width="0.22em"/>
</mml:math>
</inline-formula> is the probability assigned to the correct class of the voxel&#x20;<italic>n</italic>.</p>
<p>Notice that <italic>J</italic>2 &#x2208; (0, 1), thus it can, conveniently, be expressed as a percentage, and it is a measure of <italic>dissimilarity</italic>&#x2014;<italic>J</italic>2 &#x3d; 100<italic>%</italic> is a completely uncorrelated estimation, and <italic>J</italic>2 &#x3d; 0<italic>%</italic> is a perfect replication of the ground truth&#x2014;while <xref ref-type="disp-formula" rid="e1">Eq. 1</xref> measures the similarity between two&#x20;sets.</p>
</sec>
<sec id="s2-2-3-2">
<title>2.2.3.2 Data Augmentation</title>
<p>In order to increase the variability of the data, random crops were selected from the data volume, then a random geometric transformation (flip, 90&#xb0; rotation, transposition, etc) is applied. As our training dataset is reasonably large, we used a simple data augmentation scheme, but richer transformations could be applied as long as the transformations result in credible samples.</p>
</sec>
<sec id="s2-2-3-3">
<title>2.2.3.3 Optimization</title>
<p>We used AdaBelief, an optimizer that combines the training stability and fast convergence of adaptive optimizers [e.g., Adam (<xref ref-type="bibr" rid="B16">Kingma and Ba, 2015</xref>)<xref ref-type="fn" rid="FN1">
<sup>1</sup>
</xref>] and good generalization capabilities of accelerated schemes (e.g., stochastic gradient descent).</p>
<p>Each model was trained for 200 epochs, each with 10 batches of 10 crops. The learning rate is set to 10<sup>&#x2212;3</sup> during the first 100 epochs then linearly decays until 10<sup>&#x2212;4</sup> during the last 100 epochs. This is done to quickly converge in the initial phase, then stabilize the optimization once it&#x2019;s close to convergence. We verified that 200 epochs was enough to achieve convergence for all the models (the validation loss plateaus, oscilating randomly), and select the model with the lowest loss (i.e.,&#x20;not necessarily the one from the last epoch).</p>
</sec>
<sec id="s2-2-3-4">
<title>2.2.3.4 Software and Hardware</title>
<p>We trained our models using Keras with TensorFlow&#x2019;s (<xref ref-type="bibr" rid="B1">Abadi et&#x20;al., 2015</xref>) GPU-enabled version with CUDA 10.1 running on two NVIDIA Quadro P4000<xref ref-type="fn" rid="FN2">
<sup>2</sup>
</xref> (2 &#xd7; 8&#xa0;GN).</p>
</sec>
</sec>
</sec>
</sec>
<sec id="s3">
<title>3 Results</title>
<p>In this section we present a compilation of qualitative and quantitative results obtained. The training process took one to 3&#xa0;h for each model, up to 8&#xa0;h for the largest 3D model. The segmentations from the three Modular U-Net versions presented in <xref ref-type="sec" rid="s2-2">Section 2.2</xref> are quantitatively compared, then an ablation analysis and the learning curve of the 2D model are presented. All the quantitative analyses were made on the test split (see <xref ref-type="sec" rid="s2-1">Section 2.1</xref>), which contains 1300 &#xd7; 1040&#x20;&#xd7; 300&#x20;&#x2248; 406&#xd7;10<sup>6</sup>&#x2009;voxels. The trained models and the data used to produce our results are publicly available online (<xref ref-type="bibr" rid="B6">Bertoldo et&#x20;al., 2021a</xref>,<xref ref-type="bibr" rid="B7">b</xref>). A further detailed analysis is provided as <xref ref-type="sec" rid="s11">Supplementary Material</xref>.</p>
<sec id="s3-1">
<title>3.1 Qualitative Results</title>
<p>
<xref ref-type="fig" rid="F4">Figure&#x20;4</xref> shows two snapshots of the volume Train-Val-Test in the test partition with a superposed segmentation generated with the 2D model. The zoomed area illustrates examples of misclassified regions&#x2014;including parts of data with annotation mistakes.</p>
<p>The volume <italic>Crack</italic> was used to evaluate our method&#x2019;s usability with another image (i.e.,&#x20;a volume from another specimen of the same material); the segmentation obtained with the 2D model is presented in <xref ref-type="fig" rid="F5">Figures 5</xref>, <xref ref-type="fig" rid="F6">6</xref>. In terms of processing speed, the results show that our method could be carried out almost in real time using typical hardware available at a synchrotron beamline: using an NVIDIA Quadro P2000 (5&#xa0;GB), it took 32&#xa0;min to process the volume Crack, with approximately 5.8 billion voxels (1579 &#xd7; 1845&#x20;&#xd7; 2002).</p>
</sec>
<sec id="s3-2">
<title>3.2 Baseline</title>
<p>For the sake of comparison, we considered two theoretical models: the Zero<sup>
<italic>th</italic>
</sup> Order Classifier (ZeroOC) and the Bin-wise ZeroOC. Both are described below, and their expected performances are summarized in <xref ref-type="table" rid="T3">Table&#x20;3</xref>. This analysis shows us the minimum expected Jaccard index for each class (see <xref ref-type="table" rid="T3">Table&#x20;3</xref>); notably, the mean class-wise Jaccard index of the Bin-wise ZeroOC is&#x20;76.2<italic>%</italic>.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Expected performance of baseline theoretical models in terms of class-wise Jaccard index (%).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Model</th>
<th align="center">Description</th>
<th align="center">Matrix</th>
<th align="center">Fiber</th>
<th align="center">Porosity</th>
<th align="center">Mean</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">ZeroOC</td>
<td align="left">Classify every voxel with the majority class (matrix)</td>
<td align="char" char=".">81.0</td>
<td align="char" char=".">0</td>
<td align="char" char=".">0</td>
<td align="char" char=".">27.0</td>
</tr>
<tr>
<td align="left">Bin-wise ZeroOC</td>
<td align="left">Classify a voxel based only on its value. The majority class of each value is chosen. This is equivalent to a ZeroOC model per gray level</td>
<td align="char" char=".">98.4</td>
<td align="char" char=".">94.2</td>
<td align="char" char=".">35.9</td>
<td align="char" char=".">76.2</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The ZeroOC is the simplest model possible: it classifies every voxel with the majority class (the matrix phase), ignoring all the information from the&#x20;data.</p>
<p>The Bin-wise ZeroOC takes gray level values into consideration individually (ignoring its neighbors) relying on the class imbalance on a per-value basis. <xref ref-type="fig" rid="F7">Figure&#x20;7</xref>, the histograms of gray level value per class, illustrates this model&#x2019;s principle: for a given gray level value in the <italic>x</italic>-axis, it selects the respective class of the highest curve in the <italic>y</italic>-axis&#x2014;in other words, it always chooses the majority class conditioned on a given the gray level. The Bin-wise ZeroOC can be seen as a look-up table of gray level values as input and class as output, completely ignoring the contextual information available (the neighbor voxels).</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Glass fiber-reinforced Polyamide 66 gray value (normalized) histograms (one per class). The histogram is normalized globally, i.e.,&#x20;a bin&#x2019;s value is the proportion of voxels out of all the voxels (all classes confounded). The superposition of the classes&#x2019; value ranges make it impossible to segment the image with a threshold on the gray values.</p>
</caption>
<graphic xlink:href="fmats-08-761229-g007.tif"/>
</fig>
</sec>
<sec id="s3-3">
<title>3.3 Quantitative Results</title>
<p>
<xref ref-type="fig" rid="F8">Figure&#x20;8</xref> presents a comparison of the three model variations (2D, 2.5D, and 3D), where they are evaluated with varying sizes. The models are scaled with the hyperparameter <italic>f</italic>
<sub>0</sub> (values inside the parentheses in <xref ref-type="fig" rid="F8">Figure&#x20;8</xref>).</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>Modular U-Net variations comparison. On the <italic>x</italic>-axis, the number of parameters; on the <italic>y</italic>-axis the mean class-wise Jaccard indices. The Modular U-Net 2D, 2.5D, and 3D versions are scaled with <italic>f</italic>
<sub>0</sub> (in parentheses), the number of filters of the first convolution of the first Convolutional Block.</p>
</caption>
<graphic xlink:href="fmats-08-761229-g008.tif"/>
</fig>
<p>The performance is measured using the Jaccard index, also known as Intersection over Union (IoU), on each phase (class). Our main metric is the arithmetic mean of the three class-wise indices, and the baseline minimum is 76.2<italic>%</italic>, Bin-wise ZeroOC&#x2019;s performance. This metric provides a good visibility of the performance differences and resumes the precision-recall trade off; other classic metrics&#x2014;even the area under the Receiver Operating Characteristic (ROC) curve (<xref ref-type="bibr" rid="B14">Hanley and McNeil, 1982</xref>)&#x2014;are close to 100% (see the <xref ref-type="sec" rid="s11">Supplementary Material</xref>), so the differences are hard to compare.</p>
</sec>
<sec id="s3-4">
<title>3.4 Ablation Study</title>
<p>
<xref ref-type="fig" rid="F9">Figure&#x20;9</xref> presents a component ablation analysis of the 2D model with <italic>f</italic>
<sub>0</sub> &#x3d; 16 in terms of number of parameters and performance. Starting with the 2D model with the default hyperparameters (see <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>; <xref ref-type="table" rid="T1">Table&#x20;1</xref>), we retrained other models removing one component at a time. The learnable up/down-samplings were replaced by &#x201c;rigid&#x201d; ones (see <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>), the 2D convolutions were replaced by separable ones, and the batch normalization was replaced by layer normalization. Finally, we varied the U-depth from 2 to&#x20;4.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>Modular U-Net ablation (2D model). On the <italic>x</italic>-axis, the number of parameters; on the <italic>y</italic>-axis the mean class-wise Jaccard indices. Components were removed individually, or replaced by alternatives. Removals: dropout, gaussian noise, residual (skip connection), batch normalization. Replacements: convolutions by separable ones, learnable Down-sampling Block/Up-sampling Block by rigid ones (<xref ref-type="fig" rid="F2">Figure&#x20;2</xref>), and batch normalization by layer normalization. We also compare the effect of the U-depth, i.e.,&#x20;number of levels in the U structure. Notice that the data point no batch norm is out of scale in the <italic>y</italic>-axis for the sake of visualization.</p>
</caption>
<graphic xlink:href="fmats-08-761229-g009.tif"/>
</fig>
<p>Notice that the model without dropout performed better than the reference model, but we kept it in our default parameters because the same thing did not occur with other variations and&#x20;sizes.</p>
</sec>
<sec id="s3-5">
<title>3.5 Learning Curve</title>
<p>Finally, we computed the learning curve of the 2D model in order to assess the trade-off between performance and amount of annotated data. As the data annotation is time-consuming and requires expertise (therefore expensive), this analysis intends to estimate how much annotation is necessary to achieve our results. The number of (consecutive) slices in the training dataset was progressively decreased from 1024 to 1 while the validation dataset is kept the same (300 slices) for evaluation. This experiment was run with the 2D Modular U-Net using the default hyperparameters (presented in <xref ref-type="table" rid="T1">Table&#x20;1</xref>), and the results presented in <xref ref-type="fig" rid="F10">Figure&#x20;10</xref> in terms of class-wise Jaccard index and their arithmetic&#x20;mean.</p>
<fig id="F10" position="float">
<label>FIGURE 10</label>
<caption>
<p>Learning curve of the 2D model with default hyperparameters (<xref ref-type="table" rid="T1">Table&#x20;1</xref>). 1024 consecutive slices from the training set were used with progressively fewer data by eliminating slices from the bottom (i.e.,&#x20;closer to the validation set, see <xref ref-type="sec" rid="s2-1-4">Section 2.1.4</xref>) of the stack. All the models are evaluated with the validation&#x20;set.</p>
</caption>
<graphic xlink:href="fmats-08-761229-g010.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<title>4 Discussions</title>
<sec id="s4-1">
<title>4.1 Overview</title>
<p>Our models achieved, qualitatively, very satisfactory results from a Materials Science application point of view, with 87% of mean class-wise Jaccard index and an F1-score macro average of 92.4<italic>%</italic> (<xref ref-type="sec" rid="s11">Supplementary Material</xref>). We stress the fact that these results were achieved without any strategy to compensate the (heavy) class imbalance (82.3% of the voxels belong to the class matrix); they may be further improved using, for instance, re-sampling strategies (<xref ref-type="bibr" rid="B3">Ando and Huang, 2017</xref>; <xref ref-type="bibr" rid="B23">Pouyanfar et&#x20;al., 2018</xref>), class-balanced loss functions (<xref ref-type="bibr" rid="B10">Cao et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B15">Khan et&#x20;al., 2019</xref>), or self-supervised pre-training (<xref ref-type="bibr" rid="B36">Yang and Xu, 2020</xref>).</p>
<p>The results obtained with another specimen (the Crack volume), thus with slight variations in the acquisition conditions, were way faster than the manual process while keeping good quality&#x2014;inspection by an expert showed no relevant error in the segmentation as compared to a human-made one. The fracture was mostly, and correctly, segmented as porosity without retraining the model, showing its capacity to generalize&#x2014;an important feature for its practical use, although some misclassified regions can be seen as holes (missing pieces) in the fracture&#x2019;s surface (<xref ref-type="fig" rid="F5">Figure&#x20;5</xref>).</p>
<p>Moreover, the processing time achieved (32&#xa0;min) is indeed a promising prospect compared to classic approaches.</p>
</sec>
<sec id="s4-2">
<title>4.2 Segmentation Errors</title>
<p>As highlighted in <xref ref-type="fig" rid="F4">Figure&#x20;4</xref>, the 2D model&#x2019;s mistakes (in red), are mostly on the interfaces of the phases, which are fairly comparable to a human annotator&#x2019;s. We (informally) estimate that they are in the error margin because, in some regions, there is no clear definition of the phases&#x2019; limits. The fibers may show smooth, blurred phase boundaries with the matrix, while some pores are smaller than the image resolution.</p>
<p>Another issue is the loss of information (all-zero regions) in some rings (e.g., <xref ref-type="fig" rid="F3">Figure&#x20;3</xref>). In such cases, even though one could deduce that there is indeed a pore, it is practically impossible to draw a well-defined interface.</p>
<p>Finally, we reiterate that the ground truth remains slightly imperfect despite our efforts to mitigate these issues. For instance, in <xref ref-type="fig" rid="F4">Figure&#x20;4</xref> we see, inside the C-like shaped porosity, a blue region, meaning that it was &#x201c;correctly&#x201d; segmented as a fiber&#x2014;yet, there is no fiber in&#x20;it.</p>
</sec>
<sec id="s4-3">
<title>4.3 Model Variations</title>
<p>
<xref ref-type="fig" rid="F10">Figure&#x20;10</xref> confirms that the porosity class is harder to detect. Although, the qualitative results are reasonable, and we underline that the Jaccard index is more sensitive on underrepresented classes because the size of the union will always be smaller (see Equations in the <xref ref-type="sec" rid="s11">Supplementary Material</xref>).</p>
<p>Contrarily to our expectations, <xref ref-type="fig" rid="F8">Figure&#x20;8</xref> shows that the 2D model performed systematically better than the 3D (albeit the difference is admittedly small). We expected the 3D model to perform better because the morphology of the objects in the image are naturally three-dimensional; besides, other work (<xref ref-type="bibr" rid="B13">Furat et&#x20;al., 2019</xref>) have obtained better results in binary segmentation problems. We raise three hypotheses about this result: 1) the set of hyperparameters is not optimal, 2) the performance metric is biased because the annotation process uses a 2D algorithm, and 3) the 3D models suffer more from overfitting because they have higher variance (more parameters, see <xref ref-type="fig" rid="F8">Figure&#x20;8</xref>).</p>
</sec>
<sec id="s4-4">
<title>4.4 Model Ablation</title>
<p>
<xref ref-type="fig" rid="F9">Figure&#x20;9</xref> contains a few interesting findings about the hyperparameters of the Modular U-Net:<list list-type="simple">
<list-item>
<p>1. using learnable up/down-sampling operations indeed gives more flexibility to the model, improving its performance compared to &#x201c;rigid&#x201d; (not learnable) operations;</p>
</list-item>
<list-item>
<p>2. separable convolutions slightly hurt the performance, but it reduces the number of parameters by&#x20;60%;</p>
</list-item>
<list-item>
<p>3. decreasing the U-depth, therefore shrinking the receptive field, improved the performance while reducing 75<italic>%</italic> of the model size; on the other hand, increasing the depth had the opposite effect, multiplying the model size by four, while degrading the performance;</p>
</list-item>
<list-item>
<p>4. batch normalization is essential for the training&#x2014;notice that the version without batch normalization is out of scale in <xref ref-type="fig" rid="F9">Figure&#x20;9</xref>, and its performance corresponds to the ZeroOC model (see the <xref ref-type="sec" rid="s11">Supplementary Material</xref>);</p>
</list-item>
</list>
</p>
<p>Model depth (item 3): this finding gives a valuable information for our future work because using shallower models require less memory (i.e.,&#x20;bigger crops can be processed at once), making it possible to accelerate the processing time. We hypothesize that the necessary receptive field for image segmentation is smaller than the model with U-depth three. Therefore, a spatially bigger input captures irrelevant, spurious context to the classification.</p>
</sec>
<sec id="s4-5">
<title>4.5 Learning Curve</title>
<p>
<xref ref-type="fig" rid="F10">Figure&#x20;10</xref> highlights a very promising finding in our results: very little annotation is necessary to train a reasonable model. The 2D Modular U-Net was capable of learning with only 1% of the train dataset (13 z-slices) without any smart strategy to select the slices (i.e.,&#x20;one could pick sparsely separated slices to achieve more data variability).</p>
<p>Furthermore, even a single slice was sufficient to achieve a reasonable performance. Notice that the Jaccard index of the porosity is still much higher than its baseline even though this class represents 0.5<italic>%</italic> of the voxels. On the other hand, this is possibly because it is an &#x201c;easier&#x201d; class, as its gray level values are concetrated around zero (see <xref ref-type="fig" rid="F7">Figure&#x20;7</xref>).</p>
<p>This result is of great interest from an application point of view because the annotation process is an important bottleneck to take our method to a practical application. Notice that the data was semi-automatically annotated with other algorithms (Seeded Region Growing and adhoc standard image transformations) and still contain imperfections (e.g., <xref ref-type="fig" rid="F4">Figure&#x20;4</xref>), so the results can improve with a more refined methodology. In other words, <xref ref-type="fig" rid="F10">Figure&#x20;10</xref> tells us that one could obtain a preliminary segmentation by annotating a single slice of this material, then iterate by refining the model&#x2019;s results in order to obtain a few tens of slices, which would provide usable results.</p>
</sec>
</sec>
<sec id="s5">
<title>5 Conclusion</title>
<p>A dataset of XCT images of glass fiber-reinforced Polyamide 66, a three-phase composite material, was presented and segmented with U-Nets; requiring the annotation of only a few tomography slices, results show a promising venue to automate processing pipelines for synchrotron&#x20;XCT.</p>
<p>We proposed the Modular U-Net, a reinterpretation of its precursor that provides a more abstract, conceptually more compact representation of this family of neural network architectures by identifying three essential blocks of a U-Net. As many variations of U-Net have been presented in the literature, we expect this re-description of the model to facilitate the communication of their implementation details. Three variants of the Modular U-Net (2D, 2.5D, and 3D) were compared, showing that all three were capable of learning the patterns from a human-made annotation, although the 2D version performed better than the others.</p>
<p>An ablation study of the 2D Modular U-Net provided insights about its components revealing that we might further accelerate the processing with smaller models, and its learning curve revealed that only 13 annotated slices are necessary to achieve our results&#x2014;in addition, a single slice was sufficient to obtain preliminary results. Finally, we note that further improvements could be achieved by adding other techniques on top of our work, for instance refining the annotations, compensating the data imbalance, and using more advanced architectures like UNet&#x2b;&#x2b; (<xref ref-type="bibr" rid="B38">Zhou et&#x20;al., 2018</xref>), which adds new, more complex, skip connections to U-Net.</p>
</sec>
</body>
<back>
<sec id="s6">
<title>Data Availability Statement</title>
<p>The datasets presented in this study can be found in an online repository (<xref ref-type="bibr" rid="B6">Bertoldo et al., 2021a</xref>). The names of the files and their respective volumes mentioned here are summarized in <xref ref-type="table" rid="T4">Table&#x20;4</xref>.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Published 3D volumes: all the data necessary to train and test the models presented in this paper are publicly available on Zenodo (<xref ref-type="bibr" rid="B6">Bertoldo et&#x20;al., 2021a</xref>). The raw files have complementary raw.info files containing metadata (volume dimensions and data type) about its respective volume. Notice that the volumes (a) and (b) contain all the three splits (train, validation, test) together (1900 z-slices), while the volumes (c) and (d) correspond to their last 300&#x20;z-slices. A demo of how to read the data is available on GitHub. The values in the data volumes are integers from 0 to 255, where the former is black and the latter is white. The values in the segmentation volumes (predictions and ground truth) are 0, 1, and 2, which respectively correspond to the phases matrix, fiber, and porosity. The volumes (e) and (f) correspond to the volume in <xref ref-type="fig" rid="F5">Figures 5</xref>, <xref ref-type="fig" rid="F6">6</xref>. <xref ref-type="fig" rid="F3">Figure&#x20;3</xref> was generated in Fiji (<xref ref-type="bibr" rid="B27">Schindelin et&#x20;al., 2012</xref>) with the volume (a). <xref ref-type="fig" rid="F4">Figure&#x20;4</xref> was generated in Avizo with volumes (a) and (d), which is derived from volumes (b) and (c). <xref ref-type="fig" rid="F6">Figure&#x20;6</xref> was generated in Avizo with volumes (e) and (f).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Zip file</th>
<th align="center">Raw file</th>
<th align="center">Description</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="2" align="left">pa66.zip</td>
<td align="left">pa66.raw <sup>(a)</sup>
</td>
<td align="left">Data (gray level image stack) of the Train-Val-Test volume</td>
</tr>
<tr>
<td align="left">pa66.ground_truth.raw <sup>(b)</sup>
</td>
<td align="left">Ground truth segmentation of the Train-Val-Test volume</td>
</tr>
<tr>
<td rowspan="2" align="left">pa66_test.zip</td>
<td align="left">pa66.test.prediction.raw <sup>(c)</sup>
</td>
<td align="left">Segmentation generated by the best 2D model on the test set.</td>
</tr>
<tr>
<td align="left">pa66.test.error_volume.raw <sup>(d)</sup>
</td>
<td align="left">Disagreement between the ground truth and the model&#x2019;s prediction on the test set: 1 means incorrect, 0 means correct</td>
</tr>
<tr>
<td rowspan="2" align="left">crack.zip</td>
<td align="left">crack.raw <sup>(e)</sup>
</td>
<td align="left">Data of the non-annotated volume containing a crack inside</td>
</tr>
<tr>
<td align="left">crack.prediction.raw <sup>(f)</sup>
</td>
<td align="left">Segmentation generated with the best 2D model on the Crack volume</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>JPCB conducted the experiments and wrote this manuscript. ED assisted the code development, helped design the experiment, and reviewed this manuscript. DR reviewed this manuscript. HP acquired the data, helped design the experiment, and reviewed this manuscript.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>JPCB was funded by the BIGM&#xc9;CA research initiative (<ext-link ext-link-type="uri" xlink:href="ttps://bigmeca.minesparis.psl.eu">ttps://bigmeca.minesparis.psl.eu</ext-link>), funded by Safran and MINES ParisTech.</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>JB would like to thank the funding from the BIGM&#xc9;CA. HP would like to acknowledge beam time allocation at the Psich&#xe9; beamline of Synchrotron SOLEIL for the proposal number 20150371.</p>
</ack>
<sec id="s11">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fmats.2021.761229/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fmats.2021.761229/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="DataSheet1.pdf" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<fn-group>
<fn id="FN1">
<label>1</label>
<p>Adam yielded equivalent results but took longer (more epochs) to converge</p>
</fn>
<fn id="FN2">
<label>2</label>
<p>
<ext-link ext-link-type="uri" xlink:href="http://pny.com/nvidia-quadro-p4000">pny.com/nvidia-quadro-p4000</ext-link>
</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Abadi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Agarwal</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Barham</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Brevdo</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Citro</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <source>TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems</source>. <publisher-loc>Savannah, GA</publisher-loc>: <publisher-name>USENIX Association</publisher-name>. </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Adams</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Bischof</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>1994</year>). <article-title>Seeded Region Growing</article-title>. <source>IEEE Trans. Pattern Anal. Machine Intell.</source> <volume>16</volume>, <fpage>641</fpage>&#x2013;<lpage>647</lpage>. <pub-id pub-id-type="doi">10.1109/34.295913</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ando</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>C. Y.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Deep Over-sampling Framework for Classifying Imbalanced Data</article-title>,&#x201d; in <source>Machine Learning and Knowledge Discovery in Databases</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Ceci</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hollm&#xe9;n</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Todorovski</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Vens</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>D&#x17e;eroski</surname>
<given-names>S.</given-names>
</name>
</person-group> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <fpage>770</fpage>&#x2013;<lpage>785</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-71249-9_46</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Badrinarayanan</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Kendall</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Cipolla</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>39</volume>, <fpage>2481</fpage>&#x2013;<lpage>2495</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2016.2644615</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Bengio</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Lecun</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>1997</year>). <source>Convolutional Networks for Images, Speech, and Time-Series</source>. <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://dl.acm.org/doi/10.5555/303568.303704">https://dl.acm.org/doi/10.5555/303568.303704</ext-link>
</comment> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bertoldo</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Decenci&#xe8;re</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Ryckelynck</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Proudhon</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2021a</year>). <article-title>Glass Fiber-Reinforced Polyamide 66&#x20;3d X-ray Computed Tomography Dataset for Deep Learning Segmentation</article-title>. <comment>[Dataset]</comment>. <pub-id pub-id-type="doi">10.5281/zenodo.4587827</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bertoldo</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Decenci&#xe8;re</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Ryckelynck</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Proudhon</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2021b</year>). <article-title>Glass Fiber-Reinforced Polyamide 66&#x20;3d X-ray Computed Tomography Segmentation Segmentation U-Nets</article-title>. <comment>[Dataset]</comment>. <pub-id pub-id-type="doi">10.5281/zenodo.4601560</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Beucher</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>1994</year>). &#x201c;<article-title>Watershed, Hierarchical Segmentation and waterfall Algorithm</article-title>,&#x201d; in <source>Morphology and its Applications to Image Processing</source>. (<publisher-loc>Dordrecht</publisher-loc>: <publisher-name>Springer</publisher-name>), <volume>2</volume>, <fpage>69</fpage>&#x2013;<lpage>76</lpage>. <pub-id pub-id-type="doi">10.1007/978-94-011-1040-2_10</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Cao</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Gaidon</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Ar&#xe9;chiga</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss</article-title>,&#x201d; in <conf-name>Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019</conf-name>, <conf-loc>Vancouver, BC, Canada</conf-loc>, <conf-date>December 8-14, 2019</conf-date>. Editors <person-group person-group-type="editor">
<name>
<surname>Wallach</surname>
<given-names>H. M.</given-names>
</name>
<name>
<surname>Larochelle</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Beygelzimer</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>d&#x2019;Alch&#xe9; Buc</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Fox</surname>
<given-names>E. B.</given-names>
</name>
<name>
<surname>Garnett</surname>
<given-names>R.</given-names>
</name>
</person-group>, <fpage>1565</fpage>&#x2013;<lpage>1576</lpage>. </citation>
</ref>
<ref id="B11">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>&#xc7;i&#xe7;ek</surname>
<given-names>&#xd6;.</given-names>
</name>
<name>
<surname>Abdulkadir</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Lienkamp</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Brox</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Ronneberger</surname>
<given-names>O.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation</article-title>,&#x201d; in <source>Medical Image Computing and Computer-Assisted Intervention (MICCAI)</source>. (<publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>). <pub-id pub-id-type="doi">10.1007/978-3-319-46723-8_49</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Duque-Arias</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Velasco-Forero</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Deschaud</surname>
<given-names>J.-E.</given-names>
</name>
<name>
<surname>Goulette</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Serna</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Decenci&#xe8;re</surname>
<given-names>E.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). &#x201c;<article-title>On Power Jaccard Losses for Semantic Segmentation</article-title>,&#x201d; in <conf-name>VISAPP 2021: 16th International Conference on Computer Vision Theory and Applications</conf-name>, <conf-loc>Vienne (on line), Austria</conf-loc>, <conf-date>February 2021</conf-date>. <pub-id pub-id-type="doi">10.5220/0010304005610568</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Furat</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Neumann</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Petrich</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Weber</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Krill III</surname>
<given-names>C. E.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Machine Learning Techniques for the Segmentation of Tomographic Image Data of Functional Materials</article-title>. <source>Front. Mater.</source> <volume>6</volume>, <fpage>145</fpage>. <pub-id pub-id-type="doi">10.3389/fmats.2019.00145</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hanley</surname>
<given-names>J.&#x20;A.</given-names>
</name>
<name>
<surname>McNeil</surname>
<given-names>B. J.</given-names>
</name>
</person-group> (<year>1982</year>). <article-title>The Meaning and Use of the Area under a Receiver Operating Characteristic (ROC) Curve</article-title>. <source>Radiology</source> <volume>143</volume>, <fpage>29</fpage>&#x2013;<lpage>36</lpage>. <pub-id pub-id-type="doi">10.1148/radiology.143.1.7063747</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Khan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Hayat</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zamir</surname>
<given-names>S. W.</given-names>
</name>
<name>
<surname>Shen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Shao</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Striking the Right Balance with Uncertainty</article-title>,&#x201d; in <conf-name>2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>, <conf-loc>Long Beach, CA</conf-loc>, <conf-date>June 2019</conf-date>, <fpage>103</fpage>&#x2013;<lpage>112</lpage>. <pub-id pub-id-type="doi">10.1109/cvpr.2019.00019</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kingma</surname>
<given-names>D. P.</given-names>
</name>
<name>
<surname>Ba</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Adam: A Method for Stochastic Optimization</article-title>,&#x201d; in <conf-name>3rd International Conference on Learning Representations, ICLR 2015</conf-name>, <conf-loc>San Diego, CA, USA</conf-loc>, <conf-date>May 7-9, 2015</conf-date>. Editors <person-group person-group-type="editor">
<name>
<surname>Bengio</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>LeCun</surname>
<given-names>Y.</given-names>
</name>
</person-group>. </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Maire</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Le Bourlot</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Adrien</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Mortensen</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Mokso</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>20&#x20;Hz X-ray Tomography during an <italic>In Situ</italic> Tensile Test</article-title>. <source>Int. J.&#x20;Fract.</source> <volume>200</volume>, <fpage>3</fpage>&#x2013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1007/s10704-016-0077-y</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Maire</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Withers</surname>
<given-names>P. J.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Quantitative X-ray Tomography</article-title>. <source>Int. Mater. Rev.</source> <volume>59</volume>, <fpage>1</fpage>&#x2013;<lpage>43</lpage>. <pub-id pub-id-type="doi">10.1179/1743280413Y.0000000023</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mirone</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Brun</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Gouillart</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Tafforeau</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Kieffer</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>The PyHST2 Hybrid Distributed Code for High Speed Tomographic Reconstruction with Iterative Reconstruction and A Priori Knowledge Capabilities</article-title>. <source>Nucl. Instr. Methods Phys. Res. Section B: Beam Interactions Mater. Atoms</source> <volume>324</volume>, <fpage>41</fpage>&#x2013;<lpage>48</lpage>. <pub-id pub-id-type="doi">10.1016/j.nimb.2013.09.030</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Oktay</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Schlemper</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Folgoc</surname>
<given-names>L. L.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Heinrich</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Misawa</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <source>Attention U-Net: Learning where to Look for the Pancreas</source>. <publisher-name>arXiv</publisher-name>. <comment>arXiv:1804.03999 [cs] [Preprint]</comment>. </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pacchioni</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>An Upgrade to a Bright Future</article-title>. <source>Nat. Rev. Phys.</source> <volume>1</volume>, <fpage>100</fpage>&#x2013;<lpage>101</lpage>. <pub-id pub-id-type="doi">10.1038/s42254-019-0019-5</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Paganin</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Mayo</surname>
<given-names>S. C.</given-names>
</name>
<name>
<surname>Gureyev</surname>
<given-names>T. E.</given-names>
</name>
<name>
<surname>Miller</surname>
<given-names>P. R.</given-names>
</name>
<name>
<surname>Wilkins</surname>
<given-names>S. W.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Simultaneous Phase and Amplitude Extraction from a Single Defocused Image of a Homogeneous Object</article-title>. <source>J.&#x20;Microsc.</source> <volume>206</volume>, <fpage>33</fpage>&#x2013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1046/j.1365-2818.2002.01010.x</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pouyanfar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Tao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Mohan</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Tian</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Kaseb</surname>
<given-names>A. S.</given-names>
</name>
<name>
<surname>Gauen</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). &#x201c;<article-title>Dynamic Sampling in Convolutional Neural Networks for Imbalanced Data Classification</article-title>,&#x201d; in <conf-name>2018 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR)</conf-name>, <conf-loc>Miami, Florida</conf-loc>, <conf-date>April, 2018</conf-date>, <fpage>112</fpage>&#x2013;<lpage>117</lpage>. <pub-id pub-id-type="doi">10.1109/MIPR.2018.00027</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Qin</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Dehghan</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zaiane</surname>
<given-names>O. R.</given-names>
</name>
<name>
<surname>Jagersand</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>U2-Net: Going Deeper with Nested U-Structure for Salient Object Detection</article-title>. <source>Pattern Recognit.</source> <volume>106</volume>, <fpage>107404</fpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2020.107404</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ronneberger</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Fischer</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Brox</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>U-net: Convolutional Networks for Biomedical Image Segmentation</article-title>,&#x201d; in <source>Medical Image Computing and Computer-Assisted Intervention</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Navab</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Hornegger</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wells</surname>
<given-names>W. M.</given-names>
</name>
<name>
<surname>Frangi</surname>
<given-names>A. F.</given-names>
</name>
</person-group> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <fpage>234</fpage>&#x2013;<lpage>241</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-24574-4_28</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rosenblatt</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>1958</year>). <article-title>The Perceptron: a Probabilistic Model for Information Storage and Organization in the Brain</article-title>. <source>Psychol. Rev.</source> <volume>65</volume> (<issue>6</issue>), <fpage>386</fpage>&#x2013;<lpage>408</lpage>. <pub-id pub-id-type="doi">10.1037/h0042519</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schindelin</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Arganda-Carreras</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Frise</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Kaynig</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Longair</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Pietzsch</surname>
<given-names>T.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>Fiji: An Open-Source Platform for Biological-Image Analysis</article-title>. <source>Nat. Methods</source> <volume>9</volume>, <fpage>676</fpage>&#x2013;<lpage>682</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.2019</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schneider</surname>
<given-names>C. A.</given-names>
</name>
<name>
<surname>Rasband</surname>
<given-names>W. S.</given-names>
</name>
<name>
<surname>Eliceiri</surname>
<given-names>K. W.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>NIH Image to ImageJ: 25&#x20;Years of Image Analysis</article-title>. <source>Nat. Methods</source> <volume>9</volume>, <fpage>671</fpage>&#x2013;<lpage>675</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.2089</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shashank Kaira</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>De Andrade</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>De Carlo</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Scullin</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Gursoy</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Automated Correlative Segmentation of Large Transmission X-ray Microscopy (TXM) Tomograms Using Deep Learning</article-title>. <source>Mater. Charact.</source> <volume>142</volume>, <fpage>203</fpage>&#x2013;<lpage>210</lpage>. <pub-id pub-id-type="doi">10.1016/j.matchar.2018.05.053</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shelhamer</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Long</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Darrell</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Fully Convolutional Networks for Semantic Segmentation</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>39</volume>, <fpage>640</fpage>&#x2013;<lpage>651</lpage>. <pub-id pub-id-type="doi">10.1109/tpami.2016.2572683</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shuai</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Phillion</surname>
<given-names>A. B.</given-names>
</name>
<name>
<surname>Callaghan</surname>
<given-names>M. D.</given-names>
</name>
<name>
<surname>Jing</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>P. D.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Fast Synchrotron X-ray Tomographic Quantification of Dendrite Evolution during the Solidification of Mg Sn Alloys</article-title>. <source>Acta Material.</source> <volume>118</volume>, <fpage>260</fpage>&#x2013;<lpage>269</lpage>. <pub-id pub-id-type="doi">10.1016/j.actamat.2016.07.047</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Stan</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Thompson</surname>
<given-names>Z. T.</given-names>
</name>
<name>
<surname>Voorhees</surname>
<given-names>P. W.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Optimizing Convolutional Neural Networks to Perform Semantic Segmentation on Large Materials Imaging Datasets: X-ray Tomography and Serial Sectioning</article-title>. <source>Mater. Charact.</source> <volume>160</volume>, <fpage>110119</fpage>. <pub-id pub-id-type="doi">10.1016/j.matchar.2020.110119</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Stoller</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Ewert</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Dixon</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Wave-U-Net: A Multi-Scale Neural Network for End-To-End Audio Source Separation</article-title>,&#x201d; in <conf-name>Proceedings of the 19th International Society for Music Information Retrieval Conference (ISMIR 2018)</conf-name>, <conf-loc>Calgary, AB</conf-loc>, <conf-date>April, 2018</conf-date>. </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Strohmann</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Bugelnig</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Breitbarth</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Wilde</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Steffens</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Germann</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Semantic Segmentation of Synchrotron Tomography of Multiphase Al-Si Alloys Using a Convolutional Neural Network with a Pixel-wise Weighted Loss Function</article-title>. <source>Sci. Rep.</source> <volume>9</volume>, <fpage>1</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1038/s41598-019-56008-7</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Withers</surname>
<given-names>P. J.</given-names>
</name>
<name>
<surname>Bouman</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Carmignato</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Cnudde</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Grimaldi</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Hagen</surname>
<given-names>C. K.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>X-Ray Computed Tomography</article-title>. <source>Nat. Rev. Methods Primers</source> <volume>1</volume>, <fpage>1</fpage>&#x2013;<lpage>21</lpage>. <pub-id pub-id-type="doi">10.1038/s43586-021-00015-4</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Rethinking the Value of Labels for Improving Class-Imbalanced Learning</article-title>,&#x201d; in <source>Advances in Neural Information Processing Systems</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Larochelle</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Ranzato</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hadsell</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Balcan</surname>
<given-names>M. F.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>H.</given-names>
</name>
</person-group> (<publisher-name>Curran Associates, Inc.</publisher-name>), <volume>Vol. 33</volume>, <fpage>19290</fpage>&#x2013;<lpage>19301</lpage>. <ext-link ext-link-type="uri" xlink:href="http://file:///Users/daniellehiraldo/Downloads/Best-Practices-AIAN-Data-Collection-20200826.pdf">file:///Users/daniellehiraldo/Downloads/Best-Practices-AIAN-Data-Collection-20200826.pdf</ext-link>. </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Road Extraction by Deep Residual U-Net</article-title>. <source>IEEE Geosci. Remote Sensing Lett.</source> <volume>15</volume>, <fpage>749</fpage>&#x2013;<lpage>753</lpage>. <pub-id pub-id-type="doi">10.1109/lgrs.2018.2802944</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Rahman Siddiquee</surname>
<given-names>M. M.</given-names>
</name>
<name>
<surname>Tajbakhsh</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Liang</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>UNet&#x2b;&#x2b;: A Nested U-Net Architecture for Medical Image Segmentation</article-title>,&#x201d; in <source>Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Stoyanov</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Taylor</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Carneiro</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Syeda-Mahmood</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Martel</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Maier-Hein</surname>
<given-names>L.</given-names>
</name>
<etal/>
</person-group> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <fpage>3</fpage>&#x2013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-00889-5_1</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>