<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Astron. Space Sci.</journal-id>
<journal-title>Frontiers in Astronomy and Space Sciences</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Astron. Space Sci.</abbrev-journal-title>
<issn pub-type="epub">2296-987X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1197358</article-id>
<article-id pub-id-type="doi">10.3389/fspas.2023.1197358</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Astronomy and Space Sciences</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Efficient galaxy classification through pretraining</article-title>
<alt-title alt-title-type="left-running-head">Schneider et al.</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fspas.2023.1197358">10.3389/fspas.2023.1197358</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Schneider</surname>
<given-names>Jesse</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2265456/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Stenning</surname>
<given-names>David C.</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/2000667/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Elliott</surname>
<given-names>Lloyd T.</given-names>
</name>
</contrib>
</contrib-group>
<aff>
<institution>Department of Statistics and Actuarial Science</institution>, <institution>Simon Fraser University</institution>, <addr-line>Burnaby</addr-line>, <addr-line>BC</addr-line>, <country>Canada</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/223813/overview">Didier Fraix-Burnet</ext-link>, UMR5274 Institut de Plan&#xe9;tologie et d&#x27;Astrophysique de Grenoble (IPAG), France</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2274502/overview">Sarvesh Gharat</ext-link>, Indian Institute of Technology Bombay, India</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1698845/overview">Clement Nyirenda</ext-link>, University of the Western Cape, South Africa</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Jesse Schneider, <email>jesse_schneider@sfu.ca</email>
</corresp>
</author-notes>
<pub-date pub-type="epub">
<day>10</day>
<month>08</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>10</volume>
<elocation-id>1197358</elocation-id>
<history>
<date date-type="received">
<day>30</day>
<month>03</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>13</day>
<month>07</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Schneider, Stenning and Elliott.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Schneider, Stenning and Elliott</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Deep learning has increasingly been applied to supervised learning tasks in astronomy, such as classifying images of galaxies based on their apparent shape (i.e., galaxy morphology classification) to gain insight regarding the evolution of galaxies. In this work, we examine the effect of pretraining on the performance of the classical AlexNet convolutional neural network (CNN) in classifying images of 14,034 galaxies from the Sloan Digital Sky Survey Data Release 4. Pretraining involves designing and training CNNs on large labeled image datasets unrelated to astronomy, which takes advantage of the vast amounts of such data available compared to the relatively small amount of labeled galaxy images. We show a statistically significant benefit of using pretraining, both in terms of improved overall classification success and reduced computational cost to achieve such performance.</p>
</abstract>
<kwd-group>
<kwd>convolutional neural networks</kwd>
<kwd>machine learning</kwd>
<kwd>galaxy morphology</kwd>
<kwd>astrostatistics</kwd>
<kwd>transfer learning</kwd>
<kwd>pretraining</kwd>
</kwd-group>
<contract-num rid="cn001">RGPIN/03985-2021 RGPIN/05484-2019 DGECR/00118-2019</contract-num>
<contract-sponsor id="cn001">Natural Sciences and Engineering Research Council of Canada<named-content content-type="fundref-id">10.13039/501100000038</named-content>
</contract-sponsor>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Astrostatistics</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Convolutional neural networks (CNNs) are a type of deep learning that is particularly well suited to computer vision tasks (<xref ref-type="bibr" rid="B4">Aggarwal, 2018</xref>). Originally inspired by research on the visual cortex by neurophysiologists Hubel and Wiesel in the mid-20th century (<xref ref-type="bibr" rid="B19">Hubel and Wiesel, 1959</xref>; <xref ref-type="bibr" rid="B4">Aggarwal, 2018</xref>), CNNs have been successfully applied to a variety of computer vision tasks, such as image classification (<xref ref-type="bibr" rid="B22">Krizhevsky et al., 2012</xref>), facial recognition (<xref ref-type="bibr" rid="B36">Taigman et al., 2014</xref>), and classification of various forms of interstitial lung disease (<xref ref-type="bibr" rid="B23">Li et al., 2014</xref>). CNNs have also been applied to tasks outside of computer vision, such as forecasting prices in financial stock markets (<xref ref-type="bibr" rid="B38">Tsantekidis et al., 2017</xref>).</p>
<p>As astronomy and astrophysics increasingly rely on large sets of image data, CNNs have increasingly been used to tackle a variety of interesting astronomical and astrophysical problems. These include identifying gravitational lenses (<xref ref-type="bibr" rid="B10">Davies et al., 2019</xref>), identifying contamination in astronomical images, e.g., by cosmic rays and diffraction spikes (<xref ref-type="bibr" rid="B27">Paillassa et al., 2020</xref>), and for supernovae detection (<xref ref-type="bibr" rid="B7">Cabrera-Vives et al., 2016</xref>). One particular task for which CNNs have proved successful, which will be discussed in more depth below, is galaxy morphology classification as in <xref ref-type="bibr" rid="B8">Cavanagh et al. (2021)</xref>. We will therefore use galaxy morphology classification to explore how <italic>pretraining</italic> [a type of transfer learning (<xref ref-type="bibr" rid="B37">Tan et al., 2018</xref>; <xref ref-type="bibr" rid="B30">Ribani and Marengoni, 2019</xref>)] can potentially benefit many tasks in astronomy and astrophysics that use CNNs.</p>
<p>When applied to a particular task, CNNs (and neural networks in general) can be trained from scratch for the given task, or instead we can use a pretrained network. A pretrained CNN is one which has already been trained on a separate data set prior to its application to the given task (<xref ref-type="bibr" rid="B4">Aggarwal, 2018</xref>). The training of neural networks is, in general, an energy-intensive activity (<xref ref-type="bibr" rid="B35">Strubell et al., 2020</xref>; <xref ref-type="bibr" rid="B6">Borowiec et al., 2022</xref>), and research is ongoing to quantify and reduce energy use [e.g., (<xref ref-type="bibr" rid="B42">Yang et al., 2017</xref>; <xref ref-type="bibr" rid="B13">Garc&#xed;a-Mart&#xed;n et al., 2019</xref>)]. Training neural networks from scratch can require more training and therefore greater expense and resource usage compared to using a pretrained network. Furthermore, this requirement for a large amount of training presupposes the availability of a sufficient amount of data within the intended domain for the desired amount of training. Thus, data availability itself can be a limitation which may preclude the possibility of a large amount of training being performed from scratch.</p>
<p>In astronomy, obtaining large, labelled training data sets is expensive, impractical, or both. As a result of the demands required to train deep learning models from scratch, pretraining may be an attractive alternative for astronomy. In recent years, transfer learning has been adopted for classification tasks in astronomy and astrophysics involving galaxy morphologies (<xref ref-type="bibr" rid="B12">Dom&#xed;nguez S&#xe1;nchez et al., 2018</xref>), variable stars (<xref ref-type="bibr" rid="B21">Kim et al., 2021</xref>), and star clusters (<xref ref-type="bibr" rid="B41">Wei et al., 2020</xref>). However, the learning that is &#x201c;transferred&#x201d; in these cases is between different astronomical surveys. That is, a classifier is trained using data from one survey, and deployed, perhaps with modification, on test data arising from a different survey. For example, a model may be trained on Sloan Digital Sky Survey images and deployed on Dark Energy Survey images (<xref ref-type="bibr" rid="B12">Dom&#xed;nguez S&#xe1;nchez et al., 2018</xref>).</p>
<p>An alternative type of transfer learning, which we refer to specifically as pretraining hereafter, involves the practice of training a neural network on another <italic>unrelated</italic> data set before applying the neural network to the particular data set of interest. This means that pretraining, using our definition, can exploit the vast effort undertaken to design CNNs for classifying large volumes of natural (everyday) images. Specifically, we will use a CNN trained on millions of non-astronomical images that comprise the ImageNet database (<xref ref-type="bibr" rid="B11">Deng et al., 2009</xref>); this CNN is known as AlexNet (<xref ref-type="bibr" rid="B22">Krizhevsky et al., 2012</xref>). We will demonstrate, through a series of numerical experiments, that for the task of galaxy morphology classification, a pretrained AlexNet outperforms an architecturally identical CNN that is trained from scratch using only galaxy morphology images. Transfer learning using training on a large data set of natural images has been explored in the context of analysis of data from the Laser Interferometer Gravitational-Wave Observatory (<xref ref-type="bibr" rid="B14">George et al., 2018</xref>) and in galaxy merger detection (<xref ref-type="bibr" rid="B2">Ackermann et al., 2018</xref>). However, as far as we are aware, such transfer learning has not been explored in the context of galaxy morphology classification, which is the focus of the present paper.</p>
<p>The rest of this paper explores the utility of pretraining in the application of a CNN to galaxy morphology image classification. This domain represents a potential use case for a pretrained network due to the expense and difficulty of gathering and labeling images of galaxies (<xref ref-type="bibr" rid="B8">Cavanagh et al., 2021</xref>). We begin in <xref ref-type="sec" rid="s2">Section 2</xref> with an overview of galaxy morphology classification and the data we will use for our experiments. In <xref ref-type="sec" rid="s3">Section 3</xref> we describe CNNs, including the particular CNN used in this paper, as well as data preparation procedures and tooling. Our numerical experiments and results are detailed in <xref ref-type="sec" rid="s4">Section 4</xref>. We discuss and summarize our contributions in <xref ref-type="sec" rid="s5">Section 5</xref>, and also discuss directions for future research.</p>
<p>All code and materials necessary to reproduce our work can be found at: <ext-link ext-link-type="uri" xlink:href="https://github.com/jsa378/01_masters">https://github.com/jsa378/01_masters</ext-link>.</p>
</sec>
<sec id="s2">
<title>2 Galaxy morphology classification and data</title>
<p>Images of galaxies are captured using either earthbound equipment or spacecraft. Traditionally, the images are classified by groups of experts who examine each image and come to a consensus regarding its classification (<xref ref-type="bibr" rid="B8">Cavanagh et al., 2021</xref>). In order to speed up the classification of images, other strategies have been used, such as the recruitment of enthusiastic amateurs (<xref ref-type="bibr" rid="B24">Lintott et al., 2010</xref>), and various automated classification techniques (<xref ref-type="bibr" rid="B9">Cheng et al., 2020</xref>).</p>
<p>CNNs in particular and deep learning more generally are well suited to tackle challenges in galaxy morphology classification, and as such there is a large body of literature devoted to such aims. We describe a few of these below, but note that the list is non-exhaustive.<list list-type="simple">
<list-item>
<p>&#x2022; <xref ref-type="bibr" rid="B9">Cheng et al. (2020)</xref> use CNNs and other machine/deep learning techniques to classify galaxy images from the Sloan Digital Sky Survey Data Release 7 into two categories (i.e., they perform two-way classification)&#x2014;elliptical and spiral&#x2014;reporting a best overall accuracy of over 99%, achieved with a CNN.</p>
</list-item>
<list-item>
<p>&#x2022; <xref ref-type="bibr" rid="B5">Barchi et al. (2020)</xref> combine data from Galaxy Zoo 1 (<xref ref-type="bibr" rid="B24">Lintott et al., 2010</xref>) and the Dark Energy Survey (<xref ref-type="bibr" rid="B1">Abbott et al., 2018</xref>) when performing galaxy morphology classification. Using deep learning, they also achieve an overall accuracy of over 99% for two-way classification. However, when a third class (barred galaxies) is added, overall accuracy drops to 82%.</p>
</list-item>
<list-item>
<p>&#x2022; <xref ref-type="bibr" rid="B15">Gharat and Dandawate (2022)</xref> also use Galaxy Zoo data, but adopt an extended Hubble tuning fork classification scheme, which places galaxies into 10 categories. Using a CNN, they achieve a best overall accuracy of about 85%.</p>
</list-item>
<list-item>
<p>&#x2022; <xref ref-type="bibr" rid="B8">Cavanagh et al. (2021)</xref> use galaxy images from Sloan Digital Sky Survey Data Release 4 (SDSS DR4) (<xref ref-type="bibr" rid="B43">York et al., 2000</xref>; <xref ref-type="bibr" rid="B34">Stoughton et al., 2002</xref>; <xref ref-type="bibr" rid="B3">Adelman-McCarthy et al., 2006</xref>), and compare the performance of different CNNs on three-way and four-way classification, reporting best overall accuracy results of 83% and 81%, respectively. The four-way classification task is especially interesting and important because the fourth class is for &#x201c;irregular and miscellaneous&#x201d; galaxies that do not belong to one of the other three classes.</p>
</list-item>
</list>
</p>
<p>In the near future, an expected deluge of data obtained by new spacecraft such as the European Space Agency&#x2019;s <italic>Euclid</italic> will overwhelm available resources for classification by humans (<xref ref-type="bibr" rid="B32">Silva et al., 2018</xref>). This adds urgency to the search for accurate, rapid, automated classification techniques, such as those cited above, among others. However, the galaxy image data currently available for training deep learning models remains relatively small. Following <xref ref-type="bibr" rid="B8">Cavanagh et al. (2021)</xref>, for this work we used 14,034 (labeled) galaxy images from the SDSS DR4 (<xref ref-type="bibr" rid="B43">York et al., 2000</xref>; <xref ref-type="bibr" rid="B34">Stoughton et al., 2002</xref>; <xref ref-type="bibr" rid="B3">Adelman-McCarthy et al., 2006</xref>); the data are described in detail in <xref ref-type="bibr" rid="B26">Nair and Abraham (2010)</xref>. Each galaxy is labeled according to its morphology, or shape, as belonging to the class of:<list list-type="simple">
<list-item>
<p>1) <italic>elliptical galaxies</italic>, having a smooth, diffuse, and elliptical shape;</p>
</list-item>
<list-item>
<p>2) <italic>spiral galaxies</italic>, disk-like in appearance and with spiral arms;</p>
</list-item>
<list-item>
<p>3) <italic>lenticular galaxies</italic>, an intermediate class of galaxies between the elliptical and spiral categories; or</p>
</list-item>
<list-item>
<p>4) <italic>irregular &#x2b; miscellaneous (Irr &#x2b; Misc) galaxies</italic>, not meeting the membership criteria for any of the above three categories.</p>
</list-item>
</list>
</p>
<p>An example of each type of galaxy, taken from the set of 14,034, is presented in <xref ref-type="fig" rid="F1">Figure 1</xref>. The 14,034 galaxy morphology images varied in size, but were often quite small&#x2014;around 100 &#xd7; 100 pixels, or 0.01 megapixel. The class breakdown of the images is as follows:<list list-type="simple">
<list-item>
<p>1) 2,738 images of elliptical galaxies (19.4% of the total),</p>
</list-item>
<list-item>
<p>2) 7,708 images of spiral galaxies (54.9% of the total),</p>
</list-item>
<list-item>
<p>3) 3,215 images of lenticular galaxies (22.9% of the total), and</p>
</list-item>
<list-item>
<p>4) 373 images of Irr &#x2b; Misc galaxies (2.7% of the total).</p>
</list-item>
</list>While this is clearly a highly imbalanced data set as governed by the distribution of galaxies in the regions imaged, we did not attempt to correct for these imbalances because for this work we are only concerned with demonstrating the benefit of pretraining. Further discussion of class differences is in <xref ref-type="sec" rid="s4-1">Subsection 4.1</xref>.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Examples of the four categories of galaxy from the Sloan Digital Sky Survey Data Release 4 (<xref ref-type="bibr" rid="B26">Nair and Abraham, 2010</xref>). <bold>(A)</bold> Elliptical galaxy. <bold>(B)</bold> Lenticular galaxy. <bold>(C)</bold> Spiral galaxy. <bold>(D)</bold> Irr &#x2b; Misc galaxy.</p>
</caption>
<graphic xlink:href="fspas-10-1197358-g001.tif"/>
</fig>
<p>Although <xref ref-type="bibr" rid="B8">Cavanagh et al. (2021)</xref> initially performed three-way classification, excluding Irr &#x2b; Misc galaxies, we will limit the current work to considering the (more challenging) four-way classification task. The reason for this is two-fold: (1): we do not want to pre-filter Irr &#x2b; Misc galaxies as they will undoubtedly be present in the future survey data, and (2) it is precisely for Irr &#x2b; Misc galaxies that we notice the greatest benefit of pretraining, perhaps due to the relative lack of training examples; we will demonstrate and discuss this in more detail in <xref ref-type="sec" rid="s4">Sections 4</xref>, <xref ref-type="sec" rid="s5">5</xref>.<xref ref-type="fn" rid="fn1">
<sup>1</sup>
</xref>
</p>
</sec>
<sec id="s3">
<title>3 Methods and data preparation</title>
<sec id="s3-1">
<title>3.1 Convolutional neural networks</title>
<sec id="s3-1-1">
<title>3.1.1 Introduction to CNNs</title>
<p>CNNs are a type of neural network that are well suited to image data (<xref ref-type="bibr" rid="B16">Goodfellow et al., 2016</xref>). They are so named because of the &#x201c;convolution&#x201d; operations applied within the network, although strictly speaking these operations are cross-correlations (<xref ref-type="bibr" rid="B16">Goodfellow et al., 2016</xref>). CNNs are perhaps the archetypal example of biologically-inspired artificial intelligence, because their conception was influenced by exploration of the visual cortex in the mid-20th century (<xref ref-type="bibr" rid="B4">Aggarwal, 2018</xref>).</p>
<p>Although formal mathematical justification for CNNs is lacking, the common explanation is that CNNs function by detecting relatively crude features of an input image, such as lines, in the early layers of the network, and superimpose these features into progressively more complex features in later layers (<xref ref-type="bibr" rid="B4">Aggarwal, 2018</xref>).</p>
</sec>
<sec id="s3-1-2">
<title>3.1.2 CNN operations</title>
<p>Like feedforward neural networks, CNNs consist of an input layer, one or more hidden layers and an output layer. The primary differences are the types of operations that the layers perform. The fundamental principles of neural networks&#x2014;the forwards and backwards phases, and gradient-based optimization&#x2014;also apply to CNNs. Furthermore, training methods consisting of feeding the entire training data set to the network multiple times, each time referred to as an <italic>epoch</italic>, are similar for both types of networks. For brevity, therefore, we will focus on the unique mathematical operations employed in CNNs; these unique operations are convolution and pooling operations. We will also briefly discuss the regularization technique known as dropout, which is commonly used to avoid overfitting deep neural networks.</p>
<sec id="s3-1-2-1">
<title>3.1.2.1 The convolution (cross-correlation)</title>
<p>Consider a color image <bold>I</bold> of size <italic>h</italic> &#xd7; <italic>w</italic> pixels, represented numerically as an array with dimensions <italic>h</italic> &#xd7; <italic>w</italic> &#xd7; 3. (The depth of 3 is for storage of the red, green and blue color values.) The convolution operation involves placing a smaller <italic>h</italic>&#x2032; &#xd7; <italic>w</italic>&#x2032; &#xd7; 3 (<italic>h</italic>&#x2032; &#x2264; <italic>h</italic>, <italic>w</italic>&#x2032; &#x2264; <italic>w</italic>) array <bold>K</bold>, called the <italic>kernel</italic> or <italic>filter</italic>, at all possible positions overlaid on <bold>I</bold> and computing the component-wise dot product between <bold>I</bold> and <bold>K</bold>.<xref ref-type="fn" rid="fn2">
<sup>2</sup>
</xref>
</p>
<p>More formally, the convolution of <italic>h</italic> &#xd7; <italic>w</italic> &#xd7; 3 image <bold>I</bold> (having <italic>i</italic>, <italic>j</italic>, <italic>l</italic> entry <italic>I</italic>
<sub>
<italic>i</italic>,<italic>j</italic>,<italic>l</italic>
</sub>) with <italic>h</italic>&#x2032; &#xd7; <italic>w</italic>&#x2032; &#xd7; 3 kernel <bold>K</bold> (having <italic>i</italic>, <italic>j</italic>, <italic>l</italic> entry <italic>K</italic>
<sub>
<italic>i</italic>,<italic>j</italic>,<italic>l</italic>
</sub>) is a function<disp-formula id="equ1">
<mml:math id="m1">
<mml:mo>&#x2a;</mml:mo>
<mml:mo>:</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>h</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>w</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2192;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="double-struck">R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>h</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#xd7;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
</mml:math>
</disp-formula>defined by<disp-formula id="equ2">
<mml:math id="m2">
<mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi mathvariant="bold">I</mml:mi>
<mml:mo>&#x2a;</mml:mo>
<mml:mi mathvariant="bold">K</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>:</mml:mo>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="true">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:munderover>
<mml:munderover accentunder="false" accent="true">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:munderover>
<mml:munderover accentunder="false" accent="true">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>I</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
<mml:mi>s</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x22c5;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>K</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>.</mml:mo>
</mml:math>
</disp-formula>
</p>
</sec>
<sec id="s3-1-2-2">
<title>3.1.2.2 The max pool</title>
<p>The max pool operation involves a smaller array <bold>P</bold> similar to the kernel <bold>K</bold>, except that <bold>P</bold> has a depth of 1. If <bold>P</bold> has dimensions <italic>p</italic> &#xd7; <italic>q</italic> &#xd7; 1 and acts on a layer <bold>L</bold> having dimensions <italic>h</italic> &#xd7; <italic>w</italic> &#xd7; <italic>d</italic>, then the pooling operation produces a layer <bold>P</bold>(<bold>L</bold>) having dimensions (<italic>h</italic> &#x2212; <italic>p</italic> &#x2b; 1) &#xd7; (<italic>w</italic> &#x2212; <italic>q</italic> &#x2b; 1) &#xd7; <italic>d</italic>. In particular,<disp-formula id="equ3">
<mml:math id="m3">
<mml:mtable class="align-star" columnalign="left">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mi mathvariant="bold">P</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi mathvariant="bold">L</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>s</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mo>:</mml:mo>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>max</mml:mi>
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="bold">L</mml:mi>
<mml:mo>:</mml:mo>
<mml:mi>r</mml:mi>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right"/>
<mml:mtd columnalign="left">
<mml:mspace width="0.5em"/>
<mml:mspace width="1em"/>
<mml:mspace width="1em"/>
<mml:mspace width="1em"/>
<mml:mspace width="1em"/>
<mml:mo>&#x2264;</mml:mo>
<mml:mfenced open="" close="}">
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>r</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
<mml:mi>s</mml:mi>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>&#x2264;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>q</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
</disp-formula>
</p>
</sec>
<sec id="s3-1-2-3">
<title>3.1.2.3 Dropout</title>
<p>Dropout is a regularization technique applied layer-wise during CNN training, which involves training random subsets of the overall network (<xref ref-type="bibr" rid="B18">Hinton et al., 2012</xref>; <xref ref-type="bibr" rid="B33">Srivastava et al., 2014</xref>). If dropout is applied to layer <italic>l</italic> in the network, then each time the network is fed a training image, independent draws from a Bernoulli(<italic>p</italic>) distribution are made to determine which nodes in layer <italic>l</italic> are kept or discarded. The training prediction and backpropagation are then carried out only over the sub-network containing the remaining nodes. During training, repeated samples from the Bernoulli(<italic>p</italic>) distribution are drawn, which means that different subsets of the original network are trained. During testing, dropout is not applied, so the entire network is used. Using dropout significantly reduces network overfitting, and thus improves the network&#x2019;s ability to generalize to the test data (<xref ref-type="bibr" rid="B33">Srivastava et al., 2014</xref>).</p>
</sec>
</sec>
</sec>
<sec id="s3-2">
<title>3.2 AlexNet</title>
<p>The CNN that underpins our work is called <italic>AlexNet</italic> (<xref ref-type="bibr" rid="B22">Krizhevsky et al., 2012</xref>). Although CNNs date back to the 1980s, in 2012 the AlexNet CNN ushered in what is arguably the modern era in computer vision by winning the ImageNet Large Scale Visual Recognition Challenge, thoroughly surpassing past performers and challengers (<xref ref-type="bibr" rid="B4">Aggarwal, 2018</xref>). Since then, CNNs have served as the standard for image classification, as can be seen by the fact that subsequent winners of the ImageNet competition have also been CNNs (<xref ref-type="bibr" rid="B4">Aggarwal, 2018</xref>). Below, a brief summary of the AlexNet architecture is provided. Further details can be found in <xref ref-type="bibr" rid="B22">Krizhevsky et al. (2012)</xref>.</p>
<p>The AlexNet network can be broken into two broad parts: an early convolutional part and a later, more conventional feed-forward part. In the convolutional part, AlexNet takes as input a 224 &#xd7; 224 &#xd7; 3 image and, over a number of layers, applies an increasing number of convolution filters which decrease in height and width. The first convolutional layer applies 64 filters of size 11 &#xd7; 11 &#xd7; 3; a later convolutional layer applies 384 filters of size 3 &#xd7; 3 &#xd7; 256. (As previously mentioned, formal mathematical justification for CNN operation is lacking, but the intuition is that larger numbers of smaller filters in later layers capture more complex features of the input image.) AlexNet uses five convolutional layers in total. After every convolutional layer the ReLU operation is used, and a handful of max-pooling operations are layered in as well. In the feed-forward part, AlexNet contains 3 linear layers with ReLU activation functions before delivering its class probability calculations.<xref ref-type="fn" rid="fn3">
<sup>3</sup>
</xref> To reduce overfitting, dropout is used.</p>
<p>Because the ImageNet Challenge requires classifying images belonging to one of 1,000 categories, AlexNet by default has 1,000 nodes in its final layer. However, this can be modified when applying AlexNet for other purposes, such as classifying galaxy morphology into four categories.</p>
<p>The architecture of AlexNet is shown in a standard visual representation in <xref ref-type="fig" rid="F2">Figure 2</xref>. Note that the original AlexNet architecture (shown in <xref ref-type="fig" rid="F2">Figure 2</xref>) has 1,000 nodes in its output layer, corresponding to the 1,000 object categories in the ImageNet dataset. For this paper, the number of output nodes was reduced to 4, in accordance with the number of categories of galaxy morphology.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>The AlexNet architecture. After receiving a 224 &#xd7; 224 &#xd7; 3 input image, AlexNet applies a sequence of convolutions, ReLU operations, and pooling operations (in the portions represented by the deep rectangular prisms) before applying a set of linear layers with dropout (in the portions represented by the tall rectangular prisms). Figure created using NN-SVG.</p>
</caption>
<graphic xlink:href="fspas-10-1197358-g002.tif"/>
</fig>
</sec>
<sec id="s3-3">
<title>3.3 Data preparation, procedures and tooling</title>
<p>For our work, all galaxy images were scaled to 256 &#xd7; 256 in height and width to allow room for cropping to 224 &#xd7; 224, which is the input size required by AlexNet. Regarding data augmentation, PyTorch is naturally set up to use random data augmentation, which involves the probabilistic application of standard data augmentation techniques. For example, the PyTorch function <monospace>transforms.RandomHorizontalFlip()</monospace> will, each epoch, horizontally flip a given image with probability 0.5.</p>
<p>The data set was split into 12,000 training images and 2,034 testing images uniformly at random. We constructed 100 such test/train splits. No hyperparameter tuning was done. (The decision to forgo hyperparameter tuning was made in order to make an even comparison between the pretrained and non-pretrained networks.) For each random split of the data set, which we refer to as a <italic>run</italic>, the pretrained AlexNet was trained for 200 epochs, while the non-pretrained AlexNet was trained for 400 epochs. This additional training time was given to the non-pretrained AlexNet so that it had a better opportunity to achieve optimal performance. The training used the standard cross-entropy loss via the PyTorch function <monospace>torch.nn.CrossEntropyLoss()</monospace>, which is described in the PyTorch documentation.</p>
<p>Standard Python-language tools including Matplotlib (<xref ref-type="bibr" rid="B20">Hunter, 2007</xref>), NumPy (<xref ref-type="bibr" rid="B17">Harris et al., 2020</xref>), pandas (<xref ref-type="bibr" rid="B25">McKinney, 2010</xref>; <xref ref-type="bibr" rid="B28">Pandas development team, 2020</xref>), and seaborn (<xref ref-type="bibr" rid="B40">Waskom, 2021</xref>) were used for performing our numerical experiments and compiling the results to be presented. In particular, PyTorch (<xref ref-type="bibr" rid="B29">Paszke et al., 2019</xref>) was used to obtain the pretrained and non-pretrained AlexNet networks, and to train and test them. The networks used are available off the shelf in PyTorch, via the <monospace>torchvision.models</monospace> subpackage. Further, the Cedar computing system at Simon Fraser University, provisioned by the Digital Research Alliance of Canada and the BC DRI Group, was used to carry out the neural network training and testing. Specifically, we used an Intel Xeon Silver 4216 processor with 12 GB RAM, and a NVIDIA Tesla V100 32 GB GPU. The total computing time used to carry out the numerical experiments described in this work was 5.64 CPU months and 0.96 GPU months.</p>
</sec>
</sec>
<sec sec-type="results" id="s4">
<title>4 Results</title>
<p>In this section, we present and discuss the results obtained over the 100 runs, and compare the pretrained and non-pretrained networks on a variety of metrics. We begin with <xref ref-type="fig" rid="F3">Figure 3</xref>, in which we present histograms of the peak test accuracy for all runs for both the pretrained and non-pretrained networks; this figure also provides the epoch in which peak test accuracy is achieved. We also compute the difference in peak test accuracy between the pretrained and non-pretrained networks, as well as the difference in the epoch number for which peak test accuracy was achieved for the two networks, and display the results in <xref ref-type="fig" rid="F4">Figure 4</xref>. (For the latter calculation, only the first 200 epochs of training were considered for the non-pretrained network.) (Note that here and elsewhere in the paper, accuracy refers to the overall accuracy&#x2014;the percentage of correct classifications out of the total.)</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>
<bold>(A)</bold> Histograms of peak test accuracies for all runs for both the pretrained and non-pretrained networks. The pretrained networks are both more accurate on average and less varied in the peak accuracy that they achieve. <bold>(B)</bold> Histograms of the epoch in which peak test accuracy is achieved for each run, for both pretrained and non-pretrained networks. These histograms appear to suggest that 200 epochs of training are probably sufficient for the pretrained networks, while the non-pretrained networks might have benefited from even more than 400 epochs of training.</p>
</caption>
<graphic xlink:href="fspas-10-1197358-g003.tif"/>
</fig>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Histograms of the gains that result from using the pretrained network as opposed to the non-pretrained network. On the left is the peak test accuracy for the pretrained network minus the respective figure for the non-pretrained network. On the right is the epoch in which peak test accuracy was achieved for the pretrained network, minus the respective figure for the non-pretrained network, if the non-pretrained network had been restricted to train for only 200 epochs. It is clear that in terms of accuracy, the non-pretrained network never outperforms the pretrained network. In terms of speed in reaching peak accuracy, the pretrained network is almost always faster than the non-pretrained network, even when restricting the latter to 200 epochs of training.</p>
</caption>
<graphic xlink:href="fspas-10-1197358-g004.tif"/>
</fig>
<p>To compare the pretrained and non-pretrained networks given a fixed training budget of 200 epochs, for each run we compute the difference between the peak accuracy of the pretrained network and the highest accuracy achieved by the non-pretrained network within its first 200 epochs of training. These results are presented in <xref ref-type="fig" rid="F5">Figure 5</xref>. Examination of the histograms displayed in <xref ref-type="fig" rid="F3">Figures 3</xref>&#x2013;<xref ref-type="fig" rid="F5">5</xref> indicates that the pretrained AlexNet is preferred over the non-pretrained version. Pretraining leads to a higher overall accuracy, with an average peak accuracy (over the 100 runs) of 84.2% versus 82.4% for pretrained and non-pretrained, respectively. Furthermore, the pretrained network required significantly fewer epochs to reach its peak accuracy.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>This figure is similar to the left histogram in <xref ref-type="fig" rid="F4">Figure 4</xref>, except that we only consider the first 200 epochs for the non-pretrained networks. The advantage for the pretrained networks roughly doubles in this case.</p>
</caption>
<graphic xlink:href="fspas-10-1197358-g005.tif"/>
</fig>
<p>To further explore the efficiency gain of pretraining, we evaluate the performance of the pretrained network over a restricted portion of training, such as the first 20 or 50 epochs. Selected results are presented in <xref ref-type="fig" rid="F6">Figure 6</xref>. <xref ref-type="fig" rid="F6">Figure 6A</xref> is a histogram of the difference in peak test accuracy for the 200-epoch pretrained network and the 50-epoch pretrained network, showing that the average gain in accuracy from the additional 150 epochs of training is only about 1%, and the maximum gain over all 100 runs is less than 3%. If the pretrained network is limited to only 20 epochs, then <xref ref-type="fig" rid="F6">Figure 6B</xref> shows that the average gain in accuracy from the additional 180 epochs of training increases to approximately 2%&#x2013;2.5%, with a maximum gain over the 100 runs of approximately 4%. This suggests that good performance can be achieved with relatively few epochs, such as 20. Depending on the specific application and resource constraints, it may therefore be sufficient to train a pretrained network for a relatively small number of epochs.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>
<bold>(A)</bold> Gain from letting the pretrained networks train for 200 epochs, as opposed to only 50 epochs. The gain is roughly 1%. <bold>(B)</bold> Gain from letting the pretrained networks train for 200 epochs, as opposed to only 20 epochs. The gain is roughly 2%&#x2013;2.5%. <bold>(C)</bold> Each line in this panel contains the progression in test accuracy for the pretrained AlexNet for 1 of the 100 runs performed. It is clear that most of the improvement occurs within the first 20&#x2013;30 epochs of training. <bold>(D)</bold> Each line in this panel contains the progression in test accuracy for the non-pretrained AlexNet for 1 of the 100 runs performed. The width of the band of lines suggests that performance of the non-pretrained AlexNet is more variable than the pretrained AlexNet. Furthermore, it usually takes around 50 epochs of training for the non-pretrained AlexNet&#x2019;s parameters to adjust enough in order for the network&#x2019;s predictions to begin shifting.</p>
</caption>
<graphic xlink:href="fspas-10-1197358-g006.tif"/>
</fig>
<p>The bottom row of <xref ref-type="fig" rid="F6">Figure 6</xref> displays the test classification accuracy curves for pretrained (<xref ref-type="fig" rid="F6">Figure 6C</xref>) and non-pretrained (<xref ref-type="fig" rid="F6">Figure 6D</xref>) models over all 100 runs. The classification accuracy curves for the pretrained network reinforce the finding that the vast majority of improvement is acquired within the first 10&#x2013;20 epochs of training. The widths of the bands of lines (an informal measure of variability) indicate that there is much less variability when using a pretrained model than when using a non-pretrained model. <xref ref-type="fig" rid="F6">Figure 6D</xref> shows that the non-pretrained models require roughly 50 epochs of training in order to make any improvement at all; presumably this is the typical amount of training necessary to adjust a model&#x2019;s parameters sufficiently in order to begin changing its classification behavior.</p>
<p>
<xref ref-type="table" rid="T1">Table 1</xref> provides a summary of the results presented in <xref ref-type="fig" rid="F3">Figures 3</xref>&#x2013;<xref ref-type="fig" rid="F6">6</xref>. The information presented in the Table is consistent with that conveyed in the Figures, namely, that the pretrained network is more accurate and less variable in its performance than the non-pretrained network. <xref ref-type="table" rid="T1">Table 1</xref> also provides more precise insight into the diminishing returns from training the pretrained network. Within the first 10% of training (20 epochs as opposed to 200), the pretrained network achieves an average peak test accuracy of 82.0%; within the first 25% of training (50 epochs), the pretrained network achieves a peak test accuracy of 83.1%. These means are within approximately 2% and 1%, respectively, of the peak test accuracy of 84.2%, and it is worth noting that despite the reduced training time, the standard deviations are essentially indistinguishable from those for the full 200 epochs&#x2019; worth of training. This implies that there is no (or very little) variance penalty when doing a comparatively small amount of training of the pretrained network.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Selected figures summarizing the results from the present paper.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center"/>
<th colspan="6" align="center">Number of training epochs</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">Network Type</td>
<td align="center">400</td>
<td align="center">200</td>
<td align="center">50</td>
<td align="center">20</td>
<td align="center">10</td>
<td align="center">5</td>
</tr>
<tr>
<td align="center">Pretrained</td>
<td align="center">&#x2014;</td>
<td align="center">84.2%, 0.7%</td>
<td align="center">83.1%, 0.7%</td>
<td align="center">82.0%, 0.8%</td>
<td align="center">80.8%, 1.1%</td>
<td align="center">79.3%, 1.2%</td>
</tr>
<tr>
<td align="center">Non-Pretrained</td>
<td align="center">82.4%, 0.9%</td>
<td align="center">79.7%, 1.5%</td>
<td align="center">&#x2014;</td>
<td align="center">&#x2014;</td>
<td align="center">&#x2014;</td>
<td align="center">&#x2014;</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The percentages are the average peak accuracy (over 100 runs) and associated standard deviations. Certain figures are excluded from the table because they are not meaningful. Percentages are rounded to the nearest tenth of a percent. The pretrained AlexNet clearly outperforms the non-pretrained AlexNet, but analyzed in terms of efficiency its advantage is even more striking: with just 20 epochs of training it is clearly superior to the non-pretrained AlexNet with 10 times as much training, and almost tied with the non-pretrained AlexNet with 20 times as much training.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>
<xref ref-type="fig" rid="F7">Figure 7</xref> is presented as typical output for one of the 100 runs for both networks. (Run 49 was selected arbitrarily.) The upper panels present the train and test accuracy progression over all epochs of training, while the lower panels present train and test performance for both networks over all epochs from the perspective of the loss function. In general, the values of the loss function did not appear to provide any information not already apparent from the accuracy information, but the loss information is nevertheless presented in <xref ref-type="fig" rid="F7">Figure 7</xref> for additional illustration.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>
<bold>(A)</bold> Train and test accuracies of run 49 over all epochs for the pretrained AlexNet. The blue line in this panel is one of the lines in <xref ref-type="fig" rid="F6">Figure 6C</xref>. <bold>(B)</bold> Equivalent information is presented for the non-pretrained AlexNet. The blue line in this panel is one of the lines in <xref ref-type="fig" rid="F6">Figure 6D</xref>. <bold>(C)</bold> Train and test loss values of run 49 over all epochs for the pretrained AlexNet. <bold>(D)</bold> Equivalent information is presented for the non-pretrained AlexNet. Run 49 was chosen as an arbitrary representative of the 100 runs performed, although naturally there is some variation. Across most runs, performance of the pretrained network in training eventually surpasses performance in testing, although the point at which this occurs, and the eventual gap between the two, vary from run to run. This phenomenon is much weaker, perhaps nonexistent, for the non-pretrained networks, which may suggest that they would have benefited from more than 400 epochs of training. The non-pretrained networks also often take at least 50 epochs before the weights have adjusted enough for predictions to begin improving (Initially the non-pretrained networks appear to classify all test images as being of spiral galaxies).</p>
</caption>
<graphic xlink:href="fspas-10-1197358-g007.tif"/>
</fig>
<sec id="s4-1">
<title>4.1 Classification accuracies by class, and analysis of models</title>
<p>In order to develop a deeper understanding of model performance, the top-performing pretrained and non-pretrained models (by test accuracy) from each run were fed the test data set once again, and per-class accuracy figures were recorded. <xref ref-type="fig" rid="F8">Figures 8</xref>, <xref ref-type="fig" rid="F9">9</xref> display histograms of this result.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>Histograms of class accuracies for the most accurate pretrained model of each run. There is a clear hierarchy in performance that is somewhat consistent with the distribution of the test data set (i.e., highest performance in the most frequently occurring images), although the fact that the lenticular galaxies are in some sense &#x201c;between&#x201d; the elliptical and spiral galaxies seems to reduce lenticular accuracy.</p>
</caption>
<graphic xlink:href="fspas-10-1197358-g008.tif"/>
</fig>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>Histograms of class accuracies for the most accurate non-pretrained model of each run. These histograms are broadly similar to those in <xref ref-type="fig" rid="F8">Figure 8</xref>, with the main exception being much poorer Irr &#x2b; Misc performance.</p>
</caption>
<graphic xlink:href="fspas-10-1197358-g009.tif"/>
</fig>
<p>The models were all roughly the same in that they were quite accurate when presented with images of spiral galaxies, less accurate when presented with images of elliptical galaxies, and less accurate still when presented with images of lenticular galaxies. Furthermore, they perform quite poorly when presented with images of irregular or miscellaneous galaxies, presumably due to the fact that there are few such images in the data set, and moreover the category itself is ill defined in the sense that it mostly serves as a grab-bag of images that fit in none of the preceding three categories.</p>
<p>However, besides the fact that pretrained networks classify images of elliptical, lenticular and spiral galaxies slightly more accurately than do the non-pretrained networks, there is one striking difference, namely, that the pretrained networks classify images of Irr &#x2b; Misc galaxies almost twice as accurately as the non-pretrained networks. Even though the accuracy is still below 50%, this result might suggest that pretraining is potentially valuable for acquiring knowledge of uncommon or irregular examples in the application at hand. The fact that the pretrained networks offer a negligible improvement over the non-pretrained on spiral galaxies, by far the most common in the data set, might bolster this hypothesis.</p>
<p>
<xref ref-type="table" rid="T2">Table 2</xref> summarizes the results presented in <xref ref-type="fig" rid="F8">Figures 8</xref>, <xref ref-type="fig" rid="F9">9</xref>. This Table makes clear that not only is the pretrained network superior across all classes to the non-pretrained network in average class accuracy, but the pretrained network is also less variable in its classification performance. The only exception to this is the Irr &#x2b; Misc class, for which the pretrained network is more variable than the non-pretrained network, but the pretrained network&#x2019;s almost twofold superiority in average accuracy for this class offsets its slightly higher variability.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Class accuracy averages and standard deviations across all 100 runs for both pretrained and non-pretrained networks.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Class</th>
<th align="center">Pretrained</th>
<th align="center">Non-pretrained</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">Elliptical</td>
<td align="center">83.3%, 3.2%</td>
<td align="center">81.5%, 4.0%</td>
</tr>
<tr>
<td align="center">Lenticular</td>
<td align="center">66.4%, 3.0%</td>
<td align="center">64.2%, 3.0%</td>
</tr>
<tr>
<td align="center">Spiral</td>
<td align="center">92.9%, 1.2%</td>
<td align="center">92.4%, 1.7%</td>
</tr>
<tr>
<td align="center">Irr &#x2b; Misc</td>
<td align="center">38.1%, 7.9%</td>
<td align="center">21.4%, 6.7%</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Percentages are rounded to the nearest tenth of a percent. The pretrained networks are more accurate on average across all categories, and also have smaller or equal standard deviations with the exception of the Irr &#x2b; Misc category. Despite the larger standard deviation in that case, it seems clear that the pretrained networks are far superior for the Irr &#x2b; Misc category.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>
<xref ref-type="fig" rid="F10">Figure 10</xref> presents confusion matrices for the top-performing pretrained and non-pretrained models from run 49. (This is the same run as that presented in <xref ref-type="fig" rid="F7">Figure 7</xref>). The information in this figure is consistent with that presented in <xref ref-type="fig" rid="F8">Figures 8</xref>, <xref ref-type="fig" rid="F9">9</xref>, namely, the hierarchy in classification performance across the four categories of galaxy morphology, and the general superiority of the pretrained network. However, the confusion matrices also provide some insight into the nature of the <italic>mis</italic>classifications made by the networks. In particular, both pretrained and non-pretrained models tend to misclassify galaxies into adjacent morphological categories. For example, the majority of the misclassified spiral galaxies are classified as lenticular, as opposed to being classified as elliptical galaxies.</p>
<fig id="F10" position="float">
<label>FIGURE 10</label>
<caption>
<p>
<bold>(A)</bold> Confusion matrix for the top-performing model from run 49 of the pretrained AlexNet. The sum of the numbers within a row is the number of images of that type within the test set for that run. For example, looking at the second row, in run 49 there were 110 &#x2b; 309 &#x2b; 77 &#x2b; 5 &#x3d; 501 images of lenticular galaxies in the test set, and 309/501 &#x2248; 61.7% of those images were classified correctly. Furthermore, 110/501 &#x2248; 22.0% of images of lenticular galaxies were mis-classified as elliptical galaxies. The confusion matrices for all runs are roughly similar in that the models tend to mis-classify images of galaxies into adjacent categories, which is relatively sensible. <bold>(B)</bold> Confusion matrix for the top-performing model from run 49 of the non-pretrained AlexNet. Again, the sum of the numbers within a row is the number of images of that type within the test set for that run. For example, looking at the second row, in run 49 there were 95 &#x2b; 319 &#x2b; 86 &#x2b; 1 &#x3d; 501 images of lenticular galaxies in the test set, and 319/501 &#x2248; 63.7% of those images were classified correctly. Furthermore, 95/501 &#x2248; 19.0% of images of lenticular galaxies were mis-classified as elliptical galaxies. As with the pretrained network, the confusion matrices for all runs are roughly similar in that the models tend to mis-classify images of galaxies into adjacent categories.</p>
</caption>
<graphic xlink:href="fspas-10-1197358-g010.tif"/>
</fig>
<p>
<xref ref-type="fig" rid="F11">Figure 11</xref> presents confusion matrix data over all 100 runs for the pretrained and non-pretrained AlexNets. (Note that as opposed to <xref ref-type="fig" rid="F10">Figure 10</xref>, the totals have been converted to proportions). Similar to <xref ref-type="fig" rid="F10">Figure 10</xref>, <xref ref-type="fig" rid="F11">Figure 11</xref> provides information not only concerning how the models classify images of galaxies, but also how they <italic>mis</italic>classify images of galaxies. In <xref ref-type="fig" rid="F11">Figure 11A</xref>, the value in each cell of the matrix is the average value for that cell from the 100 individual confusion matrices for the pretrained network. <xref ref-type="fig" rid="F11">Figure 11C</xref> contains an equivalent matrix for the non-pretrained network. <xref ref-type="fig" rid="F11">Figures 11B, D</xref> contain the associated cell-wise standard deviations for the pretrained and non-pretrained networks, respectively. As with <xref ref-type="fig" rid="F10">Figure 10</xref>, the most salient feature of <xref ref-type="fig" rid="F11">Figure 11</xref> is the demonstration that both pretrained and non-pretrained networks tend to misclassify images of galaxies into adjacent categories. This is sensible given that the morphological characteristics of these galaxies are thought to occur on a continuum, at least to some extent.</p>
<fig id="F11" position="float">
<label>FIGURE 11</label>
<caption>
<p>
<bold>(A)</bold> The average confusion matrix for the pretrained AlexNet is shown. The entries in this confusion matrix are the means across all 100 confusion matrices for the pretrained network, one pertaining to each run. This figure shows that the network tends to misclassify images of galaxies into adjacent categories. <bold>(B)</bold> A matrix showing the standard deviations of the individual confusion matrices for the pretrained network over all 100 runs. <bold>(C)</bold> The average confusion matrix for the non-pretrained AlexNet. The entries in this confusion matrix are the means across all 100 confusion matrices for the non-pretrained network, one pertaining to each run. Like the top-left figure, the bottom-left figure shows that the non-pretrained network tends to misclassify images of galaxies into adjacent categories. <bold>(D)</bold> A matrix showing the standard deviations of the individual confusion matrices for the non-pretrained network over all 100 runs.</p>
</caption>
<graphic xlink:href="fspas-10-1197358-g011.tif"/>
</fig>
</sec>
<sec id="s4-2">
<title>4.2 Statistical significance tests</title>
<p>In order to provide a quantitative measure of the differences in performance between the pretrained and non-pretrained AlexNets, <italic>sign tests</italic> are conducted below. [The sign test makes no assumptions about the distribution of the quantities being compared (<xref ref-type="bibr" rid="B31">Roussas, 1997</xref>)].</p>
<p>Let <italic>X</italic>
<sub>0</sub>, <italic>X</italic>
<sub>1</sub>, <italic>&#x2026;</italic>, <italic>X</italic>
<sub>99</sub> be independent and identically distributed random variables with distribution function <italic>F</italic>, representing the distribution of peak test accuracy for the 100 runs of the pretrained AlexNet (over 200 epochs of training). Similarly, let <italic>Y</italic>
<sub>0</sub>, <italic>Y</italic>
<sub>1</sub>, <italic>&#x2026;</italic>, <italic>Y</italic>
<sub>99</sub> be independent and identically distributed random variables with distribution function <italic>G</italic>, representing the distribution of peak test accuracy for the 100 runs of the <italic>non</italic>-pretrained AlexNet (over 400 epochs of training). We wish to test the hypothesis<disp-formula id="equ4">
<mml:math id="m4">
<mml:mi>H</mml:mi>
<mml:mo>: </mml:mo>
<mml:mspace width=".17em"/>
<mml:mi>F</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>G</mml:mi>
<mml:mo>.</mml:mo>
</mml:math>
</disp-formula>The result of the two-sided sign test is a <italic>p</italic>-value of approximately 1.58 &#xd7; 10<sup>&#x2212;30</sup>, which provides strong evidence against the null hypothesis. We therefore have statistically significant evidence that the pretrained model is more accurate, even though the differences in accuracy may appear slight.</p>
<p>We can perform a similar test on differences in training time and energy used, using the number of epochs of training required to reach peak test accuracy as a proxy. Using similar definitions as above, the result of a two-sided sign test is the same <italic>p</italic>-value of approximately 1.58 &#xd7; 10<sup>&#x2212;30</sup>, which again provides strong evidence against the null hypothesis. We therefore have statistically significant evidence that the pretrained model is not only more accurate but also more efficient.</p>
</sec>
<sec id="s4-3">
<title>4.3 Persistence of improvement to increased number of training epochs</title>
<p>To test how well the non-pretrained network might perform if it were given more than 400 epochs of training, runs 45&#x2013;49 were repeated (i.e., the same train and test data splits were used), and both the pretrained and non-pretrained networks were trained for 1,000 epochs, instead of 200 and 400 epochs, respectively. (Runs 45&#x2013;49 were arbitrarily chosen; time and computational resource limitations prevent using 1,000 epochs for all 100 runs). <xref ref-type="table" rid="T3">Table 3</xref> contains peak test accuracies (rounded to the nearest tenth of a percent) and the epoch in which the peak test accuracy occurred.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Peak test accuracies (rounded to the nearest tenth of a percent) and the epoch in which the peak test accuracy occurred, using 1,000 epochs of training for runs 45&#x2013;49.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Run</th>
<th align="center">Pretrained</th>
<th align="center">Non-pretrained</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">45</td>
<td align="center">85.1%, 227</td>
<td align="center">84.5%, 878</td>
</tr>
<tr>
<td align="center">46</td>
<td align="center">84.9%, 244</td>
<td align="center">84.1%, 922</td>
</tr>
<tr>
<td align="center">47</td>
<td align="center">84.4%, 243</td>
<td align="center">83.4%, 994</td>
</tr>
<tr>
<td align="center">48</td>
<td align="center">84.5%, 466</td>
<td align="center">83.7%, 847</td>
</tr>
<tr>
<td align="center">49</td>
<td align="center">84.5%, 380</td>
<td align="center">84.4%, 932</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Examining <xref ref-type="table" rid="T3">Table 3</xref> we observe that, for each run, the pretrained network achieves a higher peak test accuracy than the non-pretrained network, and the difference is usually in the 0.5%&#x2013;1.0% range. This evidence suggests that it is not simply a matter of increasing the training time for the non-pretrained network in order to close the gap to the pretrained network, at least not up to 1,000 epochs of training. In other words, <xref ref-type="table" rid="T3">Table 3</xref> suggests that, at least up to 1,000 epochs of training, the advantage of pretraining on unrelated image data is not only greater efficiency but also greater ultimate performance in terms of overall accuracy.</p>
</sec>
</sec>
<sec id="s5">
<title>5 Summary and outlook</title>
<p>The main objective of this work is to compare pretrained (on ImageNet) and non-pretrained versions of AlexNet by training them on galaxy images from the Sloan Digital Sky Survey Data Release 4 [as described in <xref ref-type="bibr" rid="B26">Nair and Abraham (2010)</xref>] and comparing their performance and efficiency. We note that while the overall classification accuracies achieved are comparable to or slightly surpass similar attempts [e.g., those described in <xref ref-type="bibr" rid="B8">Cavanagh et al. (2021)</xref>], chasing the highest possible classification accuracy would lead us to consider other network architectures, hyperparameter tuning, etc. Rather, we have demonstrated the benefit to considering pretrained deep learning models for certain tasks. Our results are as follows:<list list-type="simple">
<list-item>
<p>1. The pretrained AlexNet had a consistent edge (compared to the non-pretrained AlexNet) in peak classification accuracy. It had an 84.2% average peak test accuracy, compared to an average peak test accuracy of 82.4% for the non-pretrained AlexNet.</p>
</list-item>
<list-item>
<p>2. The pretrained AlexNet was much more efficient (compared to the non-pretrained AlexNet) in that it attained peak test accuracy much more quickly. On average, the pretrained AlexNet achieved peak test accuracy in epoch 155 (standard deviation of 34 epochs), compared to epoch 367 (standard deviation of 33 epochs) for the non-pretrained AlexNet.</p>
</list-item>
<list-item>
<p>3. When considering only the first 200 epochs of training for the non-pretrained AlexNet, in order to provide a comparison with the pretrained AlexNet given an equal amount of training, the peak classification accuracy advantage for the pretrained AlexNet more than doubles, to about 4.6%.</p>
</list-item>
<list-item>
<p>4. The pretrained AlexNet achieves comparable performance to state-of-the-art methods, such as <xref ref-type="bibr" rid="B8">Cavanagh et al. (2021)</xref>, rather quickly. The pretrained AlexNet&#x2019;s average peak test accuracy after just 20 epochs of training is 82.0%, comparable with the headline 81%&#x2013;83% figures from <xref ref-type="bibr" rid="B8">Cavanagh et al. (2021)</xref>. After 50 epochs of training, the pretrained AlexNet&#x2019;s figure is 83.1%. This suggests that, taking advantage of pretraining, peak performance comparable to that from <xref ref-type="bibr" rid="B8">Cavanagh et al. (2021)</xref> can be achieved in as little as 10&#x2013;60 min, depending on the computational resources at hand.</p>
</list-item>
<list-item>
<p>5. Regarding per-class accuracies, the most striking advantage for the pretrained AlexNet is that it often classifies the Irr &#x2b; Misc images more than twice as accurately as the non-pretrained AlexNet. (Gains in classification accuracy for the other three categories are much smaller).</p>
</list-item>
</list>
</p>
<p>Regarding the last point, as a neural network is somewhat of a black box, it is hard to know precisely how the pretrained AlexNet becomes so much more adept at identifying images of Irr &#x2b; Misc galaxies. It seems reasonable to speculate that there is some sort of generalizable information within the unrelated ImageNet (pre)training set that is nevertheless applicable to classifying images of galaxies. Further speculating, it may be that pretraining on large, general, but unrelated data sets is of particular value in maximizing the ability to identify or classify rare cases in the particular application of interest, particularly when those rare cases are considered significant.</p>
<p>To further explore the benefit of pretraining for classifying Irr &#x2b; Misc galaxies, we compute precision, recall, and <italic>F</italic>
<sub>1</sub>-score for all galaxy types for both the pretrained and non-pretrained networks (evaluated on the unseen test data for each of the 100 runs); the results are presented in <xref ref-type="table" rid="T4">Table 4</xref>. The <italic>F</italic>
<sub>1</sub> score, as the harmonic mean of precision and recall, is a more holistic measure of model performance than either of its constituent components individually. The <italic>F</italic>
<sub>1</sub> score's sensitivity to class imbalances makes it a useful measure of model performance given an imbalanced dataset, as in the present case (<xref ref-type="bibr" rid="B44">Murphy, 2022</xref>). While there is only marginal improvement in precision, recall and <italic>F</italic>
<sub>1</sub> for elliptical, lenticular, and spiral galaxies, there is a significant improvement in recall and <italic>F</italic>
<sub>1</sub> for Irr &#x2b; Misc galaxies (but only a slight improvement in precision).</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Precision, recall and <italic>F</italic>
<sub>1</sub> for each class for both pretrained and non-pretrained networks.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Class</th>
<th align="center">Pretrained</th>
<th align="center">Non-pretrained</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">Elliptical precision</td>
<td align="center">79.0%, 2.7%</td>
<td align="center">78.2%, 3.6%</td>
</tr>
<tr>
<td align="center">Lenticular precision</td>
<td align="center">69.4%, 2.5%</td>
<td align="center">66.5%, 2.6%</td>
</tr>
<tr>
<td align="center">Spiral precision</td>
<td align="center">91.7%, 1.2%</td>
<td align="center">90.0%, 1.4%</td>
</tr>
<tr>
<td align="center">Irr &#x2b; Misc precision</td>
<td align="center">58.5%, 9.0%</td>
<td align="center">56.6%, 10.9%</td>
</tr>
<tr>
<td align="center">Elliptical recall</td>
<td align="center">83.3%, 3.3%</td>
<td align="center">81.5%, 4.0%</td>
</tr>
<tr>
<td align="center">Lenticular recall</td>
<td align="center">66.4%, 3.0%</td>
<td align="center">64.2%, 3.0%</td>
</tr>
<tr>
<td align="center">Spiral recall</td>
<td align="center">92.9%, 1.2%</td>
<td align="center">92.4%, 1.7%</td>
</tr>
<tr>
<td align="center">Irr &#x2b; Misc recall</td>
<td align="center">38.1%, 7.9%</td>
<td align="center">21.4%, 6.7%</td>
</tr>
<tr>
<td align="center">Elliptical <italic>F</italic>
<sub>1</sub>
</td>
<td align="center">81.0%, 1.5%</td>
<td align="center">79.7%, 1.6%</td>
</tr>
<tr>
<td align="center">Lenticular <italic>F</italic>
<sub>1</sub>
</td>
<td align="center">67.8%, 1.7%</td>
<td align="center">65.3%, 2.0%</td>
</tr>
<tr>
<td align="center">Spiral <italic>F</italic>
<sub>1</sub>
</td>
<td align="center">92.3%, 0.6%</td>
<td align="center">91.1%, 0.6%</td>
</tr>
<tr>
<td align="center">Irr &#x2b; Misc <italic>F</italic>
<sub>1</sub>
</td>
<td align="center">45.5%, 6.8%</td>
<td align="center">30.3%, 7.2%</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>In each cell, the first percentage is an average and the second a standard deviations, computed over all 100 runs. Percentages are rounded to the nearest tenth of a percent.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>A challenge for galaxy morphology classification and many areas of astrophysical image classification more generally is the relative lack of training data. The number of galaxy images available for the present work, 14,034, is much smaller than the amount of data typically available for training deep learning models; ImageNet alone contains more than 14,000,000 images, the labeling of which is trivial compared to classifying galaxy morphologies by hand. While upcoming surveys such as those by <italic>Euclid</italic> will generate more data, proper labelling remains a challenge. However, the use of pretrained models, as we described in this paper, offers the community a way of leveraging the significant effort already spent on developing and training deep learning models without sacrificing accuracy. Indeed, we suggest that the accuracy of a pretrained model may be slightly superior on common examples and vastly superior on rare examples, with much greater efficiency to boot.</p>
<p>Looking ahead we note that, in the deep learning community, AlexNet in particular and perhaps CNNs in general are no longer considered state of the art. This can be seen, for instance, in the progression of performance on the ImageNet data set over time. AlexNet is no longer close to the top-performing models on ImageNet, most of which are no longer CNNs. Transformer models (<xref ref-type="bibr" rid="B39">Vaswani et al., 2017</xref>) are often the highest-performing architectures currently and are much closer to the current state of the art. Exploration of their properties and performance is the subject of future work, especially the potential benefit of pretraining with them.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s7">
<title>Author contributions</title>
<p>JS performed the numerical experiments described in this article and drafted the manuscript. DCS co-supervised this work, and assisted with drafting and editing the manuscript. LTE co-supervised this work, and contributed to manuscript preparation. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>DCS is supported by NSERC grant number RGPIN/03985-2021. LTE is supported by a Michael Smith Health Research BC Scholar Award, and NSERC grant numbers RGPIN/05484-2019 and DGECR/00118-2019.</p>
</sec>
<ack>
<p>Thank you to the Digital Research Alliance of Canada and the BC DRI Group for their provision of the Cedar system at Simon Fraser University, which was used in the preparation of this paper.</p>
</ack>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<fn-group>
<fn id="fn1">
<label>1</label>
<p>Methods and tools used in the present work are similar although not identical to those used in (<xref ref-type="bibr" rid="B8">Cavanagh et al., 2021</xref>), so while the peak accuracy results obtained in the present paper are slightly superior to those obtained in (<xref ref-type="bibr" rid="B8">Cavanagh et al., 2021</xref>), the results are not directly comparable and we therefore do not make such a comparison.</p>
</fn>
<fn id="fn2">
<label>2</label>
<p>The convolution operation, for which CNNs are named, is actually a cross-correlation since neither of the functions in the operation&#x2019;s arguments are reflected. Despite this, this paper will adhere to the convention of referring to this operation as a convolution.</p>
</fn>
<fn id="fn3">
<label>3</label>
<p>
<xref ref-type="bibr" rid="B22">Krizhevsky et al. (2012)</xref>, which introduced the AlexNet architecture, makes clear that the size of AlexNet was limited by computational power and memory, and patience to endure long training times.</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Abbott</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Abdalla</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Alarcon</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Aleksic</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Allam</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Allen</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Dark energy survey year 1 results: cosmological constraints from galaxy clustering and weak lensing</article-title>. <source>Phys. Rev. D.</source> <volume>98</volume>, <fpage>043526</fpage>. <pub-id pub-id-type="doi">10.1103/physrevd.98.043526</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ackermann</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Schawinski</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Weigel</surname>
<given-names>A. K.</given-names>
</name>
<name>
<surname>Turp</surname>
<given-names>M. D.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Using transfer learning to detect galaxy mergers</article-title>. <source>Mon. Notices R. Astronomical Soc.</source> <volume>479</volume>, <fpage>415</fpage>&#x2013;<lpage>425</lpage>. <pub-id pub-id-type="doi">10.1093/mnras/sty1398</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Adelman-McCarthy</surname>
<given-names>J. K.</given-names>
</name>
<name>
<surname>Ag&#xfc;eros</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Allam</surname>
<given-names>S. S.</given-names>
</name>
<name>
<surname>Anderson</surname>
<given-names>K. S. J.</given-names>
</name>
<name>
<surname>Anderson</surname>
<given-names>S. F.</given-names>
</name>
<name>
<surname>Annis</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2006</year>). <article-title>The fourth data release of the sloan digital sky survey</article-title>. <source>Astrophysical J. Suppl. Ser.</source> <volume>162</volume>, <fpage>38</fpage>&#x2013;<lpage>48</lpage>. <pub-id pub-id-type="doi">10.1086/497917</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Aggarwal</surname>
<given-names>C. C.</given-names>
</name>
</person-group> (<year>2018</year>). <source>
<italic>Neural Networks and deep learning: A textbook</italic> (gewerbestrasse 11</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>, <fpage>6330</fpage>.</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Barchi</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>de Carvalho</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Rosa</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Sautter</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Soares-Santos</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Marques</surname>
<given-names>B.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Machine and deep learning applied to galaxy morphology - a comparative study</article-title>. <source>Astronomy Comput.</source> <volume>30</volume>, <fpage>100334</fpage>. <pub-id pub-id-type="doi">10.1016/j.ascom.2019.100334</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Borowiec</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Harper</surname>
<given-names>R. R.</given-names>
</name>
<name>
<surname>Garraghan</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>The environmental consequence of deep learning</article-title>. <source>ITNOW</source> <volume>63</volume>, <fpage>10</fpage>&#x2013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1093/itnow/bwab099</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Cabrera-Vives</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Reyes</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>F&#xf6;rster</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Est&#xe9;vez</surname>
<given-names>P. A.</given-names>
</name>
<name>
<surname>Maureira</surname>
<given-names>J.-C.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Supernovae detection by using convolutional neural networks</article-title>,&#x201d; in <conf-name>2016 International Joint Conference on Neural Networks (IJCNN)</conf-name>, <conf-loc>Vancouver, BC, Canada</conf-loc>, <conf-date>July 24-29, 2016</conf-date>, <fpage>251</fpage>&#x2013;<lpage>258</lpage>. <pub-id pub-id-type="doi">10.1109/IJCNN.2016.7727206</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cavanagh</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Bekki</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Groves</surname>
<given-names>B. A.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Morphological classification of galaxies with deep learning: comparing 3-way and 4-way CNNs</article-title>. <source>Mon. Notices R. Astronomical Soc.</source> <volume>506</volume>, <fpage>659</fpage>&#x2013;<lpage>676</lpage>. <pub-id pub-id-type="doi">10.1093/mnras/stab1552</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cheng</surname>
<given-names>T.-Y.</given-names>
</name>
<name>
<surname>Conselice</surname>
<given-names>C. J.</given-names>
</name>
<name>
<surname>Arag&#xf3;n-Salamanca</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Bluck</surname>
<given-names>A. F. L.</given-names>
</name>
<name>
<surname>Hartley</surname>
<given-names>W. G.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Optimizing automatic morphological classification of galaxies with machine learning and deep learning using Dark Energy Survey imaging</article-title>. <source>Mon. Notices R. Astronomical Soc.</source> <volume>493</volume>, <fpage>4209</fpage>&#x2013;<lpage>4228</lpage>. <pub-id pub-id-type="doi">10.1093/mnras/staa501</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Davies</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Serjeant</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bromley</surname>
<given-names>J. M.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Using convolutional neural networks to identify gravitational lenses in astronomical images</article-title>. <source>Mon. Notices R. Astronomical Soc.</source> <volume>487</volume>, <fpage>5263</fpage>&#x2013;<lpage>5271</lpage>. <pub-id pub-id-type="doi">10.1093/mnras/stz1288</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Deng</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Dong</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Socher</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>L.-J.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Fei-Fei</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2009</year>). &#x201c;<article-title>Imagenet: a large-scale hierarchical image database</article-title>,&#x201d; in <conf-name>2009 IEEE conference on computer vision and pattern recognition (Ieee)</conf-name>, <conf-loc>Miami, Florida</conf-loc>, <conf-date>Held 20-25 June 2009</conf-date>, <fpage>248</fpage>&#x2013;<lpage>255</lpage>.</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dom&#xed;nguez S&#xe1;nchez</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Huertas-Company</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Bernardi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kaviraj</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Fischer</surname>
<given-names>J. L.</given-names>
</name>
<name>
<surname>Abbott</surname>
<given-names>T. M. C.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Transfer learning for galaxy morphology from one survey to another</article-title>. <source>Mon. Notices R. Astronomical Soc.</source> <volume>484</volume>, <fpage>93</fpage>&#x2013;<lpage>100</lpage>. <pub-id pub-id-type="doi">10.1093/mnras/sty3497</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Garc&#xed;a-Mart&#xed;n</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Rodrigues</surname>
<given-names>C. F.</given-names>
</name>
<name>
<surname>Riley</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Grahn</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Estimation of energy consumption in machine learning</article-title>. <source>J. Parallel Distributed Comput.</source> <volume>134</volume>, <fpage>75</fpage>&#x2013;<lpage>88</lpage>. <pub-id pub-id-type="doi">10.1016/j.jpdc.2019.07.007</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>George</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Shen</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Huerta</surname>
<given-names>E. A.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Classification and unsupervised clustering of ligo data with deep transfer learning</article-title>. <source>Phys. Rev. D.</source> <volume>97</volume>, <fpage>101501</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevD.97.101501</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gharat</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Dandawate</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Galaxy classification: a deep learning approach for classifying sloan digital sky survey images</article-title>. <source>Mon. Notices R. Astronomical Soc.</source> <volume>511</volume>, <fpage>5120</fpage>&#x2013;<lpage>5124</lpage>. <pub-id pub-id-type="doi">10.1093/mnras/stac457</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Goodfellow</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Bengio</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Courville</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2016</year>). <source>
<italic>Deep learning</italic> (one broadway 12th floor cambridge, MA 02142</source>. <publisher-loc>United States</publisher-loc>: <publisher-name>The MIT Press</publisher-name>.</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Harris</surname>
<given-names>C. R.</given-names>
</name>
<name>
<surname>Millman</surname>
<given-names>K. J.</given-names>
</name>
<name>
<surname>van der Walt</surname>
<given-names>S. J.</given-names>
</name>
<name>
<surname>Gommers</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Virtanen</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Cournapeau</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Array programming with NumPy</article-title>. <source>Nature</source> <volume>585</volume>, <fpage>357</fpage>&#x2013;<lpage>362</lpage>. <pub-id pub-id-type="doi">10.1038/s41586-020-2649-2</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Hinton</surname>
<given-names>G. E.</given-names>
</name>
<name>
<surname>Srivastava</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Krizhevsky</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sutskever</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Salakhutdinov</surname>
<given-names>R. R.</given-names>
</name>
</person-group> (<year>2012</year>). <source>Improving neural networks by preventing co-adaptation of feature detectors</source>. <comment>arXiv</comment>. <pub-id pub-id-type="doi">10.48550/ARXIV.1207.0580</pub-id>
</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hubel</surname>
<given-names>D. H.</given-names>
</name>
<name>
<surname>Wiesel</surname>
<given-names>T. N.</given-names>
</name>
</person-group> (<year>1959</year>). <article-title>Receptive fields of single neurones in the cat&#x2019;s striate cortex</article-title>. <source>J. Physiology</source> <volume>148</volume>, <fpage>574</fpage>&#x2013;<lpage>591</lpage>. <pub-id pub-id-type="doi">10.1113/jphysiol.1959.sp006308</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hunter</surname>
<given-names>J. D.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Matplotlib: a 2d graphics environment</article-title>. <source>Comput. Sci. Eng.</source> <volume>9</volume>, <fpage>90</fpage>&#x2013;<lpage>95</lpage>. <pub-id pub-id-type="doi">10.1109/MCSE.2007.55</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kim</surname>
<given-names>D-W.</given-names>
</name>
<name>
<surname>Yeo</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Bailer-Jones</surname>
</name>
<name>
<surname>Coryn</surname>
<given-names>A. L.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Deep transfer learning for the classification of variable sources</article-title>. <source>A&#x26;A</source> <volume>653</volume>, <fpage>A22</fpage>. <pub-id pub-id-type="doi">10.1051/0004-6361/202140369</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Krizhevsky</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sutskever</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Hinton</surname>
<given-names>G. E.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Imagenet classification with deep convolutional neural networks</article-title>,&#x201d; in <source>Advances in neural information processing systems</source> Editors <person-group person-group-type="editor">
<name>
<surname>Pereira</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Burges</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Bottou</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Weinberger</surname>
<given-names>K.</given-names>
</name>
</person-group> (<publisher-loc>United States</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>).</citation>
</ref>
<ref id="B23">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Cai</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Feng</surname>
<given-names>D. D.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2014</year>). &#x201c;<article-title>Medical image classification with convolutional neural network</article-title>,&#x201d; in <conf-name>2014 13th International Conference on Control Automation Robotics Vision (ICARCV)</conf-name>, <conf-loc>Singapore</conf-loc>, <conf-date>10-12 December 2014</conf-date>, <fpage>844</fpage>. <pub-id pub-id-type="doi">10.1109/ICARCV.2014.7064414</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lintott</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Schawinski</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Bamford</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Slosar</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Land</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Thomas</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2010</year>). <article-title>Galaxy Zoo 1: data release of morphological classifications for nearly 900 000 galaxies</article-title>. <source>Mon. Notices R. Astronomical Soc.</source> <volume>410</volume>, <fpage>166</fpage>&#x2013;<lpage>178</lpage>. <pub-id pub-id-type="doi">10.1111/j.1365-2966.2010.17432.x</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>McKinney</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2010</year>). &#x201c;<article-title>Data structures for statistical computing in Python</article-title>,&#x201d; in <conf-name>Proceedings of the 9th Python in Science Conference</conf-name>, <conf-loc>Austin, Texas</conf-loc>, <conf-date>June 28 - July 3, 2010</conf-date>, <fpage>56</fpage>&#x2013;<lpage>61</lpage>. <pub-id pub-id-type="doi">10.25080/Majora-92bf1922-00a</pub-id>
</citation>
</ref>
<ref id="B44">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Murphy</surname>
<given-names>K. P.</given-names>
</name>
</person-group> (<year>2022</year>). <source>Probabilistic machine learning: an introduction</source>. <publisher-name>MIT Press</publisher-name>.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nair</surname>
<given-names>P. B.</given-names>
</name>
<name>
<surname>Abraham</surname>
<given-names>R. G.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>A catalog of detailed visual morphological classifications for 14,034 galaxies in the sloan</article-title>. <source>Digit. Sky Surv.</source> <volume>186</volume>, <fpage>427</fpage>&#x2013;<lpage>456</lpage>. <pub-id pub-id-type="doi">10.1088/0067-0049/186/2/427</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Paillassa</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Bertin</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Bouy</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Maximask and maxitrack: two new tools for identifying contaminants in astronomical images using convolutional neural networks</article-title>. <source>A&#x26;A</source> <volume>634</volume>, <fpage>A48</fpage>. <pub-id pub-id-type="doi">10.1051/0004-6361/201936345</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="book">
<collab>Pandas development team</collab> (<year>2020</year>). <source>pandas-dev/pandas: Pandas</source>. <pub-id pub-id-type="doi">10.5281/zenodo.3509134</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Paszke</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Gross</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Francisco</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Adam</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>James</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Gregory</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>PyTorch: an imperative style, high-performance deep learning library</article-title>. <source>arXiv e-prints</source>, <fpage>arXiv:1912.01703</fpage>. <pub-id pub-id-type="doi">10.48550/arXiv.1912.01703</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ribani</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Marengoni</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>A survey of transfer learning for convolutional neural networks</article-title>,&#x201d; in <conf-name>2019 32nd SIBGRAPI Conference on Graphics, Patterns and Images Tutorials (SIBGRAPI-T)</conf-name>, <conf-loc>Rio de Janeiro, Brazil</conf-loc>, <conf-date>Oct. 28 2019 to Oct. 31 2019</conf-date>, <fpage>47</fpage>&#x2013;<lpage>57</lpage>. <pub-id pub-id-type="doi">10.1109/SIBGRAPI-T.2019.00010</pub-id>
</citation>
</ref>
<ref id="B31">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Roussas</surname>
<given-names>G. G.</given-names>
</name>
</person-group> (<year>1997</year>). <source>A course in mathematical statistics</source>. <edition>Second Edition</edition>. <publisher-loc>San Diego, CA 92101-4495, United States</publisher-loc>: <publisher-name>Academic Press</publisher-name>. <comment>525 B Street, Suite 1900</comment>.</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Silva</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>L. T.</given-names>
</name>
<name>
<surname>Hayes</surname>
<given-names>W. B.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Sparcfire: enhancing spiral galaxy recognition using arm analysis and random forests</article-title>. <source>Galaxies</source> <volume>6</volume>, <fpage>95</fpage>. <pub-id pub-id-type="doi">10.3390/galaxies6030095</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Srivastava</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Hinton</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Krizhevsky</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sutskever</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Salakhutdinov</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Dropout: a simple way to prevent neural networks from overfitting</article-title>. <source>J. Mach. Learn. Res.</source> <volume>15</volume>, <fpage>1929</fpage>&#x2013;<lpage>1958</lpage>.</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Stoughton</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Lupton</surname>
<given-names>R. H.</given-names>
</name>
<name>
<surname>Bernardi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Blanton</surname>
<given-names>M. R.</given-names>
</name>
<name>
<surname>Burles</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Castander</surname>
<given-names>F. J.</given-names>
</name>
<etal/>
</person-group> (<year>2002</year>). <article-title>Sloan digital sky survey: early data release</article-title>. <source>Sloan Digit. Sky Surv. Early Data Release</source> <volume>123</volume>, <fpage>485</fpage>&#x2013;<lpage>548</lpage>. <pub-id pub-id-type="doi">10.1086/324741</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Strubell</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Ganesh</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>McCallum</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Energy and policy considerations for modern deep learning research</article-title>. <source>Proc. AAAI Conf. Artif. Intell.</source> <volume>34</volume>, <fpage>13693</fpage>&#x2013;<lpage>13696</lpage>. <pub-id pub-id-type="doi">10.1609/aaai.v34i09.7123</pub-id>
</citation>
</ref>
<ref id="B36">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Taigman</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ranzato</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Wolf</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2014</year>). &#x201c;<article-title>Deepface: closing the gap to human-level performance in face verification</article-title>,&#x201d; in <conf-name>2014 IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Columbus, OH, USA</conf-loc>, <conf-date>June 23 2014 to June 28 2014</conf-date>, <fpage>1701</fpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2014.220</pub-id>
</citation>
</ref>
<ref id="B37">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Tan</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Kong</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>A survey on deep transfer learning</article-title>,&#x201d; in <source>Artificial neural networks and machine learning &#x2013; ICANN 2018</source> Editors <person-group person-group-type="editor">
<name>
<surname>K&#x16f;rkov&#xe1;</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Manolopoulos</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Hammer</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Iliadis</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Maglogiannis</surname>
<given-names>I.</given-names>
</name>
</person-group> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <fpage>270</fpage>&#x2013;<lpage>279</lpage>.</citation>
</ref>
<ref id="B38">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Tsantekidis</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Passalis</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Tefas</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Kanniainen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Gabbouj</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Iosifidis</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Forecasting stock prices from the limit order book using convolutional neural networks</article-title>.&#x201d; in <conf-name>2017 IEEE 19th Conference on Business Informatics (CBI)</conf-name>, <conf-loc>Thessaloniki, Greece</conf-loc>, <conf-date>24-27 July 2017</conf-date>, <fpage>7</fpage>&#x2013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1109/CBI.2017.23</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Vaswani</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Shazeer</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Parmar</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Uszkoreit</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Jones</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Gomez</surname>
<given-names>A. N.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). &#x201c;<article-title>Attention is all you need</article-title>,&#x201d; in <source>Advances in neural information processing systems</source> Editors <person-group person-group-type="editor">
<name>
<surname>Guyon</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Luxburg</surname>
<given-names>U. V.</given-names>
</name>
<name>
<surname>Bengio</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Wallach</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Fergus</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Vishwanathan</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<publisher-loc>Unites States</publisher-loc>: <publisher-name>Curran Associates, Inc.</publisher-name>).</citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Waskom</surname>
<given-names>M. L.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>seaborn: statistical data visualization</article-title>. <source>J. Open Source Softw.</source> <volume>6</volume>, <fpage>3021</fpage>. <pub-id pub-id-type="doi">10.21105/joss.03021</pub-id>
</citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Huerta</surname>
<given-names>E. A.</given-names>
</name>
<name>
<surname>Whitmore</surname>
<given-names>B. C.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>J. C.</given-names>
</name>
<name>
<surname>Hannon</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Chandar</surname>
<given-names>R.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Deep transfer learning for star cluster classification: &#x2160;. Application to the PHANGS&#x2013;HST survey</article-title>. <source>Mon. Notices R. Astronomical Soc.</source> <volume>493</volume>, <fpage>3178</fpage>&#x2013;<lpage>3193</lpage>. <pub-id pub-id-type="doi">10.1093/mnras/staa325</pub-id>
</citation>
</ref>
<ref id="B42">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>T.-J.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y.-H.</given-names>
</name>
<name>
<surname>Emer</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Sze</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>A method to estimate the energy consumption of deep neural networks</article-title>,&#x201d; in <conf-name>2017 51st Asilomar Conference on Signals, Systems, and Computers</conf-name>, <conf-loc>Pacific Grove, CA, USA</conf-loc>, <conf-date>29 October&#x2013;1 November 2017</conf-date>, <fpage>1916</fpage>&#x2013;<lpage>1920</lpage>. <pub-id pub-id-type="doi">10.1109/ACSSC.2017.8335698</pub-id>
</citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>York</surname>
<given-names>D. G.</given-names>
</name>
<name>
<surname>Adelman</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>John</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Anderson</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Anderson</surname>
<given-names>S. F.</given-names>
</name>
<name>
<surname>Annis</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2000</year>). <article-title>The sloan digital sky survey: technical summary</article-title>. <source>Astronomical J.</source> <volume>120</volume>, <fpage>1579</fpage>&#x2013;<lpage>1587</lpage>. <pub-id pub-id-type="doi">10.1086/301513</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>