<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="methods-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Aging Neurosci.</journal-id>
<journal-title>Frontiers in Aging Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Aging Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1663-4365</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnagi.2021.635945</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Comparing Pre-trained and Feature-Based Models for Prediction of Alzheimer&#x00027;s Disease Based on Speech</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Balagopalan</surname> <given-names>Aparna</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1033074/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Eyre</surname> <given-names>Benjamin</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1260322/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Robin</surname> <given-names>Jessica</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1223771/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Rudzicz</surname> <given-names>Frank</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Novikova</surname> <given-names>Jekaterina</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1157893/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Winterlight Labs Inc.</institution>, <addr-line>Toronto, ON</addr-line>, <country>Canada</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Computer Science, University of Toronto</institution>, <addr-line>Toronto, ON</addr-line>, <country>Canada</country></aff>
<aff id="aff3"><sup>3</sup><institution>Vector Institute for Artificial Intelligence</institution>, <addr-line>Toronto, ON</addr-line>, <country>Canada</country></aff>
<aff id="aff4"><sup>4</sup><institution>Unity Health Toronto</institution>, <addr-line>Toronto, ON</addr-line>, <country>Canada</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Saturnino Luz, University of Edinburgh, United Kingdom</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Kewei Chen, Banner Alzheimer&#x00027;s Institute, United States; Juan Jos&#x000E9; Garc&#x000ED;a Meil&#x000E1;n, University of Salamanca, Spain</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Jekaterina Novikova <email>jekaterina&#x00040;winterlightlabs.com</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>27</day>
<month>04</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>13</volume>
<elocation-id>635945</elocation-id>
<history>
<date date-type="received">
<day>30</day>
<month>11</month>
<year>2020</year>
</date>
<date date-type="accepted">
<day>24</day>
<month>03</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2021 Balagopalan, Eyre, Robin, Rudzicz and Novikova.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Balagopalan, Eyre, Robin, Rudzicz and Novikova</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract><p><bold>Introduction:</bold> Research related to the automatic detection of Alzheimer&#x00027;s disease (AD) is important, given the high prevalence of AD and the high cost of traditional diagnostic methods. Since AD significantly affects the content and acoustics of spontaneous speech, natural language processing, and machine learning provide promising techniques for reliably detecting AD. There has been a recent proliferation of classification models for AD, but these vary in the datasets used, model types and training and testing paradigms. In this study, we compare and contrast the performance of two common approaches for automatic AD detection from speech on the same, well-matched dataset, to determine the advantages of using domain knowledge vs. pre-trained transfer models.</p>
<p><bold>Methods:</bold> Audio recordings and corresponding manually-transcribed speech transcripts of a picture description task administered to 156 demographically matched older adults, 78 with Alzheimer&#x00027;s Disease (AD) and 78 cognitively intact (healthy) were classified using machine learning and natural language processing as &#x0201C;AD&#x0201D; or &#x0201C;non-AD.&#x0201D; The audio was acoustically-enhanced, and post-processed to improve quality of the speech recording as well control for variation caused by recording conditions. Two approaches were used for classification of these speech samples: (1) using domain knowledge: extracting an extensive set of clinically relevant linguistic and acoustic features derived from speech and transcripts based on prior literature, and (2) using transfer-learning and leveraging large pre-trained machine learning models: using transcript-representations that are automatically derived from state-of-the-art pre-trained language models, by fine-tuning Bidirectional Encoder Representations from Transformer (BERT)-based sequence classification models.</p>
<p><bold>Results:</bold> We compared the utility of speech transcript representations obtained from recent natural language processing models (i.e., BERT) to more clinically-interpretable language feature-based methods. Both the feature-based approaches and fine-tuned BERT models significantly outperformed the baseline linguistic model using a small set of linguistic features, demonstrating the importance of extensive linguistic information for detecting cognitive impairments relating to AD. We observed that fine-tuned BERT models numerically outperformed feature-based approaches on the AD detection task, but the difference was not statistically significant. Our main contribution is the observation that when tested on the same, demographically balanced dataset and tested on independent, unseen data, both domain knowledge and pretrained linguistic models have good predictive performance for detecting AD based on speech. It is notable that linguistic information alone is capable of achieving comparable, and even numerically better, performance than models including both acoustic and linguistic features here. We also try to shed light on the inner workings of the more black-box natural language processing model by performing an interpretability analysis, and find that attention weights reveal interesting patterns such as higher attribution to more important information content units in the picture description task, as well as pauses and filler words.</p>
<p><bold>Conclusion:</bold> This approach supports the value of well-performing machine learning and linguistically-focussed processing techniques to detect AD from speech and highlights the need to compare model performance on carefully balanced datasets, using consistent same training parameters and independent test datasets in order to determine the best performing predictive model.</p></abstract>
<kwd-group>
<kwd>Alzheimer&#x00027;s disease</kwd>
<kwd>dementia detection</kwd>
<kwd>MMSE regression</kwd>
<kwd>BERT</kwd>
<kwd>feature engineering</kwd>
<kwd>transfer learning</kwd>
</kwd-group>
<counts>
<fig-count count="2"/>
<table-count count="11"/>
<equation-count count="0"/>
<ref-count count="54"/>
<page-count count="12"/>
<word-count count="8950"/>
</counts>
</article-meta>
</front>
<body>

<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Alzheimer&#x00027;s disease (AD) is a progressive neurodegenerative disease that causes problems with memory, thinking, and behavior. AD affects over 40 million people worldwide with high costs of acute and long-term care (Prince et al., <xref ref-type="bibr" rid="B38">2016</xref>). Current forms of diagnosis are both time consuming and expensive (Prabhakaran et al., <xref ref-type="bibr" rid="B37">2018</xref>), which might explain why almost half of those living with AD do not receive a timely diagnosis (Jammeh et al., <xref ref-type="bibr" rid="B20">2018</xref>).</p>
<p>Studies have shown that valuable clinical information indicative of cognition can be obtained from spontaneous speech elicited using pictures (Goodglass et al., <xref ref-type="bibr" rid="B18">2001</xref>). Studies have capitalized on this clinical observation, using speech analysis, natural language processing (NLP), and machine learning (ML) to distinguish between speech from healthy and cognitively impaired participants in datasets including semi-structured speech tasks such as picture description. Some of the first papers on this topic reported ML methods for automatic AD-detection using speech datasets achieving high classification performance (between 82 and 93% accuracy) (K&#x000F6;nig et al., <xref ref-type="bibr" rid="B25">2015</xref>; Fraser et al., <xref ref-type="bibr" rid="B17">2016</xref>; Noorian et al., <xref ref-type="bibr" rid="B32">2017</xref>; Karlekar et al., <xref ref-type="bibr" rid="B23">2018</xref>; Zhu et al., <xref ref-type="bibr" rid="B53">2018</xref>; Gosztolya et al., <xref ref-type="bibr" rid="B19">2019</xref>). These models serve as quick, objective, and non-invasive assessments of an individual&#x00027;s cognitive status which could be developed into more accessible tools to facilitate clinical screening and diagnosis. Since these initial reports, there has been a proliferation of studies reporting classification models for AD based on speech, as described by recent reviews and meta-analyses (Slegers et al., <xref ref-type="bibr" rid="B42">2018</xref>; de la Fuente Garcia et al., <xref ref-type="bibr" rid="B12">2020</xref>; Petti et al., <xref ref-type="bibr" rid="B35">2020</xref>; Pulido et al., <xref ref-type="bibr" rid="B39">2020</xref>), but the field still lacks validation of predictive models on publicly-available, balanced, and standardized benchmark datasets.</p>
<p>The existing studies that have addressed differences between AD and non-AD speech and worked on developing speech-based AD biomarkers, are often descriptive rather than predictive. Thus, they often overlook common biases in evaluations of AD detection methods, such as repeated occurrences of speech from the same participant, variations in audio quality of speech samples, and imbalances of gender and age distribution in the used datasets, as noted in the systematic reviews and meta-analyses published on this topic (Slegers et al., <xref ref-type="bibr" rid="B42">2018</xref>; Chen et al., <xref ref-type="bibr" rid="B8">2020</xref>; Petti et al., <xref ref-type="bibr" rid="B35">2020</xref>). As such, the existing ML models may be prone to the biases introduced in available data. In addition, the performance of the previously developed predictive AD-detection models has been evaluated using either random train/test split or a cross-validation technique, which may result in artificially increased reported performance of ML models (i.e., overfitting) as compared to their evaluation on a held out unseen dataset (more details on evaluation techniques are provided in the section 2.3.1.2), especially when it comes to smaller and unbalanced datasets (Johnson et al., <xref ref-type="bibr" rid="B22">2018</xref>). Due to these reasons, it&#x00027;s difficult to compare model performance across papers and datasets, since they are rarely matched in terms of data and model characteristics.</p>
<p>To overcome the problem of bias and overfitting and introduce a common dataset to compare model performance, the ADReSS challenge (Luz et al., <xref ref-type="bibr" rid="B28">2020</xref>) was introduced in 2020, in which the organizers provided an age/sex-matched balanced speech dataset, which consisted of speech from AD and non-AD participants describing a picture. The challenge consisted of two key tasks: (1) Speech classification task: classifying speech as AD or non-AD. (2) Neuropsychological score regression task: predicting Mini-Mental State Examination (MMSE) (Cockrell and Folstein, <xref ref-type="bibr" rid="B9">2002</xref>) scores from speech. The organizers restricted access to the test dataset to make it completely unseen for participants to ensure the fair evaluation of models&#x00027; performance. The work presented in this paper is focused entirely on this new balanced dataset and follows the ADReSS challenge&#x00027;s evaluation process. As such, the models presented in this paper are more generalizable to unseen data than those developed in the previously discussed studies.</p>
<p>In this work, we develop ML models to detect AD from speech using picture description data of the demographically-matched ADReSS Challenge speech dataset (Luz et al., <xref ref-type="bibr" rid="B28">2020</xref>), and compare the following training regimes and input representations to detect AD:</p>
<list list-type="order">
<list-item><p><bold>Using domain knowledge</bold>: with this approach, we extract clinically relevant linguistic features from transcripts of speech, and acoustic features from corresponding audio files for binary AD vs. non-AD classification and MMSE score regression. The features extracted are informed by previous clinical and ML research in the space of cognitive impairment detection (Fraser et al., <xref ref-type="bibr" rid="B17">2016</xref>).</p></list-item>
<list-item><p><bold>Using transfer learning</bold>: with this approach, we fine-tune pre-trained BERT</p></list-item>
</list>
<p>We describe below the details of each approach.</p>
<sec>
<title>1.1. Domain Knowledge-Based Approach</title>
<p>The overwhelming majority of NLP and ML approaches on AD detection from speech are still based on hand-crafted engineering of clinically-relevant features (de la Fuente Garcia et al., <xref ref-type="bibr" rid="B12">2020</xref>). Previous work that focused on automatic AD detection from speech uses certain acoustic features (such as zero-crossing rate, Mel-frequency cepstral coefficients etc.) and linguistic features (such as proportions of various parts-of-speech (POS) tags (Orimaye et al., <xref ref-type="bibr" rid="B33">2015</xref>; Fraser et al., <xref ref-type="bibr" rid="B17">2016</xref>; Noorian et al., <xref ref-type="bibr" rid="B32">2017</xref>), etc.) from speech transcripts. Fraser et al. (<xref ref-type="bibr" rid="B17">2016</xref>) extracted 370 linguistic and acoustic features from picture descriptions in the DementiaBank dataset, and obtained an AD detection accuracy of 82% at transcript-level. Fraser et al.&#x00027;s model was evaluated using cross-validation. More recent studies showed the addition of normative data helped increase accuracy up to 93%, when evaluated using a random train/test split (Noorian et al., <xref ref-type="bibr" rid="B32">2017</xref>; Balagopalan et al., <xref ref-type="bibr" rid="B4">2018</xref>). Yancheva et al. (<xref ref-type="bibr" rid="B48">2015</xref>) showed ML models are capable of predicting the MMSE scores from features of speech elicited via picture descriptions, with mean absolute error of 2.91-3.83.</p>
<p>Detecting AD or predicting MMSE scores with pre-engineered features of speech and thereby infusing domain knowledge into the task has several advantages, such as more interpretable model decisions, the possibility to represent speech in different modalities (both acoustic and linguistic), and potentially lower computational resource requirements when paired with conventional ML models. However, there are also a few disadvantages, e.g., a feature engineering process is very expensive and time-consuming, it requires clinical expertise, is prone to biases in data, and carries the risk of missing highly relevant features.</p>
</sec>
<sec>
<title>1.2. Transfer Learning-Based Approach</title>
<p>In the recent years, transfer learning, or in other words, utilizing language representations from huge pre-trained neural models that learn robust representations for text, has become ubiquitous in NLP (Young et al., <xref ref-type="bibr" rid="B50">2018</xref>). One of the most popular transfer learning models is BERT (Devlin et al., <xref ref-type="bibr" rid="B13">2019</xref>), which trains &#x0201C;contextual embeddings&#x0201D; wherein a representation of a sentence (or transcript) is influenced by the context in which the words occur in sentences. This model offers enhanced parallelization and better modeling of long-range dependencies in text and as such, has achieved state-of-the-art performance on a variety of tasks in NLP. Previous research (Jawahar et al., <xref ref-type="bibr" rid="B21">2019</xref>; Rogers et al., <xref ref-type="bibr" rid="B41">2021</xref>) has suggested that it encodes language information (lexical, syntactic etc.) that is known to be important for performing complex natural language tasks, including AD detection from speech.</p>
<p>BERT uses powerful attention mechanisms to encode global dependencies between the input and output. This allows it to achieve state-of-the-art results on a suite of benchmarks (Devlin et al., <xref ref-type="bibr" rid="B13">2019</xref>). Fine-tuning BERT for a few epochs can potentially attain good performance even on small datasets.</p>
<p>The transfer learning technique in general and BERT model specifically are promising approaches to apply to the task of AD detection from speech because such a technique eliminates the need of expensive and time-consuming feature engineering, mitigates the need of big training datasets, and potentially results in more generalizable models. However, the common critique is that BERT is pre-trained on the corpus of healthy language and as such is not usable for detecting AD. In addition, BERT is not directly interpretable, unlike feature-based models. Finally, the original version of the BERT model is only able to use text as input, thus eliminating the possibility to employ the acoustic modality of speech, when detecting AD. All these may be the reasons why BERT was not previously used for developing predictive models for AD detection, even though its performance on many other NLP tasks is exceptional.</p>
</sec>
<sec>
<title>1.3. Motivation and Contributions</title>
<p>Our motivation in this work is to benchmark a BERT training procedure on transcripts from a pathological speech dataset, and evaluate the effectiveness of high-level language representations from BERT in detecting AD. We are specifically interested in understanding whether BERT has a potential to outperform traditional widely used domain-knowledge based approaches given that it does not include acoustic features, and at the same time increase the generalizability of the predictive models.</p>
<p>To eliminate the biases of unbalanced data, we perform all our experiments on the carefully demographically-matched ADReSS dataset. To understand how well the presented models generalize to unseen data, we evaluate performance of the models using both cross-validation and testing on unseen held out dataset.</p>
<p>We find that the feature-based SVM model with RBF kernel outperforms all the other models, and performs on par with BERT, when evaluated using cross-validation. When evaluation is performed on the unseen held out test data, the fine-tuned BERT text sequence classification models achieve the highest AD detection accuracy of 83.3%. This BERT model numerically, though not significantly, outperforms the SVM model that achieves 81.3% accuracy on the unseen test set. These results show that: (1) Extensive feature-based&#x02014;i.e., containing linguistic information for various aspects of language such as semantics, syntax, and lexicon&#x02014;classification models significantly outperforms the linguistic baseline provided in the challenge showing that feature engineering to capture various aspects of language such as semantics and syntax helps with reliable detection of AD from speech, (2) BERT proved to be a generalizable model comparable to feature-based ones that make use of domain knowledge via hand-crafted feature engineering as shown by its higher performance on the independent test set in our case, (3) linguistic-only information encoded in BERT is sufficient for the strong predictive performance of the AD detection models.</p>
</sec>
</sec>
<sec sec-type="materials and methods" id="s2">
<title>2. Materials and Methods</title>
<sec>
<title>2.1. ADReSS Dataset</title>
<p>Our data are derived from the ADReSS Challenge dataset (Luz et al., <xref ref-type="bibr" rid="B28">2020</xref>), which consists of 156 speech recordings and associated transcripts from non-AD (<italic>N</italic> = 78) and AD (<italic>N</italic> = 78) English-speaking participants. Speech is elicited from participants through the Cookie Theft picture from the Boston Diagnostic Aphasia exam (Goodglass et al., <xref ref-type="bibr" rid="B18">2001</xref>). Transcripts were annotated using the CHAT coding system (MacWhinney, <xref ref-type="bibr" rid="B30">2000</xref>). In contrast to other speech datasets for AD detection such as DementiaBank&#x00027;s English Pitt Corpus (Becker et al., <xref ref-type="bibr" rid="B5">1994</xref>), the ADReSS challenge dataset is carefully matched for age and gender in order to minimize risk of bias in the prediction tasks (<xref ref-type="table" rid="T1">Tables 1</xref>&#x02013;<xref ref-type="table" rid="T3">3</xref>). Recordings were acoustically enhanced by the challenge organizers with stationary noise removal and audio volume normalization was applied across all speech segments to control for variation caused by recording conditions such as microphone placement (Luz et al., <xref ref-type="bibr" rid="B28">2020</xref>). The speech dataset is divided into the train set and the unseen held out test set. MMSE (Cockrell and Folstein, <xref ref-type="bibr" rid="B9">2002</xref>) scores are available for all but one of the participants in the train set.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Basic characteristics of the patients in each group in the ADReSS challenge dataset are more balanced in comparison to DementiaBank.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Dataset</bold></th>
<th/>
<th/>
<th valign="top" align="center" colspan="2" style="border-bottom: thin solid #000000;"><bold>Class</bold></th>
</tr>
<tr>
<th/>
<th/>
<th/>
<th valign="top" align="center"><bold>AD</bold></th>
<th valign="top" align="center"><bold>Non-AD</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">ADReSS</td>
<td valign="top" align="left">Train</td>
<td valign="top" align="left">Male</td>
<td valign="top" align="center">24</td>
<td valign="top" align="center">24</td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="left">Female</td>
<td valign="top" align="center">30</td>
<td valign="top" align="center">30</td>
</tr>
<tr>
<td valign="top" align="left">ADReSS</td>
<td valign="top" align="left">Test</td>
<td valign="top" align="left">Male</td>
<td valign="top" align="center">11</td>
<td valign="top" align="center">11</td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="left">Female</td>
<td valign="top" align="center">13</td>
<td valign="top" align="center">13</td>
</tr>
<tr>
<td valign="top" align="left">DementiaBank (Becker et al., <xref ref-type="bibr" rid="B5">1994</xref>)</td>
<td valign="top" align="left">&#x02013;</td>
<td valign="top" align="left">Male</td>
<td valign="top" align="center">125</td>
<td valign="top" align="center">83</td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="left">Female</td>
<td valign="top" align="center">197</td>
<td valign="top" align="center">146</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>ADReSS Training set from Luz et al. (<xref ref-type="bibr" rid="B28">2020</xref>): basic characteristics of the patients in each group (M, male; F, female).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th/>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>AD</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>Non-AD</bold></th>
</tr>
<tr>
<th valign="top" align="left"><bold>Age</bold></th>
<th valign="top" align="center"><bold>M</bold></th>
<th valign="top" align="center"><bold>F</bold></th>
<th valign="top" align="center"><bold>MMSE (sd)</bold></th>
<th valign="top" align="center"><bold>M</bold></th>
<th valign="top" align="center"><bold>F</bold></th>
<th valign="top" align="center"><bold>MMSE (sd)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">[50, 55)</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">30.0 (n/a)</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">29.0 (n/a)</td>
</tr>
<tr>
<td valign="top" align="left">[55, 60)</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">16.3 (4.9)</td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">29.0 (1.3)</td>
</tr>
<tr>
<td valign="top" align="left">[60, 65)</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">18.3 (6.1)</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">29.3 (1.3)</td>
</tr>
<tr>
<td valign="top" align="left">[65, 70)</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">10</td>
<td valign="top" align="center">16.9 (5.8)</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">10</td>
<td valign="top" align="center">29.1 (0.9)</td>
</tr>
<tr>
<td valign="top" align="left">[70, 75)</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">8</td>
<td valign="top" align="center">15.8 (4.5)</td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">8</td>
<td valign="top" align="center">29.1 (0.8)</td>
</tr>
<tr>
<td valign="top" align="left">[75, 80)</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">17.2 (5.4)</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">28.8 (0.4)</td>
</tr> <tr>
<td valign="top" align="left">Total</td>
<td valign="top" align="center">24</td>
<td valign="top" align="center">30</td>
<td valign="top" align="center">17.0 (5.5)</td>
<td valign="top" align="center">24</td>
<td valign="top" align="center">30</td>
<td valign="top" align="center">29.1 (1.0)</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>ADReSS test set from Luz et al. (<xref ref-type="bibr" rid="B28">2020</xref>): basic characteristics of the patients in each group (M, male; F, female).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th/>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>AD</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>Non-AD</bold></th>
</tr>
<tr>
<th valign="top" align="left"><bold>Age</bold></th>
<th valign="top" align="center"><bold>M</bold></th>
<th valign="top" align="center"><bold>F</bold></th>
<th valign="top" align="center"><bold>MMSE (sd)</bold></th>
<th valign="top" align="center"><bold>M</bold></th>
<th valign="top" align="center"><bold>F</bold></th>
<th valign="top" align="center"><bold>MMSE (sd)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">[50, 55)</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">23.0 (n.a)</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">28.0 (n.a)</td>
</tr>
<tr>
<td valign="top" align="left">[55, 60)</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">18.7 (1.0)</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">28.5 (1.2)</td>
</tr>
<tr>
<td valign="top" align="left">[60, 65)</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">14.7 (3.7)</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">28.7 (0.9)</td>
</tr>
<tr>
<td valign="top" align="left">[65, 70)</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">23.2 (4.0)</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">29.4 (0.7)</td>
</tr>
<tr>
<td valign="top" align="left">[70, 75)</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">17.3 (6.9)</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">28.0 (2.4)</td>
</tr>
<tr>
<td valign="top" align="left">[75, 80)</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">21.5 (6.3)</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30.0 (0.0)</td>
</tr> <tr style="border-top: thin solid #000000;">
<td valign="top" align="left">Total</td>
<td valign="top" align="center">11</td>
<td valign="top" align="center">13</td>
<td valign="top" align="center">19.5 (5.3)</td>
<td valign="top" align="center">11</td>
<td valign="top" align="center">13</td>
<td valign="top" align="center">28.8 (1.5)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>2.2. Feature Extraction</title>
<p>The speech transcripts in the dataset are manually transcribed as per the CHAT protocol (MacWhinney, <xref ref-type="bibr" rid="B30">2000</xref>), and include speech segments from both the participant and an investigator. We only use the portion of the transcripts corresponding to the participant. Additionally, we combine all participant speech segments corresponding to a single picture description for extracting acoustic features.</p>
<p>We extract 509 manually-engineered features from transcripts and associated audio files (see <xref ref-type="table" rid="T4">Tables 4</xref>&#x02013;<xref ref-type="table" rid="T6">6</xref>). These features are identified as indicators of cognitive impairment in previous literature, and hence encode domain knowledge.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Summary of all lexico-syntactic features extracted.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Feature type</bold></th>
<th valign="top" align="center"><bold>&#x00023;Features</bold></th>
<th valign="top" align="left"><bold>Brief Description</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Syntactic complexity</td>
<td valign="top" align="center">36</td>
<td valign="top" align="left">L2 Analyzer features; utterance length, depth of syntactic parse tree</td>
</tr>
<tr>
<td valign="top" align="left">Production rules</td>
<td valign="top" align="center">104</td>
<td valign="top" align="left">Proportion of production type</td>
</tr>
<tr>
<td valign="top" align="left">Phrasal type ratios</td>
<td valign="top" align="center">13</td>
<td valign="top" align="left">Proportion, average length and rate of phrase types</td>
</tr>
<tr>
<td valign="top" align="left">Lexical norm-based</td>
<td valign="top" align="center">12</td>
<td valign="top" align="left">Average lexical norms across words for (e.g., imageability)</td>
</tr>
<tr>
<td valign="top" align="left">Lexical richness</td>
<td valign="top" align="center">6</td>
<td valign="top" align="left">Type-token ratios; brunet; Honor&#x00027;s statistic</td>
</tr>
<tr>
<td valign="top" align="left">Word category</td>
<td valign="top" align="center">5</td>
<td valign="top" align="left">Proportion of demonstratives, function words,</td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="left">Light verbs and inflected verbs, and propositions</td>
</tr>
<tr>
<td valign="top" align="left">Noun ratio</td>
<td valign="top" align="center">3</td>
<td valign="top" align="left">Ratios nouns:(nouns&#x0002B;verbs); nouns:verbs; pronouns:(nouns&#x0002B;pronouns)</td>
</tr>
<tr>
<td valign="top" align="left">Length measures</td>
<td valign="top" align="center">1</td>
<td valign="top" align="left">Average word length</td>
</tr>
<tr>
<td valign="top" align="left">Universal POS proportions</td>
<td valign="top" align="center">18</td>
<td valign="top" align="left">Proportions of Spacy universal POS tags</td>
</tr>
<tr>
<td valign="top" align="left">POS tag proportions</td>
<td valign="top" align="center">53</td>
<td valign="top" align="left">Proportions of Penn Treebank POS tags</td>
</tr>
<tr>
<td valign="top" align="left">Local coherence</td>
<td valign="top" align="center">15</td>
<td valign="top" align="left">Similarity between word2vec representations of utterances</td>
</tr>
<tr>
<td valign="top" align="left">Utterance distances</td>
<td valign="top" align="center">5</td>
<td valign="top" align="left">Fraction of pairs of utterances below a similarity threshold (0.5, 0.3, 0); avg/min distance</td>
</tr>
<tr>
<td valign="top" align="left">Speech-graph features</td>
<td valign="top" align="center">13</td>
<td valign="top" align="left">Representing words as nodes in a graph and computing density, number of loops, etc.</td>
</tr>
<tr>
<td valign="top" align="left">Utterance cohesion</td>
<td valign="top" align="center">1</td>
<td valign="top" align="left">Number of switches in verb tense across utterances divided by total number of utterances</td>
</tr>
<tr>
<td valign="top" align="left">Rate</td>
<td valign="top" align="center">2</td>
<td valign="top" align="left">Ratios&#x02014;number of words: duration of audio; number of syllables: duration of speech,</td>
</tr>
<tr>
<td valign="top" align="left">Invalid words</td>
<td valign="top" align="center">1</td>
<td valign="top" align="left">Proportion of words not in the English dictionary</td>
</tr>
<tr>
<td valign="top" align="left">Sentiment norm-based</td>
<td valign="top" align="center">9</td>
<td valign="top" align="left">Average sentiment norms across all words, noun, and verbs</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The number of features in each subtype is shown in the second column (titled &#x0201C;&#x00023;Features&#x0201D;)</italic>.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Summary of all acoustic/temporal features extracted.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Feature type</bold></th>
<th valign="top" align="center"><bold>&#x00023;Features</bold></th>
<th valign="top" align="left"><bold>Brief description</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Pauses and fillers</td>
<td valign="top" align="center">9</td>
<td valign="top" align="left">Total and mean duration of pauses; long and short pause counts;</td>
</tr>
<tr>
<td/>
<td/>
<td valign="top" align="left">pause to word ratio; fillers (um, uh); duration of pauses to word durations</td>
</tr>
<tr>
<td valign="top" align="left">Fundamental frequency</td>
<td valign="top" align="center">4</td>
<td valign="top" align="left">Avg/min/max/median fundamental frequency of audio</td>
</tr>
<tr>
<td valign="top" align="left">Duration-related</td>
<td valign="top" align="center">2</td>
<td valign="top" align="left">Duration of audio and spoken segment of audio</td>
</tr>
<tr>
<td valign="top" align="left">Zero-crossing rate</td>
<td valign="top" align="center">4</td>
<td valign="top" align="left">Avg/variance/skewness/kurtosis of zero-crossing rate</td>
</tr>
<tr>
<td valign="top" align="left">MFCC</td>
<td valign="top" align="center">168</td>
<td valign="top" align="left">Avg/variance/skewness/kurtosis of 42 MFCC coefficients</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The number of features in each subtype is shown in the second column (titled &#x0201C;&#x00023;Features&#x0201D;)</italic>.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>Summary of all semantic features extracted.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Feature type</bold></th>
<th valign="top" align="center"><bold>&#x00023;Features</bold></th>
<th valign="top" align="left"><bold>Brief description</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Word frequency</td>
<td valign="top" align="center">10</td>
<td valign="top" align="left">Proportion of lemmatized words occurrences</td>
</tr>
<tr>
<td valign="top" align="left">Global coherence</td>
<td valign="top" align="center">15</td>
<td valign="top" align="left">Cosine distances between word2vec utterances and content units</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The number of features in each subtype is shown in the second column (titled &#x0201C;&#x00023;Features&#x0201D;)</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>All the features are divided into three higher-level categories:</p>
<list list-type="order">
<list-item><p><bold>Lexico-syntactic features (297):</bold> Frequencies of various production rules from the constituency parsing tree of the transcripts (Chae and Nenkova, <xref ref-type="bibr" rid="B7">2009</xref>), speech-graph based features (Mota et al., <xref ref-type="bibr" rid="B31">2012</xref>), lexical norm-based features (e.g., average sentiment valence of all words in a transcript, average imageability of all words in a transcript; Warriner et al., <xref ref-type="bibr" rid="B46">2013</xref>), features indicative of lexical richness. We also extract syntactic features (Ai and Lu, <xref ref-type="bibr" rid="B2">2010</xref>) such as the proportion of various POS-tags, and similarity between consecutive utterances.</p></list-item>
<list-item><p><bold>Acoustic and temporal features (187):</bold> Mel-frequency cepstral coefficients (MFCCs), fundamental frequency, statistics related to zero-crossing rate, as well as proportion of various pauses (for example, filled and unfilled pauses, ratio of a number of pauses to a number of words etc.; Davis and Maclagan, <xref ref-type="bibr" rid="B11">2009</xref>).</p></list-item>
<list-item><p><bold>Semantic features based on picture description content (25):</bold> Proportions of various information content units used in the picture, identified as being relevant to memory impairment in prior literature (Croisile et al., <xref ref-type="bibr" rid="B10">1996</xref>).</p></list-item>
</list>
</sec>
<sec>
<title>2.3. Experiments</title>
<sec>
<title>2.3.1. AD vs. Non-AD Classification</title>
<sec>
<title>2.3.1.1. Training Regimes</title>
<p>We benchmark the following training regimes for classification: classifying features extracted at transcript-level and a BERT model fine-tuned on transcripts.</p>
<p><bold>Domain knowledge-based approach:</bold> We classify lexicosyntactic, semantic, and acoustic features extracted at transcript-level with four conventional ML models (SVM), neural network (NN), random forest (RF), na&#x000EF;ve Bayes (NB)<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref>.</p>
<p><italic>Hyperparameter tuning:</italic> All parameters in classification models were tuned to the best possible setting by searching within a grid of possible parameter values using 10-fold cross validation on the ADReSS challenge &#x0201C;train&#x0201D; set.</p>
<p>The random forest classifier fits 200 decision trees and considers <inline-formula><mml:math id="M1"><mml:msqrt><mml:mrow><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msqrt></mml:math></inline-formula> when looking for the best split. The minimum number of samples required to split an internal node is 2, and the minimum number of samples required to be at a leaf node is 2. Bootstrap samples are used when building trees. All other parameters are set to the default value.</p>
<p>The Gaussian Naive Bayes classifier is fit with balanced priors and variance smoothing coefficient set to 1<italic>e</italic> &#x02212; 10 and all other parameters default in each case.</p>
<p>The SVM is trained with a radial basis function kernel with kernel coefficient(&#x003B3;) 0.001, and regularization parameter set to 100.</p>
<p>The NN used consists of two layers of 10 units each (note we varied both the number of units and number of layers while tuning for the optimal hyperparameter setting). The ReLU activation function is used at each hidden layer. The model is trained using Adam (Kingma and Ba, <xref ref-type="bibr" rid="B24">2014</xref>) for 200 epochs and with a batch size of number of samples in train set in each fold. All other parameters are default.</p>
<p>We perform feature selection by choosing top-k number of features, based on ANOVA <italic>F</italic>-value between label/features. The number of features is jointly optimized with the classification model parameters.</p>
<p><bold>Transfer learning-based approach:</bold> In order to leverage the language information encoded by BERT (Devlin et al., <xref ref-type="bibr" rid="B13">2019</xref>), we use pre-trained model weights to initialize our classification model. All our experiments are based on the <italic>bert-base-uncased</italic> variant (Devlin et al., <xref ref-type="bibr" rid="B13">2019</xref>), which consists of 12 layers, each having a hidden size of 768 and 12 attention heads. Maximum input length is 512 tokens. Initial learning rate is set to 2<italic>e</italic> &#x02212; 5, and Adam optimizer (Kingma and Ba, <xref ref-type="bibr" rid="B24">2014</xref>) is used. Cross-entropy loss is used while fine-tuning for AD detection.</p>
<p>While the base BERT model is pre-trained with sentence pairs, our input to the model consists of speech transcripts with several transcribed utterances with start and separator special tokens from the BERT vocabulary at the beginning and end of each utterance respectively, following Liu and Lapata (<xref ref-type="bibr" rid="B27">2019</xref>). This is performed to ensure that utterance boundaries are easily encoded, since cross-utterance information such as coherence and utterance transitions is important for reliable AD detection (Fraser et al., <xref ref-type="bibr" rid="B17">2016</xref>). An embedding, following Devlin et al. (<xref ref-type="bibr" rid="B13">2019</xref>), pooling information across all tokenized units in the transcript is extracted as the aggregate transcript representation from the BERT base for each transcript. This is then passed to the classification layer, and the combined model is fine-tuned on the AD detection task&#x02014;all using an open-source PyTorch (Paszke et al., <xref ref-type="bibr" rid="B34">2019</xref>) implementation of BERT-based text sequence classification models and tokenizers (Wolf et al., <xref ref-type="bibr" rid="B47">2019</xref>). As noted by Devlin et al. (<xref ref-type="bibr" rid="B13">2019</xref>), this pooled embedding representation heavily depends on the fine-tuning task&#x02014;in our case, AD detection at transcript level.</p>
<p>The transcript input to the classification model consists of several transcribed utterances with corresponding start and end tokens for each utterance, following (Liu and Lapata, <xref ref-type="bibr" rid="B27">2019</xref>). The final hidden state corresponding to the first start (<italic>[CLS]</italic>) token in the transcript which summarizes the information across all tokens in the transcript using the self-attention mechanism in BERT is used as the aggregate representation, and passed to the classification layer (Devlin et al., <xref ref-type="bibr" rid="B13">2019</xref>; Wolf et al., <xref ref-type="bibr" rid="B47">2019</xref>). This model is then fine-tuned on training data.</p>
<p><italic>Hyperparameter tuning:</italic> We optimize the number of epochs to 10 by varying it from 1 to 12 during CV. Adam optimizer (Kingma and Ba, <xref ref-type="bibr" rid="B24">2014</xref>) and linear scheduling for the learning rate (Paszke et al., <xref ref-type="bibr" rid="B34">2019</xref>) are used. Learning rate and other parameters are set based on prior work on fine-tuning BERT (Devlin et al., <xref ref-type="bibr" rid="B13">2019</xref>; Wolf et al., <xref ref-type="bibr" rid="B47">2019</xref>).</p>
</sec>
<sec>
<title>2.3.1.2. Evaluation</title>
<p><bold>Cross-validation on ADReSS train set:</bold> We use two CV strategies in our work&#x02014;leave-one-subject-out CV (LOSO CV) and 10-fold CV at transcript level. We report evaluation metrics with LOSO CV for all models except fine-tuned BERT for direct comparison to challenge baselines. Due to computational constraints of GPU memory, we are unable to perform LOSO CV for the BERT model. Hence, we perform 10-fold CV to compare feature-based classification models with fine-tuned BERT. Values of performance metrics for each model are averaged across three runs with different random seeds in all cases.</p>
<p><bold>Predictions on ADReSS test set:</bold> We generate three predictions with different seeds from each hyperparameter-optimized classifier trained on the complete train set, and then produce a majority prediction to avoid overfitting. We report performance on the challenge test set, as obtained from the challenge organizers. We evaluate task performance primarily using accuracy scores, since all train/test sets are known to be balanced. We also report precision, recall, specificity, and F1 with respect to the positive class (AD).</p>
</sec>
</sec>
<sec>
<title>2.3.2. MMSE Score Regression</title>
<sec>
<title>2.3.2.1. Training regimes</title>
<p><bold>Domain knowledge-based approach:</bold> For this task, we benchmark two kinds of regression models, linear, and ridge, using pre-engineered features as input. MMSE scores are always within the range of 0&#x02013;30, and so predictions are clipped to a range between 0 and 30.</p>
<p><italic>Hyperparameter tuning:</italic> Each model&#x00027;s performance is optimized using hyperparameters selected via grid-search LOSO CV. We perform feature selection by choosing top-k number of features, based on an F-Score computed from the correlation of each feature with MMSE score. The number of features is optimized for all models. For ridge regression, the number of features is jointly optimized with the coefficient for L2 regularization, &#x003B1;.</p>
</sec>
<sec>
<title>2.3.2.2. Evaluation</title>
<p>We report root mean squared error (RMSE) and mean absolute error (MAE) for the predictions produced by each of the models on the training set with LOSO CV. In addition, we include the RMSE for two models&#x00027; predictions on the ADReSS test set. Hyperparameters for these models were selected based on performance in grid-search 10-fold cross validation on the training set, motivated by the thought that 10-fold CV better demonstrates how well a model will generalize to the test set.</p>
</sec>
</sec>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>3. Results</title>
<sec>
<title>3.1. AD vs. Non-AD Classification</title>
<p>In <xref ref-type="table" rid="T7">Table 7</xref>, the classification performance with all the models evaluated on the train set via 10-fold CV is displayed. We observe that BERT numerically outperforms all domain knowledge-based ML models with respect to all metrics, with an average accuracy of 81.8%. SVM is the best-performing domain knowledge-based model. However, accuracy of the fine-tuned BERT model is not significantly higher than that of the SVM classifier based on an Kruskal-Wallis <italic>H</italic>-test (<italic>H</italic> &#x0003D; 0.4838, <italic>p</italic> &#x0003E; 0.05). Note that we used a Kruskal-Wallis <italic>H</italic>-test here, and in performance-comparisons in sections below since we observe that accuracy is not normally distributed on varying the random seed while training/inference.</p>
<table-wrap position="float" id="T7">
<label>Table 7</label>
<caption><p>Ten-fold CV results averaged across three runs with different random seeds on the ADReSS train set.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>&#x00023;Features</bold></th>
<th valign="top" align="center"><bold>Accuracy</bold></th>
<th valign="top" align="center"><bold>Precision</bold></th>
<th valign="top" align="center"><bold>Recall</bold></th>
<th valign="top" align="center"><bold>Specificity</bold></th>
<th valign="top" align="center"><bold>F1</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">SVM</td>
<td valign="top" align="center">10</td>
<td valign="top" align="center">0.796</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.79</td>
</tr>
<tr>
<td valign="top" align="left">NN</td>
<td valign="top" align="center">10</td>
<td valign="top" align="center">0.762</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.76</td>
</tr>
<tr>
<td valign="top" align="left">RF</td>
<td valign="top" align="center">50</td>
<td valign="top" align="center">0.738</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.74</td>
</tr>
<tr>
<td valign="top" align="left">NB</td>
<td valign="top" align="center">80</td>
<td valign="top" align="center">0.750</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.75</td>
</tr>
<tr>
<td valign="top" align="left">BERT</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center"><bold>0.818</bold></td>
<td valign="top" align="center"><bold>0.84</bold></td>
<td valign="top" align="center"><bold>0.79</bold></td>
<td valign="top" align="center"><bold>0.85</bold></td>
<td valign="top" align="center"><bold>0.81</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Accuracy for BERT is higher, but not significantly so from SVM (H &#x0003D; 0.4838, p &#x0003E; 0.05 Kruskal-Wallis H-test). Bold indicates the best result</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>We also report the performance of all our classification models with LOSO CV (<bold>Table 9</bold>). Each of our classification models significantly outperform the challenge baseline, which is uses 34 simple language summary statistic measures (e.g., duration, total utterances, MLU, type-token ratio, percentages of nine parts of speech) on the CHAT transcripts by a large margin (&#x0002B;10% accuracy for the best performing model, <italic>p</italic> &#x0003D; 0.036 with Kruskal-Wallis H = 4.35 test). Feature selection results in accuracy increase of about 13% for the SVM classifier.</p>
<p>Performance results on the unseen, held out challenge test set are shown in <xref ref-type="table" rid="T8">Table 8</xref> and follow the trend of the cross-validated performance in terms of accuracy, with BERT outperforming the best feature-based classification model SVM with an accuracy of 83.33%, but not significantly so (<italic>H</italic> &#x0003D; 2.4, <italic>p</italic> &#x0003E; 0.05). The accuracy with a BERT-based classification model ranges between 85.14 and 81.25%.</p>
<table-wrap position="float" id="T8">
<label>Table 8</label>
<caption><p>AD detection results on unseen, held out ADReSS test set averaged over three runs with different random seeds.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>&#x00023;Features</bold></th>
<th valign="top" align="center"><bold>Accuracy</bold></th>
<th valign="top" align="center"><bold>Precision</bold></th>
<th valign="top" align="center"><bold>Recall</bold></th>
<th valign="top" align="center"><bold>Specificity</bold></th>
<th valign="top" align="center"><bold>F1</bold></th>
<th valign="top" align="center"><bold>AUROC</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Baseline (Luz et al., <xref ref-type="bibr" rid="B28">2020</xref>)</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.7500</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.7800</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">SVM</td>
<td valign="top" align="center">10</td>
<td valign="top" align="center">0.8125</td>
<td valign="top" align="center">0.8000</td>
<td valign="top" align="center">0.8333</td>
<td valign="top" align="center">0.7917</td>
<td valign="top" align="center">0.8124</td>
<td valign="top" align="center">0.8125</td>
</tr>
<tr>
<td valign="top" align="left">NN</td>
<td valign="top" align="center">10</td>
<td valign="top" align="center">0.7708</td>
<td valign="top" align="center">0.7671</td>
<td valign="top" align="center">0.7778</td>
<td valign="top" align="center">0.7639</td>
<td valign="top" align="center">0.7708</td>
<td valign="top" align="center">0.7708</td>
</tr>
<tr>
<td valign="top" align="left">RF</td>
<td valign="top" align="center">50</td>
<td valign="top" align="center">0.7569</td>
<td valign="top" align="center">0.8033</td>
<td valign="top" align="center">0.6806</td>
<td valign="top" align="center">0.8333</td>
<td valign="top" align="center">0.7555</td>
<td valign="top" align="center">0.7500</td>
</tr>
<tr>
<td valign="top" align="left">NB</td>
<td valign="top" align="center">80</td>
<td valign="top" align="center">0.7292</td>
<td valign="top" align="center">0.7895</td>
<td valign="top" align="center">0.6250</td>
<td valign="top" align="center">0.8333</td>
<td valign="top" align="center">0.7262</td>
<td valign="top" align="center">0.7292</td>
</tr>
<tr>
<td valign="top" align="left">BERT</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center"><bold>0.8332</bold></td>
<td valign="top" align="center"><bold>0.8389</bold></td>
<td valign="top" align="center"><bold>0.8333</bold></td>
<td valign="top" align="center"><bold>0.8333</bold></td>
<td valign="top" align="center"><bold>0.8327</bold></td>
<td valign="top" align="center"><bold>0.8333</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Bold indicates the best result</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>3.2. MMSE Score Regression</title>
<p>Performance of regression models evaluated on both train and test sets is shown in <xref ref-type="table" rid="T9">Table 9</xref>. Ridge regression with 25 features selected attains the lowest RMSE of 4.56 (with a corresponding MAE of 3.50, or 11.67% error) during LOSO-CV on the training set. The results show that feature selection is impactful for performance and helps achieve a decrease of up to 1.5 RMSE points (and up to 0.86 of MAE) for a ridge regressor. Furthermore, a ridge regressor is able to achieve an RMSE of 4.56 on the ADReSS test set, a decrease of 0.64 from the baseline. We also experimented with different non-linear regression methods&#x02014;however, given the small dataset size and the difficulty of the task, the linear regression models highlighted in <xref ref-type="table" rid="T9">Table 9</xref> performed the best.</p>
<table-wrap position="float" id="T9">
<label>Table 9</label>
<caption><p>LOSO-CV MMSE regression results on the ADReSS train and test sets.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>&#x00023;Features</bold></th>
<th valign="top" align="center"><bold>&#x003B1;</bold></th>
<th valign="top" align="center" style="border-bottom: thin solid #000000;"><bold>RMSE</bold></th>
<th valign="top" align="center" style="border-bottom: thin solid #000000;"><bold>MAE</bold></th>
<th valign="top" align="center" style="border-bottom: thin solid #000000;"><bold>RMSE</bold></th>
</tr>
<tr>
<th/>
<th/>
<th/>
<th valign="top" align="center" colspan="2"><bold>Train set</bold></th>
<th valign="top" align="center"><bold>Test set</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Baseline (Luz et al., <xref ref-type="bibr" rid="B28">2020</xref>)</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">4.38</td>
<td/>
<td valign="top" align="center">5.20</td>
</tr>
<tr>
<td valign="top" align="left">LR</td>
<td valign="top" align="center">15</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">5.37</td>
<td valign="top" align="center">4.18</td>
<td valign="top" align="center">4.94</td>
</tr>
<tr>
<td valign="top" align="left">LR</td>
<td valign="top" align="center">20</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">4.94</td>
<td valign="top" align="center">3.72</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Ridge</td>
<td valign="top" align="center">509</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">6.06</td>
<td valign="top" align="center">4.36</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Ridge</td>
<td valign="top" align="center">35</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">4.87</td>
<td valign="top" align="center">3.79</td>
<td valign="top" align="center"><bold>4.56</bold></td>
</tr>
<tr>
<td valign="top" align="left">Ridge</td>
<td valign="top" align="center">25</td>
<td valign="top" align="center">10</td>
<td valign="top" align="center"><bold>4.56</bold></td>
<td valign="top" align="center"><bold>3.50</bold></td>
<td valign="top" align="center">&#x02013;</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Bold indicates the best result</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>4. Discussion</title>
<sec>
<title>4.1. Feature Differentiation Analysis</title>
<p>While we extracted a large number of linguistic and acoustic features to capture a wide range of linguistic and acoustic changes in speech associated with AD, based on a survey of prior literature (Yancheva et al., <xref ref-type="bibr" rid="B48">2015</xref>; Fraser et al., <xref ref-type="bibr" rid="B17">2016</xref>; Pou-Prom and Rudzicz, <xref ref-type="bibr" rid="B36">2018</xref>; Zhu et al., <xref ref-type="bibr" rid="B54">2019</xref>), we are also interested in identifying the <italic>most differentiating</italic> features between AD and non-AD speech. In order to study statistically significant differences in linguistic/acoustic phenomena, we perform independent <italic>t</italic>-tests between feature means for each class in the ADReSS training set, following the methodology followed by Eyre et al. (<xref ref-type="bibr" rid="B15">2020</xref>). 87 features are significantly different between the two groups at <italic>p</italic> &#x0003C; 0.05. Seventy-nine of these are text-based lexicosyntactic and semantic features, while eight are acoustic. These eight acoustic features include the number of long pauses, pause duration, and mean/skewness/variance-statistics of various MFCC coefficients. However, after Bonferroni correction for multiple testing, we identify that only 13 features are significantly different between AD and non-AD speech at <italic>p</italic> &#x0003C; 9<italic>e</italic> &#x02212; 5, and none of these features are acoustic (<xref ref-type="table" rid="T10">Table 10</xref>). This implies that linguistic features are particularly differentiating between the AD/non-AD classes here, which explains why models trained only on linguistic features (i.e., BERT models) attain performance well above random chance.</p>
<table-wrap position="float" id="T10">
<label>Table 10</label>
<caption><p>Feature differentiation analysis results for the most important features, based on ADReSS train set.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Feature</bold></th>
<th valign="top" align="left"><bold>Feature type</bold></th>
<th valign="top" align="center"><bold>&#x003BC;<sub><italic>AD</italic></sub></bold></th>
<th valign="top" align="center"><bold>&#x003BC;<sub><italic>non</italic>&#x02212;<italic>AD</italic></sub></bold></th>
<th valign="top" align="center"><bold>Correlation</bold></th>
<th valign="top" align="center"><bold>Weight</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Average cosine distance between utterances</td>
<td valign="top" align="left">Semantic</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.94</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Fraction of pairs of utterances below a similarity threshold (0.5)</td>
<td valign="top" align="left">Semantic</td>
<td valign="top" align="center">0.03</td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Cosine distance between word2vec utterances and content units</td>
<td valign="top" align="left">Semantic</td>
<td valign="top" align="center">0.46</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">&#x02212;0.54<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">&#x02212;1.01</td>
</tr>
<tr>
<td valign="top" align="left">Distinct content units mentioned: total content units</td>
<td valign="top" align="left">Semantic</td>
<td valign="top" align="center">0.27</td>
<td valign="top" align="center">0.45</td>
<td valign="top" align="center">0.63<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">1.78</td>
</tr>
<tr>
<td valign="top" align="left">Distinct action content units mentioned: total content units</td>
<td valign="top" align="left">Semantic</td>
<td valign="top" align="center">0.15</td>
<td valign="top" align="center">0.30</td>
<td valign="top" align="center">0.49<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">1.04</td>
</tr>
<tr>
<td valign="top" align="left">Distinct object content units mentioned: total content units</td>
<td valign="top" align="left">Semantic</td>
<td valign="top" align="center">0.28</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">0.59<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">1.72</td>
</tr>
<tr>
<td valign="top" align="left">Cosine distance between GloVe utterances and content units</td>
<td valign="top" align="left">Semantic</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02212;0.42<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">&#x02212;0.03</td>
</tr>
<tr>
<td valign="top" align="left">Average word length (in letters)</td>
<td valign="top" align="left">Lexico-syntactic</td>
<td valign="top" align="center">3.57</td>
<td valign="top" align="center">3.78</td>
<td valign="top" align="center">0.45<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">1.07</td>
</tr>
<tr>
<td valign="top" align="left">Proportion of pronouns</td>
<td valign="top" align="left">Lexico-syntactic</td>
<td valign="top" align="center">0.09</td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Ratio (pronouns):(pronouns&#x0002B;nouns)</td>
<td valign="top" align="left">Lexico-syntactic</td>
<td valign="top" align="center">0.35</td>
<td valign="top" align="center">0.23</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Proportion of personal pronouns</td>
<td valign="top" align="left">Lexico-syntactic</td>
<td valign="top" align="center">0.09</td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Proportion of adverbs</td>
<td valign="top" align="left">Lexico-syntactic</td>
<td valign="top" align="center">0.06</td>
<td valign="top" align="center">0.04</td>
<td valign="top" align="center">&#x02212;0.41<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">&#x02212;0.41</td>
</tr>
<tr>
<td valign="top" align="left">Proportion of adverbial phrases amongst all rules</td>
<td valign="top" align="left">Lexico-syntactic</td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">&#x02212;0.37</td>
<td valign="top" align="center">&#x02212;0.74</td>
</tr>
<tr>
<td valign="top" align="left">Proportion of non-dictionary words</td>
<td valign="top" align="left">Lexico-syntactic</td>
<td valign="top" align="center">0.11</td>
<td valign="top" align="center">0.08</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
</tr>
<tr>
<td valign="top" align="left">Proportion of gerund verbs</td>
<td valign="top" align="left">Lexico-syntactic</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.37</td>
<td valign="top" align="center">1.08</td>
</tr>
<tr>
<td valign="top" align="left">Proportion of words in adverb category</td>
<td valign="top" align="left">Lexico-syntactic</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">&#x02212;0.4<xref ref-type="table-fn" rid="TN1"><sup>&#x0002A;</sup></xref></td>
<td valign="top" align="center">&#x02212;0.49</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>&#x003BC;<sub>AD</sub> and &#x003BC;<sub>non&#x02212;AD</sub> show the means of the 13 significantly different features at p &#x0003C; 9e-5 (after Bonferroni correction) for the AD and non-AD group, respectively. We also show Spearman correlation between MMSE score and features, and regression weights of the features associated with the five greatest and five lowest regression weights from our regression experiments</italic>.</p>
<fn id="TN1">
<label>&#x0002A;</label>
<p><italic>Next to correlation indicates significance at p &#x0003C; 9e-5</italic>.</p></fn>
</table-wrap-foot>
</table-wrap>
<p>The features that differentiate the AD and non-AD groups largely indicate semantic impairments in AD, reflected in the types of words used and the content of their picture descriptions. Importantly, many of the differentiating features replicate findings from Fraser et al. (<xref ref-type="bibr" rid="B17">2016</xref>), suggesting that despite the present dataset being more demographically balanced, many of the previous findings maintain. In addition, the differentiating features are consistent with other previous clinical literature documenting decreased specificity and information content in AD. For example, the features relating to the content units in the picture and the cosine similarity between utterances and picture content units show that the picture descriptions produced in AD have fewer relevant content words and that the words used are less semantically related to the themes of the picture. Lower average cosine distance in AD signifies more repetition in speech. These findings are consistent with previous studies reporting reduced information content and coherence in AD (Croisile et al., <xref ref-type="bibr" rid="B10">1996</xref>; Snowdon et al., <xref ref-type="bibr" rid="B43">1996</xref>; Dijkstra et al., <xref ref-type="bibr" rid="B14">2004</xref>; Forbes-McKay and Venneri, <xref ref-type="bibr" rid="B16">2005</xref>; Riley et al., <xref ref-type="bibr" rid="B40">2005</xref>; Le et al., <xref ref-type="bibr" rid="B26">2011</xref>; Ahmed et al., <xref ref-type="bibr" rid="B1">2013</xref>; Boschi et al., <xref ref-type="bibr" rid="B6">2017</xref>). Other differentiating features related to the use of shorter words, and increased use of pronouns, adverbs, and words not found in the dictionary. These features may all reflect the use of less specific and simpler language, and replicate previous findings of decreased specificity of language in AD (Le et al., <xref ref-type="bibr" rid="B26">2011</xref>; Ahmed et al., <xref ref-type="bibr" rid="B1">2013</xref>; Szatloczki et al., <xref ref-type="bibr" rid="B44">2015</xref>; Fraser et al., <xref ref-type="bibr" rid="B17">2016</xref>). Interestingly, while Fraser et al. (<xref ref-type="bibr" rid="B17">2016</xref>) found differences in acoustic features, none of those findings survived Bonferroni correction in the present study, which may indicate that this age/sex-balanced dataset reduced the acoustic differences between groups.</p>
<p>In order to visualize the class-separability of the feature-based representations, we visualize (t-SNE) t-Distributed Stochastic Neighbor Embedding (Maaten and Hinton, <xref ref-type="bibr" rid="B29">2008</xref>) plots in <xref ref-type="fig" rid="F1">Figure 1</xref>. t-SNE is a non-linear dimensionality reduction algorithm used for exploring high-dimensional data. It maps multi-dimensional data to two or more dimensions suitable for human observation. We observe strong class-separation between the two classes, indicating that a non-linear model would be capable of good AD detection performance with these representations.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>A t-SNE plot showing class separation. Note we only use the 13 features significantly different between classes (see <xref ref-type="table" rid="T10">Table 10</xref>) in feature representation for this plot.</p></caption>
<graphic xlink:href="fnagi-13-635945-g0001.tif"/>
</fig>
</sec>
<sec>
<title>4.2. Interpreting Attention Patterns in BERT-Based Models</title>
<p>We look at multi-scale attention visualizations of BERT fine-tuned for the AD detection task, using the BertViz library (Vig, <xref ref-type="bibr" rid="B45">2019</xref>) (<xref ref-type="fig" rid="F2">Figure 2</xref>). Self-attention is an important component of BERT-based models, and looking at attention patterns can help us interpret model decisions. We used the BERT-base model which consists of 12 layers, and 12 attention heads in each layer. We visualize, for both AD and healthy speech transcripts, the attention weights for the final &#x0201C;[CLS]&#x0201D; token, whose representation is passed to the fully-connected layer for classification. On analyzing the attention weights attributed to words in both healthy and AD transcripts, we find that:</p>
<list list-type="order">
<list-item><p>attention weights are often attributed to a few important &#x0201C;information content units.&#x0201D; which have been identified to be important speech indicators of AD in prior work (Fraser et al., <xref ref-type="bibr" rid="B17">2016</xref>) such as &#x0201C;water,&#x0201D; &#x0201C;boy,&#x0201D; etc.</p></list-item>
<list-item><p>attention weights are also sometimes attributed to pauses and fillers, such as &#x0201C;uh&#x0201D; and &#x0201C;um.&#x0201D;</p></list-item>
<list-item><p>attention weights are also attributed to the sentence separator tokens, and we think this approximates to roughly counting the number of utterances in the transcript.</p></list-item>
</list>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>An attention visualization plot showing attention contributions of embeddings corresponding to each word to the &#x0201C;pooled&#x0201D; representation. This example is a sub-sample (first two utterances) of a speech transcript from a healthy person.</p></caption>
<graphic xlink:href="fnagi-13-635945-g0002.tif"/>
</fig>
<p>Hence, as seen in sections 4.1 and 4.2, we observe that for both the feature-based classification models and BERT-based models, information units and fillers such as &#x0201C;uh&#x0201D; and &#x0201C;um&#x0201D; seem to be important predictors, similar to findings observed by Yuan et al. (<xref ref-type="bibr" rid="B52">2020</xref>).</p>
</sec>
<sec>
<title>4.3. Analysing AD Detection Performance Differences</title>
<p>We observe that both feature-based and BERT-based classification models are significantly better than the linguistic baseline, showing the importance of an extensive amount of linguistic features for detecting AD-related differences. When compared on this well-matched dataset, BERT tended to have higher performance, but the difference was not significant. Based on feature differentiation analysis, we hypothesize that good performance with a text-focused BERT model on this speech classification task is due to the strong utility of linguistic features on this dataset. BERT captures a wide range of linguistic phenomena due to its training methodology, potentially encapsulating most of the important lexico-syntactic and semantic features. It is thus able to use information present in the lexicon, syntax, and semantics of the transcribed speech after fine-tuning (Jawahar et al., <xref ref-type="bibr" rid="B21">2019</xref>).</p>
<p>We also see a trend of better performance when increasing the number of folds (see SVM in <xref ref-type="table" rid="T7">Tables 7</xref>, <xref ref-type="table" rid="T11">11</xref>) in cross-validation. We postulate that this is due to the small size of the dataset, and hence differences in training set size in each fold (<italic>N</italic><sub><italic>train</italic></sub> &#x0003D; 107 with LOSO, <italic>N</italic><sub><italic>train</italic></sub> &#x0003D; 98 with 10-fold CV). Note that, in this dataset, both feature-based and BERT-based classification methods rely on linguistic features to achieve better classification than baseline. This implies that the linguistic features from speech transcripts are quite informative for the AD detection task. Hence, an interesting direction of future research is expanding our current set of features to incorporate more discourse-related features (which could be getting captured to some degree in fine-tuned BERT models).</p>
<table-wrap position="float" id="T11">
<label>Table 11</label>
<caption><p>LOSO-CV results averaged across three runs with different random seeds on the ADReSS train set.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>&#x00023;Features</bold></th>
<th valign="top" align="center"><bold>Accuracy</bold></th>
<th valign="top" align="center"><bold>Precision</bold></th>
<th valign="top" align="center"><bold>Recall</bold></th>
<th valign="top" align="center"><bold>Specificity</bold></th>
<th valign="top" align="center"><bold>F1</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Baseline (Luz et al., <xref ref-type="bibr" rid="B28">2020</xref>)</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.768</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">&#x02013;</td>
<td valign="top" align="center">0.77</td>
</tr>
<tr>
<td valign="top" align="left">SVM</td>
<td valign="top" align="center">509</td>
<td valign="top" align="center">0.741</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.74</td>
</tr>
<tr>
<td valign="top" align="left">SVM</td>
<td valign="top" align="center">10</td>
<td valign="top" align="center"><bold>0.870</bold></td>
<td valign="top" align="center"><bold>0.90</bold></td>
<td valign="top" align="center"><bold>0.83</bold></td>
<td valign="top" align="center"><bold>0.91</bold></td>
<td valign="top" align="center"><bold>0.87</bold></td>
</tr>
<tr>
<td valign="top" align="left">NN</td>
<td valign="top" align="center">10</td>
<td valign="top" align="center">0.836</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.83</td>
</tr>
<tr>
<td valign="top" align="left">RF</td>
<td valign="top" align="center">50</td>
<td valign="top" align="center">0.778</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">0.78</td>
</tr>
<tr>
<td valign="top" align="left">NB</td>
<td valign="top" align="center">80</td>
<td valign="top" align="center">0.787</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.78</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Accuracy for SVM is significantly higher than NN (H &#x0003D; 4.50, p &#x0003D; 0.034 Kruskal-Wallis H-test). Bold indicates the best result</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>4.4. Regression Weights for MMSE Prediction</title>
<p>To assess the relative importance of individual input features for MMSE prediction, we report features with the five highest and five lowest regression weights reflecting the five strongest positive and negative relationships with MMSE scores (<xref ref-type="table" rid="T10">Table 10</xref>). Each presented value is the average weight assigned to that feature across each of the LOSO CV folds. We also present the correlation with MMSE score coefficients for those 10 features, as well as their significance, in <xref ref-type="table" rid="T10">Table 10</xref>. We observe that for each of these highly weighted features, a positive or negative correlation coefficient is accompanied by a positive or negative regression weight, respectively. This demonstrates that these 10 features are so distinguishing that, even in the presence of other regressors, their relationship with MMSE score remains the same. We also note that all 10 of these are linguistic features, further demonstrating that linguistic information is particularly distinguishing when it comes to predicting the severity of a patient&#x00027;s AD. Notably, seven of the ten features were among those that differentiated between AD and non-AD groups, demonstrating that there is high overlap between the features relevant to group differentiation and MMSE score prediction. These features included those relating to the information content and the coherence of picture descriptions, reflected by content unit and cosine distance features. Word length and use of adverbs were also relevant to MMSE prediction, with longer words and fewer adverbs correlating with higher MMSE scores. The use of gerund verbs was found to have a high regression weight for MMSE prediction and positively correlated with MMSE scores, despite not being significantly different between AD and non-AD groups after Bonferroni correction. Reduced use of inflected verbs has been found in some previous research (Ahmed et al., <xref ref-type="bibr" rid="B1">2013</xref>; Fraser et al., <xref ref-type="bibr" rid="B17">2016</xref>), and is thought to reflect an grammatic impairment.</p>
</sec>
</sec>
<sec sec-type="conclusions" id="s5">
<title>5. Conclusions</title>
<p>In this paper, we rigorously compare two widely used approaches&#x02014;linguistic and acoustic feature engineering based on domain knowledge, and text-only transfer learning using fine-tuned BERT classification model. Our results show that pre-trained models that are fine-tuned for the AD classification task are capable of performing well on AD detection, achieving comparable, or even slightly improved performance compared to hand-crafted feature engineering. We observe that linguistic features are capable of attaining predictive performance well above chance on this acoustically and demographically balanced speech dataset, and posit this to be the reason why a text-only approach with BERT numerically outperforms a multi-modal feature-engineering based approach. The present findings highlight the importance of measuring the linguistic, and especially semantic content of speech, in addition to acoustic analyses. In future work, it would be interesting to study methods that combine feature-based and pre-trained neural LM-based prediction models to optimize AD detection from speech&#x02014;this could potentially help harness complementary benefits of both approaches. It is interesting to note that the winners of the ADReSS challenge also used a pre-trained language model, augmented with additional information about speech disfluencies (Yuan et al., <xref ref-type="bibr" rid="B52">2020</xref>), which outperforms our best model by 6% in accuracy and F1-score, further indicating the degree of promise in such an approach. These results build on previous work to demonstrate how automated speech analysis can be used to help characterize AD. Speech samples can be collected quickly and non-invasively, and as demonstrated in the present results, yield measures relating to the presence and severity of AD.</p>
<p>Further work will build on these results to develop improved tools for disease screening and monitoring in AD, improving the efficiency of clinical research and treatment. In the future, we will experiment with different neural models such as XLNet (Yang et al., <xref ref-type="bibr" rid="B49">2019</xref>), and with different tokenization and encoding strategies for transcript representations. A direction for future work is developing ML models that combine representations from BERT and hand-crafted features (Yu et al., <xref ref-type="bibr" rid="B51">2015</xref>). Such feature-fusion approaches could potentially boost performance on the cognitive impairment detection task.</p>
</sec>
<sec sec-type="data-availability-statement" id="s6">
<title>Data Availability Statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found at: <ext-link ext-link-type="uri" xlink:href="https://dementia.talkbank.org/">https://dementia.talkbank.org/</ext-link>.</p>
</sec>
<sec id="s7">
<title>Ethics Statement</title>
<p>The studies involving human participants were reviewed and approved by DementiaBank consortium. The patients/participants provided their written informed consent to participate in this study.</p>
</sec>
<sec id="s8">
<title>Author Contributions</title>
<p>All authors contributed to writing and edits. Methods and analyses were performed by AB, JN, and BE.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>Authors AB, BE, JR and JN were employed by company Winterlight Labs Inc. The remaining author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<ack><p>The results shown in this manuscript were also presented at INTERSPEECH, 2020 as a part of the ADReSS challenge track (Balagopalan et al., <xref ref-type="bibr" rid="B3">2020</xref>).</p>
</ack>

<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ahmed</surname> <given-names>S.</given-names></name> <name><surname>Haigh</surname> <given-names>A.-M. F.</given-names></name> <name><surname>de Jager</surname> <given-names>C. A.</given-names></name> <name><surname>Garrard</surname> <given-names>P.</given-names></name></person-group> (<year>2013</year>). <article-title>Connected speech as a marker of disease progression in autopsy-proven Alzheimer&#x00027;s disease</article-title>. <source>Brain</source> <volume>136</volume>, <fpage>3727</fpage>&#x02013;<lpage>3737</lpage>. <pub-id pub-id-type="doi">10.1093/brain/awt269</pub-id><pub-id pub-id-type="pmid">24142144</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ai</surname> <given-names>H.</given-names></name> <name><surname>Lu</surname> <given-names>X.</given-names></name></person-group> (<year>2010</year>). <article-title>A web-based system for automatic measurement of lexical complexity,</article-title> in <source>27th Annual Symposium of the Computer-Assisted Language Consortium (CALICO-10)</source> (<publisher-loc>Amherst, MA</publisher-loc>), <fpage>8</fpage>&#x02013;<lpage>12</lpage>.</citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Balagopalan</surname> <given-names>A.</given-names></name> <name><surname>Eyre</surname> <given-names>B.</given-names></name> <name><surname>Rudzicz</surname> <given-names>F.</given-names></name> <name><surname>Novikova</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>To BERT or not to BERT: comparing speech and language-based approaches for Alzheimer&#x00027;s disease detection</article-title>. <source>Proc. Interspeech</source> <volume>2020</volume>, <fpage>2167</fpage>&#x02013;<lpage>2171</lpage>. <pub-id pub-id-type="doi">10.21437/Interspeech.2020-2557</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Balagopalan</surname> <given-names>A.</given-names></name> <name><surname>Novikova</surname> <given-names>J.</given-names></name> <name><surname>Rudzicz</surname> <given-names>F.</given-names></name> <name><surname>Ghassemi</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>The effect of heterogeneous data for Alzheimer&#x00027;s disease detection from speech</article-title>. <source>arXiv preprint arXiv:1811.12254</source>.</citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Becker</surname> <given-names>J. T.</given-names></name> <name><surname>Boiler</surname> <given-names>F.</given-names></name> <name><surname>Lopez</surname> <given-names>O. L.</given-names></name> <name><surname>Saxton</surname> <given-names>J.</given-names></name> <name><surname>McGonigle</surname> <given-names>K. L.</given-names></name></person-group> (<year>1994</year>). <article-title>The natural history of Alzheimer&#x00027;s disease: description of study cohort and accuracy of diagnosis</article-title>. <source>Arch. Neurol</source>. <volume>51</volume>, <fpage>585</fpage>&#x02013;<lpage>594</lpage>. <pub-id pub-id-type="doi">10.1001/archneur.1994.00540180063015</pub-id><pub-id pub-id-type="pmid">8198470</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Boschi</surname> <given-names>V.</given-names></name> <name><surname>Catricala</surname> <given-names>E.</given-names></name> <name><surname>Consonni</surname> <given-names>M.</given-names></name> <name><surname>Chesi</surname> <given-names>C.</given-names></name> <name><surname>Moro</surname> <given-names>A.</given-names></name> <name><surname>Cappa</surname> <given-names>S. F.</given-names></name></person-group> (<year>2017</year>). <article-title>Connected speech in neurodegenerative language disorders: a review</article-title>. <source>Front. Psychol</source>. <volume>8</volume>:<fpage>269</fpage>. <pub-id pub-id-type="doi">10.3389/fpsyg.2017.00269</pub-id><pub-id pub-id-type="pmid">28321196</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chae</surname> <given-names>J.</given-names></name> <name><surname>Nenkova</surname> <given-names>A.</given-names></name></person-group> (<year>2009</year>). <article-title>Predicting the fluency of text with shallow structural features: case studies of machine translation and human-written text,</article-title> in <source>Proceedings of the 12th Conference of the European Chapter of the ACL (EACL 2009)</source> (<publisher-loc>Athens</publisher-loc>), <fpage>139</fpage>&#x02013;<lpage>147</lpage>. <pub-id pub-id-type="doi">10.3115/1609067.1609082</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>L.</given-names></name> <name><surname>Dodge</surname> <given-names>H. H.</given-names></name> <name><surname>Asgari</surname> <given-names>M.</given-names></name></person-group> (<year>2020</year>). <article-title>Topic-based measures of conversation for detecting mild cognitive impairment,</article-title> in <source>Proceedings of the First Workshop on Natural Language Processing for Medical Conversations</source> (<publisher-loc>Virtual</publisher-loc>), <fpage>63</fpage>&#x02013;<lpage>67</lpage>. <pub-id pub-id-type="pmid">33642674</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cockrell</surname> <given-names>J. R.</given-names></name> <name><surname>Folstein</surname> <given-names>M. F.</given-names></name></person-group> (<year>2002</year>). <article-title>Mini-mental state examination</article-title>. <source>Princ. Pract. Geriatr. Psychiatry</source>, <fpage>140</fpage>&#x02013;<lpage>141</lpage>. <pub-id pub-id-type="doi">10.1002/0470846410.ch27(ii)</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Croisile</surname> <given-names>B.</given-names></name> <name><surname>Ska</surname> <given-names>B.</given-names></name> <name><surname>Brabant</surname> <given-names>M.-J.</given-names></name> <name><surname>Duchene</surname> <given-names>A.</given-names></name> <name><surname>Lepage</surname> <given-names>Y.</given-names></name> <name><surname>Aimard</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>1996</year>). <article-title>Comparative study of oral and written picture description in patients with Alzheimer&#x00027;s disease</article-title>. <source>Brain Lang</source>. <volume>53</volume>, <fpage>1</fpage>&#x02013;<lpage>19</lpage>. <pub-id pub-id-type="doi">10.1006/brln.1996.0033</pub-id><pub-id pub-id-type="pmid">8722896</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Davis</surname> <given-names>B. H.</given-names></name> <name><surname>MacLagan</surname> <given-names>M.</given-names></name></person-group> (<year>2009</year>). <article-title>Examining pauses in Alzheimer&#x00027;s discourse</article-title>. <source>Am. J. Alzheimer&#x00027;s Dis. Other Dement</source>. <volume>24</volume>, <fpage>141</fpage>&#x02013;<lpage>154</lpage>. <pub-id pub-id-type="doi">10.1177/1533317508328138</pub-id><pub-id pub-id-type="pmid">19150969</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>de la Fuente Garcia</surname> <given-names>S.</given-names></name> <name><surname>Ritchie</surname> <given-names>C.</given-names></name> <name><surname>Luz</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Artificial intelligence, speech, and language processing approaches to monitoring Alzheimer&#x00027;s disease: a systematic review</article-title>. <source>J. Alzheimer&#x00027;s Dis</source>. <volume>78</volume>, <fpage>1547</fpage>&#x02013;<lpage>1574</lpage>. <pub-id pub-id-type="doi">10.3233/JAD-200888</pub-id><pub-id pub-id-type="pmid">33185605</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Devlin</surname> <given-names>J.</given-names></name> <name><surname>Chang</surname> <given-names>M.-W.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>Toutanova</surname> <given-names>K.</given-names></name></person-group> (<year>2019</year>). <article-title>Bert: pre-training of deep bidirectional transformers for language understanding,</article-title> in <source>Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Volume 1 (Long and Short Papers)</source> (<publisher-loc>Minneapolis, MN</publisher-loc>), <fpage>4171</fpage>&#x02013;<lpage>4186</lpage>.</citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dijkstra</surname> <given-names>K.</given-names></name> <name><surname>Bourgeois</surname> <given-names>M. S.</given-names></name> <name><surname>Allen</surname> <given-names>R. S.</given-names></name> <name><surname>Burgio</surname> <given-names>L. D.</given-names></name></person-group> (<year>2004</year>). <article-title>Conversational coherence: discourse analysis of older adults with and without dementia</article-title>. <source>J. Neurolinguist</source>. <volume>17</volume>, <fpage>263</fpage>&#x02013;<lpage>283</lpage>. <pub-id pub-id-type="doi">10.1016/S0911-6044(03)00048-4</pub-id><pub-id pub-id-type="pmid">8558875</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Eyre</surname> <given-names>B.</given-names></name> <name><surname>Balagopalan</surname> <given-names>A.</given-names></name> <name><surname>Novikova</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Fantastic features and where to find them: detecting cognitive impairment with a subsequence classification guided approach,</article-title> in <source>Proceedings of the Sixth Workshop on Noisy User-Generated Text (W-NUT 2020)</source> (<publisher-loc>Virtual</publisher-loc>), <fpage>193</fpage>&#x02013;<lpage>199</lpage>. <pub-id pub-id-type="doi">10.18653/v1/2020.wnut-1.25</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Forbes-McKay</surname> <given-names>K. E.</given-names></name> <name><surname>Venneri</surname> <given-names>A.</given-names></name></person-group> (<year>2005</year>). <article-title>Detecting subtle spontaneous language decline in early Alzheimer&#x00027;s disease with a picture description task</article-title>. <source>Neurol. Sci</source>. <volume>26</volume>, <fpage>243</fpage>&#x02013;<lpage>254</lpage>. <pub-id pub-id-type="doi">10.1007/s10072-005-0467-9</pub-id><pub-id pub-id-type="pmid">16193251</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fraser</surname> <given-names>K. C.</given-names></name> <name><surname>Meltzer</surname> <given-names>J. A.</given-names></name> <name><surname>Rudzicz</surname> <given-names>F.</given-names></name></person-group> (<year>2016</year>). <article-title>Linguistic features identify Alzheimer&#x00027;s disease in narrative speech</article-title>. <source>J. Alzheimer&#x00027;s Dis</source>. <volume>49</volume>, <fpage>407</fpage>&#x02013;<lpage>422</lpage>. <pub-id pub-id-type="doi">10.3233/JAD-150520</pub-id><pub-id pub-id-type="pmid">26484921</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Goodglass</surname> <given-names>H.</given-names></name> <name><surname>Kaplan</surname> <given-names>E.</given-names></name> <name><surname>Barresi</surname> <given-names>B.</given-names></name></person-group> (<year>2001</year>). <source>BDAE-3: Boston Diagnostic Aphasia Examination, 3rd Edn</source>. <publisher-loc>Philadelphia, PA</publisher-loc>: <publisher-name>Lippincott Williams &#x00026; Wilkins</publisher-name>.</citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gosztolya</surname> <given-names>G.</given-names></name> <name><surname>Vincze</surname> <given-names>V.</given-names></name> <name><surname>T&#x000F3;th</surname> <given-names>L.</given-names></name> <name><surname>P&#x000E1;k&#x000E1;ski</surname> <given-names>M.</given-names></name> <name><surname>K&#x000E1;lm&#x000E1;n</surname> <given-names>J.</given-names></name> <name><surname>Hoffmann</surname> <given-names>I.</given-names></name></person-group> (<year>2019</year>). <article-title>Identifying mild cognitive impairment and mild Alzheimer&#x00027;s disease based on spontaneous speech using ASR and linguistic features</article-title>. <source>Comput. Speech Lang</source>. <volume>53</volume>, <fpage>181</fpage>&#x02013;<lpage>197</lpage>. <pub-id pub-id-type="doi">10.1016/j.csl.2018.07.007</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jammeh</surname> <given-names>E. A.</given-names></name> <name><surname>Camille</surname> <given-names>B. C.</given-names></name> <name><surname>Stephen</surname> <given-names>W. P.</given-names></name> <name><surname>Escudero</surname> <given-names>J.</given-names></name> <name><surname>Anastasiou</surname> <given-names>A.</given-names></name> <name><surname>Zhao</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Machine-learning based identification of undiagnosed dementia in primary care: a feasibility study</article-title>. <source>BJGP Open</source> <volume>2</volume>:bjgpopen18X101589. <pub-id pub-id-type="doi">10.3399/bjgpopen18X101589</pub-id><pub-id pub-id-type="pmid">30564722</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jawahar</surname> <given-names>G.</given-names></name> <name><surname>Sagot</surname> <given-names>B.</given-names></name> <name><surname>Seddah</surname> <given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>What does bert learn about the structure of language?,</article-title> in <source>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</source> (<publisher-loc>Florence</publisher-loc>), <fpage>3651</fpage>&#x02013;<lpage>3657</lpage>. <pub-id pub-id-type="doi">10.18653/v1/P19-1356</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Johnson</surname> <given-names>A. E.</given-names></name> <name><surname>Pollard</surname> <given-names>T. J.</given-names></name> <name><surname>Naumann</surname> <given-names>T.</given-names></name></person-group> (<year>2018</year>). <article-title>Generalizability of predictive models for intensive care unit patients</article-title>. <source>arXiv preprint arXiv:1812.02275</source>.</citation></ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Karlekar</surname> <given-names>S.</given-names></name> <name><surname>Niu</surname> <given-names>T.</given-names></name> <name><surname>Bansal</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>Detecting linguistic characteristics of Alzheimer&#x00027;s dementia by interpreting neural models,</article-title> in <source>Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Volume 2 (Short Papers)</source> (<publisher-loc>New Orleans, LA</publisher-loc>), <fpage>701</fpage>&#x02013;<lpage>707</lpage>. <pub-id pub-id-type="doi">10.18653/v1/N18-2110</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D. P.</given-names></name> <name><surname>Ba</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>Adam: a method for stochastic optimization</article-title>. <source>arXiv preprint arXiv:1412.6980</source>.</citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>K&#x000F6;nig</surname> <given-names>A.</given-names></name> <name><surname>Satt</surname> <given-names>A.</given-names></name> <name><surname>Sorin</surname> <given-names>A.</given-names></name> <name><surname>Hoory</surname> <given-names>R.</given-names></name> <name><surname>Toledo-Ronen</surname> <given-names>O.</given-names></name> <name><surname>Derreumaux</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Automatic speech analysis for the assessment of patients with predementia and Alzheimer&#x00027;s disease</article-title>. <source>Alzheimer&#x00027;s Dement. Diagn. Assess. Dis. Monit</source>. <volume>1</volume>, <fpage>112</fpage>&#x02013;<lpage>124</lpage>. <pub-id pub-id-type="doi">10.1016/j.dadm.2014.11.012</pub-id><pub-id pub-id-type="pmid">27239498</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Le</surname> <given-names>X.</given-names></name> <name><surname>Lancashire</surname> <given-names>I.</given-names></name> <name><surname>Hirst</surname> <given-names>G.</given-names></name> <name><surname>Jokel</surname> <given-names>R.</given-names></name></person-group> (<year>2011</year>). <article-title>Longitudinal detection of dementia through lexical and syntactic changes in writing: a case study of three british novelists</article-title>. <source>Liter. Linguist. Comput</source>. <volume>26</volume>, <fpage>435</fpage>&#x02013;<lpage>461</lpage>. <pub-id pub-id-type="doi">10.1093/llc/fqr013</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Lapata</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <article-title>Text summarization with pretrained encoders,</article-title> in <source>Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source> (<publisher-loc>Hong Kong</publisher-loc>), <fpage>3721</fpage>&#x02013;<lpage>3731</lpage>. <pub-id pub-id-type="doi">10.18653/v1/D19-1387</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Luz</surname> <given-names>S.</given-names></name> <name><surname>Haider</surname> <given-names>F.</given-names></name> <name><surname>de la Fuente</surname> <given-names>S.</given-names></name> <name><surname>Fromm</surname> <given-names>D.</given-names></name> <name><surname>MacWhinney</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). <article-title>Alzheimer&#x00027;s dementia recognition through spontaneous speech: the address challenge</article-title>. <source>arXiv:2004.06833</source>. <pub-id pub-id-type="doi">10.21437/Interspeech.2020-2571</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maaten</surname> <given-names>L. v. d.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2008</year>). <article-title>Visualizing data using t-SNE</article-title>. <source>J. Mach. Learn. Res</source>. <volume>9</volume>, <fpage>2579</fpage>&#x02013;<lpage>2605</lpage>.</citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>MacWhinney</surname> <given-names>B.</given-names></name></person-group> (<year>2000</year>). <article-title>The CHILDES project: tools for analyzing talk: Volume I: Transcription format and programs, Volume II: the database</article-title>. <source>Comput. Linguist</source>. <volume>26</volume>:<fpage>657</fpage>. <pub-id pub-id-type="doi">10.1162/coli.2000.26.4.657</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mota</surname> <given-names>N. B.</given-names></name> <name><surname>Vasconcelos</surname> <given-names>N. A.</given-names></name> <name><surname>Lemos</surname> <given-names>N.</given-names></name> <name><surname>Pieretti</surname> <given-names>A. C.</given-names></name> <name><surname>Kinouchi</surname> <given-names>O.</given-names></name> <name><surname>Cecchi</surname> <given-names>G. A.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Speech graphs provide a quantitative measure of thought disorder in psychosis</article-title>. <source>PLoS ONE</source> <volume>7</volume>:<fpage>e34928</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0034928</pub-id><pub-id pub-id-type="pmid">22506057</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Noorian</surname> <given-names>Z.</given-names></name> <name><surname>Pou-Prom</surname> <given-names>C.</given-names></name> <name><surname>Rudzicz</surname> <given-names>F.</given-names></name></person-group> (<year>2017</year>). <article-title>On the importance of normative data in speech-based assessment</article-title>. <source>arXiv preprint arXiv:1712.00069</source>.</citation></ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Orimaye</surname> <given-names>S. O.</given-names></name> <name><surname>Tai</surname> <given-names>K. Y.</given-names></name> <name><surname>Wong</surname> <given-names>J. S.-M.</given-names></name> <name><surname>Wong</surname> <given-names>C. P.</given-names></name></person-group> (<year>2015</year>). <article-title>Learning linguistic biomarkers for predicting mild cognitive impairment using compound skip-grams</article-title>. <source>arXiv preprint arXiv:1511.02436</source>.</citation></ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Paszke</surname> <given-names>A.</given-names></name> <name><surname>Gross</surname> <given-names>S.</given-names></name> <name><surname>Massa</surname> <given-names>F.</given-names></name> <name><surname>Lerer</surname> <given-names>A.</given-names></name> <name><surname>Bradbury</surname> <given-names>J.</given-names></name> <name><surname>Chanan</surname> <given-names>G.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Pytorch: an imperative style, high-performance deep learning library,</article-title> in <source>Advances in Neural Information Processing Systems</source> (<publisher-loc>Vancouver, CA</publisher-loc>), <fpage>8024</fpage>&#x02013;<lpage>8035</lpage>.</citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Petti</surname> <given-names>U.</given-names></name> <name><surname>Baker</surname> <given-names>S.</given-names></name> <name><surname>Korhonen</surname> <given-names>A.</given-names></name></person-group> (<year>2020</year>). <article-title>A systematic literature review of automatic Alzheimer&#x00027;s disease detection from speech and language</article-title>. <source>J. Am. Med. Inform. Assoc</source>. <volume>27</volume>, <fpage>1784</fpage>&#x02013;<lpage>1797</lpage>. <pub-id pub-id-type="doi">10.1093/jamia/ocaa174</pub-id><pub-id pub-id-type="pmid">32929494</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pou-Prom</surname> <given-names>C.</given-names></name> <name><surname>Rudzicz</surname> <given-names>F.</given-names></name></person-group> (<year>2018</year>). <article-title>Learning multiview embeddings for assessing dementia,</article-title> in <source>Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing</source> (<publisher-loc>Brussels</publisher-loc>), <fpage>2812</fpage>&#x02013;<lpage>2817</lpage>. <pub-id pub-id-type="doi">10.18653/v1/D18-1304</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Prabhakaran</surname> <given-names>G.</given-names></name> <name><surname>Bakshi</surname> <given-names>R.</given-names></name></person-group> (<year>2018</year>). <article-title>Analysis of structure and cost in a longitudinal study of Alzheimer&#x00027;s disease</article-title>. <source>J. Health Care Fin</source>. <volume>8</volume>:<fpage>411</fpage>. <pub-id pub-id-type="doi">10.4172/2161-0460.1000411</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Prince</surname> <given-names>M.</given-names></name> <name><surname>Comas-Herrera</surname> <given-names>A.</given-names></name> <name><surname>Knapp</surname> <given-names>M.</given-names></name> <name><surname>Guerchet</surname> <given-names>M.</given-names></name> <name><surname>Karagiannidou</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <source>World Alzheimer Report 2016: Improving Healthcare for People Living With Dementia: Coverage, Quality and Costs Now and in the Future</source>. <publisher-name>Alzheimer&#x00027;s Disease International</publisher-name>.</citation></ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pulido</surname> <given-names>M. L. B.</given-names></name> <name><surname>Hern&#x000E1;ndez</surname> <given-names>J. B. A.</given-names></name> <name><surname>Ballester</surname> <given-names>M. &#x000C1;. F.</given-names></name> <name><surname>Gonz&#x000E1;lez</surname> <given-names>C. M. T.</given-names></name> <name><surname>Mekyska</surname> <given-names>J.</given-names></name> <name><surname>Sm&#x000E9;kal</surname> <given-names>Z.</given-names></name></person-group> (<year>2020</year>). <article-title>Alzheimer&#x00027;s disease and automatic speech analysis: a review</article-title>. <source>Expert Syst. Appl</source>. <volume>150</volume>:<fpage>113213</fpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2020.113213</pub-id><pub-id pub-id-type="pmid">33833713</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Riley</surname> <given-names>K. P.</given-names></name> <name><surname>Snowdon</surname> <given-names>D. A.</given-names></name> <name><surname>Desrosiers</surname> <given-names>M. F.</given-names></name> <name><surname>Markesbery</surname> <given-names>W. R.</given-names></name></person-group> (<year>2005</year>). <article-title>Early life linguistic ability, late life cognitive function, and neuropathology: findings from the nun study</article-title>. <source>Neurobiol. Aging</source> <volume>26</volume>, <fpage>341</fpage>&#x02013;<lpage>347</lpage>. <pub-id pub-id-type="doi">10.1016/j.neurobiolaging.2004.06.019</pub-id><pub-id pub-id-type="pmid">15639312</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rogers</surname> <given-names>A.</given-names></name> <name><surname>Kovaleva</surname> <given-names>O.</given-names></name> <name><surname>Rumshisky</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>A primer in bertology: what we know about how bert works</article-title>. <source>Trans. Assoc. Comput. Linguist</source>. <volume>8</volume>, <fpage>842</fpage>&#x02013;<lpage>866</lpage>. <pub-id pub-id-type="doi">10.1162/tacl_a_00349</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Slegers</surname> <given-names>A.</given-names></name> <name><surname>Filiou</surname> <given-names>R.-P.</given-names></name> <name><surname>Montembeault</surname> <given-names>M.</given-names></name> <name><surname>Brambati</surname> <given-names>S. M.</given-names></name></person-group> (<year>2018</year>). <article-title>Connected speech features from picture description in Alzheimer&#x00027;s disease: a systematic review</article-title>. <source>J. Alzheimer&#x00027;s Dis</source>. <volume>65</volume>, <fpage>519</fpage>&#x02013;<lpage>542</lpage>. <pub-id pub-id-type="doi">10.3233/JAD-170881</pub-id><pub-id pub-id-type="pmid">30103314</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Snowdon</surname> <given-names>D. A.</given-names></name> <name><surname>Kemper</surname> <given-names>S. J.</given-names></name> <name><surname>Mortimer</surname> <given-names>J. A.</given-names></name> <name><surname>Greiner</surname> <given-names>L. H.</given-names></name> <name><surname>Wekstein</surname> <given-names>D. R.</given-names></name> <name><surname>Markesbery</surname> <given-names>W. R.</given-names></name></person-group> (<year>1996</year>). <article-title>Linguistic ability in early life and cognitive function and Alzheimer&#x00027;s disease in late life: findings from the nun study</article-title>. <source>JAMA</source> <volume>275</volume>, <fpage>528</fpage>&#x02013;<lpage>532</lpage>. <pub-id pub-id-type="doi">10.1001/jama.1996.03530310034029</pub-id><pub-id pub-id-type="pmid">8606473</pub-id></citation></ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Szatloczki</surname> <given-names>G.</given-names></name> <name><surname>Hoffmann</surname> <given-names>I.</given-names></name> <name><surname>Vincze</surname> <given-names>V.</given-names></name> <name><surname>Kalman</surname> <given-names>J.</given-names></name> <name><surname>Pakaski</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>Speaking in Alzheimer&#x00027;s disease, is that an early sign? Importance of changes in language abilities in Alzheimer&#x00027;s disease</article-title>. <source>Front. Aging Neurosci</source>. <volume>7</volume>:<fpage>195</fpage>. <pub-id pub-id-type="doi">10.3389/fnagi.2015.00195</pub-id><pub-id pub-id-type="pmid">26539107</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Vig</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>A multiscale visualization of attention in the transformer model</article-title>. <source>arXiv preprint arXiv:1906.05714</source>. <pub-id pub-id-type="doi">10.18653/v1/P19-3007</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Warriner</surname> <given-names>A. B.</given-names></name> <name><surname>Kuperman</surname> <given-names>V.</given-names></name> <name><surname>Brysbaert</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>Norms of valence, arousal, and dominance for 13,915 English lemmas</article-title>. <source>Behav. Res. Methods</source> <volume>45</volume>, <fpage>1191</fpage>&#x02013;<lpage>1207</lpage>. <pub-id pub-id-type="doi">10.3758/s13428-012-0314-x</pub-id><pub-id pub-id-type="pmid">23404613</pub-id></citation></ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wolf</surname> <given-names>T.</given-names></name> <name><surname>Debut</surname> <given-names>L.</given-names></name> <name><surname>Sanh</surname> <given-names>V.</given-names></name> <name><surname>Chaumond</surname> <given-names>J.</given-names></name> <name><surname>Delangue</surname> <given-names>C.</given-names></name> <name><surname>Moi</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Huggingface&#x00027;s transformers: state-of-the-art natural language processing</article-title>. <source>ArXiv abs/1910.03771</source>. <pub-id pub-id-type="doi">10.18653/v1/2020.emnlp-demos.6</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yancheva</surname> <given-names>M.</given-names></name> <name><surname>Fraser</surname> <given-names>K. C.</given-names></name> <name><surname>Rudzicz</surname> <given-names>F.</given-names></name></person-group> (<year>2015</year>). <article-title>Using linguistic features longitudinally to predict clinical scores for Alzheimer&#x00027;s disease and related dementias,</article-title> in <source>Proceedings of SLPAT 2015: 6th Workshop on Speech and Language Processing for Assistive Technologies</source> (<publisher-loc>Dresden</publisher-loc>), <fpage>134</fpage>&#x02013;<lpage>139</lpage>. <pub-id pub-id-type="doi">10.18653/v1/W15-5123</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Z.</given-names></name> <name><surname>Dai</surname> <given-names>Z.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Carbonell</surname> <given-names>J.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R. R.</given-names></name> <name><surname>Le</surname> <given-names>Q. V.</given-names></name></person-group> (<year>2019</year>). <article-title>Xlnet: generalized autoregressive pretraining for language understanding,</article-title> in <source>Advances in Neural Information Processing Systems</source> (<publisher-loc>Vancouver, CA</publisher-loc>), <fpage>5753</fpage>&#x02013;<lpage>5763</lpage>.</citation></ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Young</surname> <given-names>T.</given-names></name> <name><surname>Hazarika</surname> <given-names>D.</given-names></name> <name><surname>Poria</surname> <given-names>S.</given-names></name> <name><surname>Cambria</surname> <given-names>E.</given-names></name></person-group> (<year>2018</year>). <article-title>Recent trends in deep learning based natural language processing</article-title>. <source>IEEE Comput. Intell. Mag</source>. <volume>13</volume>, <fpage>55</fpage>&#x02013;<lpage>75</lpage>. <pub-id pub-id-type="doi">10.1109/MCI.2018.2840738</pub-id></citation></ref>
<ref id="B51">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>M.</given-names></name> <name><surname>Gormley</surname> <given-names>M. R.</given-names></name> <name><surname>Dredze</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>Combining word embeddings and feature embeddings for fine-grained relation extraction,</article-title> in <source>Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics</source> (<publisher-loc>Denver, CO</publisher-loc>), <fpage>1374</fpage>&#x02013;<lpage>1379</lpage>. <pub-id pub-id-type="doi">10.3115/v1/N15-1155</pub-id></citation></ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yuan</surname> <given-names>J.</given-names></name> <name><surname>Bian</surname> <given-names>Y.</given-names></name> <name><surname>Cai</surname> <given-names>X.</given-names></name> <name><surname>Huang</surname> <given-names>J.</given-names></name> <name><surname>Ye</surname> <given-names>Z.</given-names></name> <name><surname>Church</surname> <given-names>K.</given-names></name></person-group> (<year>2020</year>). <article-title>Disfluencies and fine-tuning pre-trained language models for detection of Alzheimer&#x00027;s disease</article-title>. <source>Proc. Interspeech</source> <volume>2020</volume>, <fpage>2162</fpage>&#x02013;<lpage>2166</lpage>. <pub-id pub-id-type="doi">10.21437/Interspeech.2020-2516</pub-id></citation></ref>
<ref id="B53">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>Z.</given-names></name> <name><surname>Novikova</surname> <given-names>J.</given-names></name> <name><surname>Rudzicz</surname> <given-names>F.</given-names></name></person-group> (<year>2018</year>). <article-title>Semi-supervised classification by reaching consensus among modalities</article-title>. <source>arXiv preprint arXiv:1805.09366</source>.</citation></ref>
<ref id="B54">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>Z.</given-names></name> <name><surname>Novikova</surname> <given-names>J.</given-names></name> <name><surname>Rudzicz</surname> <given-names>F.</given-names></name></person-group> (<year>2019</year>). <article-title>Detecting cognitive impairments by agreeing on interpretations of linguistic features,</article-title> in <source>Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source> (<publisher-loc>Minneapolis, MN</publisher-loc>), <fpage>1431</fpage>&#x02013;<lpage>1441</lpage>. <pub-id pub-id-type="doi">10.18653/v1/N19-1146</pub-id></citation></ref>
</ref-list>

<fn-group>
<fn id="fn0001"><p><sup>1</sup><ext-link ext-link-type="uri" xlink:href="https://scikit-learn.org/stable/">https://scikit-learn.org/stable/</ext-link>.</p></fn>
</fn-group>
<fn-group>
<fn fn-type="financial-disclosure"><p><bold>Funding.</bold> FR was supported by a CIFAR Chair in AI.</p>
</fn>
</fn-group>
</back>
</article>