<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Big Data</journal-id>
<journal-title>Frontiers in Big Data</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Big Data</abbrev-journal-title>
<issn pub-type="epub">2624-909X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fdata.2022.863100</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Big Data</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>BurnoutEnsemble: Augmented Intelligence to Detect Indications for Burnout in Clinical Psychology</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Merhbene</surname> <given-names>Ghofrane</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Nath</surname> <given-names>Sukanya</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Puttick</surname> <given-names>Alexandre R.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1688523/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Kurpicz-Briki</surname> <given-names>Mascha</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/958744/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Applied Machine Intelligence Research Group, Department of Engineering and Information Technology, Bern University of Applied Sciences</institution>, <addr-line>Bern</addr-line>, <country>Switzerland</country></aff>
<aff id="aff2"><sup>2</sup><institution>Institute for Research in Open, Distance and eLearning (IFeL), Swiss Distance University of Applied Sciences</institution>, <addr-line>Brig</addr-line>, <country>Switzerland</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Kathiravan Srinivasan, Vellore Institute of Technology, India</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Juan Pedro Mart&#x000ED;nez-Ram&#x000F3;n, University of Murcia, Spain; Francisco Manuel Morales Rodr&#x000ED;guez, University of Granada, Spain</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Mascha Kurpicz-Briki <email>mascha.kurpicz&#x00040;bfh.ch</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Medicine and Public Health, a section of the journal Frontiers in Big Data</p></fn></author-notes>
<pub-date pub-type="epub">
<day>05</day>
<month>04</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>5</volume>
<elocation-id>863100</elocation-id>
<history>
<date date-type="received">
<day>26</day>
<month>01</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>25</day>
<month>02</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Merhbene, Nath, Puttick and Kurpicz-Briki.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Merhbene, Nath, Puttick and Kurpicz-Briki</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Burnout, a state of emotional, physical, and mental exhaustion caused by excessive and prolonged stress, is a growing concern. It is known to occur when an individual feels overwhelmed, emotionally exhausted, and unable to meet the constant demands imposed upon them. Detecting burnout is not an easy task, in large part because symptoms can overlap with those of other illnesses or syndromes. The use of natural language processing (NLP) methods has the potential to mitigate the limitations of typical burnout detection <italic>via</italic> inventories. In this article, the performance of NLP methods on anonymized free text data samples collected from the online forum/social media platform Reddit was analyzed. A dataset consisting of 13,568 samples describing first-hand experiences, of which 352 are related to burnout and 979 to depression, was compiled. This work demonstrates the effectiveness of NLP and machine learning methods in detecting indicators for burnout. Finally, it improves upon standard baseline classifiers by building and training an ensemble classifier using two methods (subreddit and random batching). The best ensemble models attain a balanced accuracy of 0.93, test F1 score of 0.43, and test recall of 0.93. Both the subreddit and random batching ensembles outperform the single classifier baselines in the experimental setup.</p></abstract>
<kwd-group>
<kwd>burnout</kwd>
<kwd>natural language processing</kwd>
<kwd>machine learning</kwd>
<kwd>augmented intelligence</kwd>
<kwd>ensemble classifier</kwd>
<kwd>psychology</kwd>
</kwd-group>
<counts>
<fig-count count="6"/>
<table-count count="9"/>
<equation-count count="0"/>
<ref-count count="42"/>
<page-count count="12"/>
<word-count count="8354"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Stress at the workplace is an increasingly relevant topic. In a study involving almost 10,000 working adults in eight territories throughout Europe, it was found that 18% of the respondents feel stressed daily, and three out of ten participants feel so stressed that they have considered finding a new job (ADP, <xref ref-type="bibr" rid="B1">2018</xref>). A Swiss study (SECO, <xref ref-type="bibr" rid="B36">2015</xref>) estimates that 24.2% of employees feel often or always stressed at their workplace, while 35.2% feel mostly (22.2%) or always (13%) exhausted at the end of the working day. In the latter group, 25.5% still feel exhausted the next morning, a circumstance which, if prolonged, can lead to various health hazards. Studies from the United States give the same indication. The Stress in America&#x00027;s Report of 2019 by the American Psychological Association shows that Americans consider a healthy stress level at an average of 3.8 (scale ranging from 1 to 10, where 10 is &#x0201C;a great deal of stress&#x0201D; and 1 is &#x0201C;little or no stress;&#x0201D;) however, they report having experienced an average stress level of 4.9 (American Psychological Association, <xref ref-type="bibr" rid="B2">2019</xref>).</p>
<p>This stress can lead to workplace burnout. In 2019, the WHO included burnout in the 11th Revision of the International Classification of Diseases (ICD-11) as a syndrome.<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref> In particular, during the pandemic crisis, burnout in the healthcare sector was an important issue: it has been shown, for instance, that the COVID-19 crisis has had an overwhelming psychological impact on intensive care workers (Azoulay et al., <xref ref-type="bibr" rid="B3">2020</xref>).</p>
<p>Identifying burnout syndrome is complex because symptoms can overlap with other diseases or syndromes (Jaggi, <xref ref-type="bibr" rid="B20">2019</xref>). In particular, the overlap between depression and burnout is an important subject of scientific discussion, e.g., (Schonfeld and Bianchi, <xref ref-type="bibr" rid="B34">2016</xref>). In clinical intervention and field research, burnout is typically detected <italic>via</italic> the use of <italic>inventories</italic>. Potential burnout patients fill out a psychological test, usually in the form of a questionnaire with scaled-response answers (e.g., not at all, sometimes, often, very often). Although such inventories are used in most studies and are well-established in the clinical environment, major limitations have been identified. For example, in personality inventories, participants are liable to fake their results, e.g., (Holden, <xref ref-type="bibr" rid="B19">2007</xref>). They may adapt their responses in high-stake situations in order to increase their chances for the desired outcome (Lambert, <xref ref-type="bibr" rid="B23">2013</xref>). A further issue with inventories is known as <italic>extreme response bias</italic> (ERB); some respondents will tend to choose (or avoid) only the highest or the lowest options in such tests (Greenleaf, <xref ref-type="bibr" rid="B16">1992</xref>; Brul&#x000E9; and Veenhoven, <xref ref-type="bibr" rid="B6">2017</xref>). It has also been shown that on self-reported tests for subjective well-being, the respondent&#x00027;s mood during testing sometimes contributes as a predictor (Diener et al., <xref ref-type="bibr" rid="B14">1991</xref>). Furthermore, defensiveness (the denial of symptoms) and social bias can influence the outcome of inventories (Williams et al., <xref ref-type="bibr" rid="B42">2019</xref>).</p>
<p>A potential way to mitigate the existing and well-known problems with inventories is to explore the use of free text questions or transcribed interviews. Previous studies have demonstrated promise in such methods (Burisch, <xref ref-type="bibr" rid="B7">2014</xref>), but, in practice, the manual effort of analyzing the resulting data often results in untenable overhead costs. Fortunately, recent developments in the field of natural language processing (NLP) make approaches using such unstructured textual data feasible. It has been shown that computational linguistic markers can be used to predict depressivity of the writer (Havigerov&#x000E1; et al., <xref ref-type="bibr" rid="B17">2019</xref>).</p>
<p>Existing work applying NLP to psychology focuses on the identification of indicators for different types of mental health disorders by using data obtained from social media, comprising the majority of available research in this area. For example, such work concentrates on suicide risk assessment (Morales et al., <xref ref-type="bibr" rid="B28">2019</xref>), (Just et al., <xref ref-type="bibr" rid="B22">2017</xref>), depression (Moreno et al., <xref ref-type="bibr" rid="B29">2011</xref>), (Schwartz et al., <xref ref-type="bibr" rid="B35">2014</xref>), post-partum depression (De Choudhury et al., <xref ref-type="bibr" rid="B11">2013</xref>), (De Choudhury et al., <xref ref-type="bibr" rid="B12">2014</xref>), or different mental health signals (Coppersmith et al., <xref ref-type="bibr" rid="B10">2014</xref>). In some cases, data from Reddit online forums have been used, for example, to detect mental health disorders (Thorstad and Wolff, <xref ref-type="bibr" rid="B40">2019</xref>), anxiety (Shen and Rudzicz, <xref ref-type="bibr" rid="B37">2017</xref>), or depression (Tadesse et al., <xref ref-type="bibr" rid="B39">2019</xref>).</p>
<p>However, very little work exists in the field of burnout detection. Burnout detection in data extracted from issues and comments posted within software development tools have been studied (M&#x000E4;ntyl&#x000E4; et al., <xref ref-type="bibr" rid="B26">2016</xref>). The authors used the valence-arousal-dominance (VAD) model to study burnout risk in a corporate setting. This model distinguishes three emotions: <italic>valence</italic> (&#x0201C;the pleasantness of a stimulus,&#x0201D;) <italic>arousal</italic> (&#x0201C;the intensity of emotion provoked by a stimulus,&#x0201D;) and <italic>dominance</italic> (&#x0201C;the degree of control exerted by a stimulus&#x0201D;) (Warriner et al., <xref ref-type="bibr" rid="B41">2013</xref>). To measure burnout risk, the metric is based on low valence and dominance and high arousal (M&#x000E4;ntyl&#x000E4; et al., <xref ref-type="bibr" rid="B26">2016</xref>). In other work, a first attempt to detect burnout based on patient and expert interviews in the German language were done; it was found that a combination of NLP and machine learning techniques in this field leads to promising results (Nath and Kurpicz-Briki, <xref ref-type="bibr" rid="B30">2021</xref>).</p>
<p>In the context of earlier work focused on gathering data from social media websites and the study of mental health conditions, this work extends state-of-the-art predictive models in the field while focusing specifically on detecting indicators for burnout in data collected from Reddit. It aims to develop the base technology for potential new directions in tool development for clinical psychology. Herein, the authors emphasize that this work is oriented toward the approach of augmented intelligence rather than artificial intelligence (Rui, <xref ref-type="bibr" rid="B32">2017</xref>); instead of replacing clinical professionals, it strives toward technology that empowers humans in the decision-making process, providing input to be considered in human decision-making.</p>
<p>The work in this article addresses the following objectives:</p>
<list list-type="bullet">
<list-item><p>It evaluates whether NLP methods applied to free text are an effective means to detect indicators for burnout, compared to a control group using general text samples, and a control group with depression-related texts.</p></list-item>
<list-item><p>In particular, it investigates how the use of an ensemble classifier can leverage the accuracy of such methods.</p></list-item>
<list-item><p>Furthermore, the approach is compared to single machine learning classifiers such as logistic regression.</p></list-item>
</list>
<p>This article is structured as follows: first, the materials and methods used in this work are discussed. In particular, this includes data collection, the characteristics of the datasets used in the experiments, and the experimental setup. Then, the results are presented, first for single classifier models and then for the ensemble models. Finally, the results are discussed and an outlook on potential future work is provided.</p>
</sec>
<sec sec-type="materials and methods" id="s2">
<title>2. Materials and Methods</title>
<sec>
<title>2.1. Reddit Data Collection</title>
<p>On Reddit, users can organize posts based on a subject, so-called <italic>subreddits</italic>, which are online micro communities dedicated to a particular topic. Reddit has the advantage of allowing the possibility to create micro communities <italic>via</italic> subreddits. As a result, in addition to topics such as gaming and music, there are thriving communities dedicated to various mental health topics, such as depression, anxiety, and bipolar disorder. In particular, there is a subreddit dedicated to burnout; unfortunately, the number of entries was too low at the time of our data collection to provide a sufficiently large dataset. However, users discuss the subject of burnout in various other subreddit threads. One can thus collect textual data related to burnout by scraping Reddit for burnout-related posts. In this work, <monospace>praw</monospace> (Boe, <xref ref-type="bibr" rid="B5">2011</xref>), a Python Reddit API Wrapper, was used to extract submissions with the keyword &#x0201C;burnout&#x0201D; and its different variations, such as &#x0201C;burnout,&#x0201D; &#x0201C;burn out,&#x0201D; &#x0201C;burned out,&#x0201D; &#x0201C;burning out,&#x0201D; &#x0201C;burnt out,&#x0201D; &#x0201C;burn-out,&#x0201D; etc., 1,536 such submissions were found.</p>
<p>However, the word <italic>burnout</italic> also widely occurs in other contexts, such as &#x0201C;The tires are burnt out.&#x0201D; It is also frequently used in informal discussions, such as having <italic>game burnout</italic> or <italic>music burnout</italic>. It was therefore necessary to isolate submissions describing burnout in the professional or educational context. A total of 677 submissions satisfying these conditions were manually identified. The replies to the selected submissions were also collected, as they were likely to contain posts by other users describing their experiences with burnout. This increased the size of the dataset to 23,371 posts. However, not all of the posts and replies were relevant to professional or educational experiences with burnout. Therefore, 352 instances were extracted manually that describe burnout experiences from a first-person perspective. This formed the test group for the data classified as <italic>burnout</italic>.</p>
<p>To create the first control group, the <italic>no burnout</italic> dataset, our method employed the strategy described in Shen and Rudzicz (<xref ref-type="bibr" rid="B37">2017</xref>). Namely, 17,025 posts from a variety of subreddits were collected: &#x0201C;askscience&#x0201D;, &#x0201C;relationships&#x0201D;, &#x0201C;writingprompts&#x0201D;, &#x0201C;teaching&#x0201D;, &#x0201C;writing&#x0201D;, &#x0201C;parenting&#x0201D;, &#x0201C;atheism&#x0201D;, &#x0201C;christianity&#x0201D;, &#x0201C;showerthoughts&#x0201D;, &#x0201C;jokes&#x0201D;, &#x0201C;lifeprotips&#x0201D;, &#x0201C;writing&#x0201D;, &#x0201C;personalfinance&#x0201D;, &#x0201C;talesfromretail&#x0201D;, &#x0201C;theoryofreddit&#x0201D;, &#x0201C;talesfromtechsupport&#x0201D;, &#x0201C;randomkindness&#x0201D;, &#x0201C;talesfromcallcenters&#x0201D;, &#x0201C;books&#x0201D;, &#x0201C;fitness&#x0201D;, &#x0201C;askdocs&#x0201D;, &#x0201C;frugal&#x0201D;, &#x0201C;legaladvice&#x0201D;, &#x0201C;youshouldknow&#x0201D;, and &#x0201C;nostupidquestions&#x0201D;, Since a number of these collected posts consisted of empty or very little text, all posts consisting of fewer than 100 characters were dropped, resulting in a final <italic>no burnout</italic> dataset consisting of 13,216 posts.</p>
<p>The second control group, the <italic>depression</italic> dataset, was collected from the subreddit for depression and contains 979 posts. As for burnout, only entries using the first-person perspective were selected.</p>
<p>The authors emphasize that no information concerning user identity (e.g., username or age) was collected.</p>
</sec>
<sec>
<title>2.2. Datasets for Experiments</title>
<p>Using the raw data consisting of 13,216 posts labeled <italic>no burnout</italic> (control group), 352 labeled <italic>burnout</italic>, and 979 labeled <italic>depression</italic>, four datasets for use in the experiments were compiled. Dataset statistics are presented in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<list list-type="simple">
<list-item><p><bold>Dataset 1: Burnout vs. No Burnout (BNB):</bold> It combines the 13,216 <italic>no burnout</italic> posts with the 352 <italic>burnout</italic> posts, resulting in a highly unbalanced dataset of size 13,568.</p></list-item>
<list-item><p><bold>Dataset 2: Burnout vs. No Burnout (Balanced) (BNB-Balanced):</bold> Balanced dataset of 704 posts, of which, 352 posts are selected from the <italic>no burnout</italic> dataset through random sampling (without replacement). Additionally, an equal number of 352 posts are added from the <italic>burnout</italic> data.</p></list-item>
<list-item><p><bold>Dataset 3: Burnout vs. No Burnout (No Keywords) (BNB-No-Keywords):</bold> It is obtained from Dataset 2 by removing the keywords from the <italic>burnout</italic> dataset that were used during data collection to search for burnout-related posts: &#x0201C;burnout,&#x0201D; &#x0201C;burn-out,&#x0201D; &#x0201C;burning out,&#x0201D; etc.</p></list-item>
<list-item><p><bold>Dataset 4: Burnout vs. Depression (BD):</bold> Balanced dataset of 704 entries, of which 352 posts are selected from the <italic>depression</italic> dataset through random sampling (without replacement). Further, an equal number of 352 posts are added from the <italic>burnout</italic> data.</p></list-item>
</list>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Dataset statistics.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Dataset Name</bold></th>
<th valign="top" align="center"><bold>No. of Samples</bold></th>
<th valign="top" align="center"><bold>Mean Text Length (chars)</bold></th>
<th valign="top" align="center"><bold>Std. Dev of Text Length</bold></th>
<th valign="top" align="center"><bold>Test Group %age</bold></th>
<th valign="top" align="center"><bold>Control Group %age</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">1. Burnout vs. No Burnout(BNB)</td>
<td valign="top" align="center">13,568</td>
<td valign="top" align="center">1158</td>
<td valign="top" align="center">1451</td>
<td valign="top" align="center">2.6%</td>
<td valign="top" align="center">97.4%</td>
</tr>
<tr>
<td valign="top" align="left">2. Burnout vs. No Burnout</td>
</tr>
<tr>
<td valign="top" align="left">(Bal.) (BNB-balanced)</td>
<td valign="top" align="center">704</td>
<td valign="top" align="center">867</td>
<td valign="top" align="center">850</td>
<td valign="top" align="center">50%</td>
<td valign="top" align="center">50%</td>
</tr>
<tr>
<td valign="top" align="left">3. Burnout vs. No Burnout</td>
</tr>
<tr>
<td valign="top" align="left">(No KWs)(BNB-no-keywords)</td>
<td valign="top" align="center">704</td>
<td valign="top" align="center">863</td>
<td valign="top" align="center">846</td>
<td valign="top" align="center">50%</td>
<td valign="top" align="center">50%</td>
</tr>
<tr>
<td valign="top" align="left">4. Burnout vs. Depression (BD)</td>
<td valign="top" align="center">704</td>
<td valign="top" align="center">1009</td>
<td valign="top" align="center">905</td>
<td valign="top" align="center">50%</td>
<td valign="top" align="center">50%</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>&#x0201C;Control&#x0201D; refers to either no burnout or depression, while test refers to burnout</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>2.3. Vectorization</title>
<p>The <monospace>spacy</monospace><xref ref-type="fn" rid="fn0002"><sup>2</sup></xref> Python NLP-library was used in order to vectorize text data for use in our NLP models. Each Reddit post was tokenized using the pre-trained <monospace>en_core_web_sm</monospace> English language pipeline and converted into a 500-dimensional bag-of-words vector, which simply counts the occurrences of each of the 500 most commonly appearing words in the text corpus.</p>
</sec>
<sec>
<title>2.4. Experimental Setup</title>
<sec>
<title>2.4.1. Single Classifier Models</title>
<p>The following experiment was repeated on Datasets 1&#x02013;4. The feature set consisted of the vectorized Reddit posts, each labeled with either 1 (burnout) or 0 (no burnout/depression). Using a 70-30% training-test split<xref ref-type="fn" rid="fn0003"><sup>3</sup></xref> and 10-fold cross-validation (CV), a variety of classifier models was trained: logistic regression, Support Vector Machine (SVM) (with linear, RBF, degree 3 polynomial and sigmoid kernels), and random forest. Each model&#x00027;s performance was measured by using the mean CV accuracy and F1 scores averaged across all folds, as well as the (balanced) accuracy, F1, and recall scores on the test data. It was chosen to specifically include recall as a metric because, in a real-world setting, it would be important to capture all possible <italic>burnout</italic> samples (recall &#x0003D; 1), even at the expense of a larger number of false positives (see Section 4.5 for further discussion).</p>
</sec>
<sec>
<title>2.4.2. Ensemble Classifier Models</title>
<p>Ensemble classifiers allow aggregating the decisions of several single classifier models. The ensemble methods presented in this work closely resemble a method known as UnderBagging (Barandela et al., <xref ref-type="bibr" rid="B4">2003</xref>). Each ensemble is built according to the template below.</p>
<p><italic>Ensemble model template:</italic></p>
<list list-type="bullet">
<list-item><p>The ensemble consists of <italic>n</italic> submodels.<xref ref-type="fn" rid="fn0004"><sup>4</sup></xref></p></list-item>
<list-item><p>Each submodel is trained with 10-folds CV on a balanced dataset of 492 posts.</p></list-item>
<list-item><p>These datasets share the same 246 <italic>burnout</italic> samples but contain pairwise disjoint sets of <italic>no burnout</italic> samples.</p></list-item>
<list-item><p>The prediction of the whole ensemble is determined by voting, i.e., for a given test sample, a label of <italic>burnout</italic> is predicted if the voting &#x0003E;<italic>p</italic>% of the submodels classify the sample as <italic>burnout</italic>.<xref ref-type="fn" rid="fn0005"><sup>5</sup></xref> Otherwise, the sample is classified as <italic>no burnout</italic>.</p></list-item>
</list>
<p>The classifier type of the submodels was restricted to logistic regression, which demonstrated the most consistent performance in our initial experiments, although RBF, linear SVMs, and random forests also showed promise. <xref ref-type="fig" rid="F1">Figure 1</xref> shows a depiction of our ensemble setup, along with a comparison to the two baseline models to which the ensemble results were compared:</p>
<list list-type="bullet">
<list-item><p><bold>Baseline 1:</bold> Logistic regression classifier trained <italic>via</italic> 70&#x02013;30% train-test split on the unbalanced Dataset 1 (BNB).</p></list-item>
<list-item><p><bold>Baseline 2:</bold> Logistic regression classifier trained on a balanced dataset obtained by randomly sampling 246 <italic>no burnout</italic> samples and combining them with the 246 <italic>burnout</italic> samples used for training.</p></list-item>
</list>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Training the baselines vs. training ensembles on balanced batches.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-05-863100-g0001.tif"/>
</fig>
<p><xref ref-type="fig" rid="F1">Figure 1</xref> depicts the setup for training the baseline classifier (logistic regression on the full unbalanced training set) and the ensemble classifiers, and <xref ref-type="fig" rid="F2">Figure 2</xref> depicts how each model makes predictions on the test data.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Computing predictions on the test set, with a voting threshold of <italic>p</italic> = 0.8 for example purposes.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-05-863100-g0002.tif"/>
</fig>
<p>Training and test data were allocated according to a 70-30% split. This was done in a stratified manner, i.e., the <italic>no burnout</italic> and <italic>burnout</italic> class distribution in the training and test data were approximately equal to the distribution in Dataset 1 (BNB) (as shown in <xref ref-type="fig" rid="F3">Figure 3</xref>). Note that the same test data were used for both baselines and ensembles.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Constructing the unbalanced test set of 4,071 samples.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-05-863100-g0003.tif"/>
</fig>
<p>Two types of data batching were tested in our ensembles (Ensemble 1 and Ensemble 2, see description below) and the following metrics were measured:</p>
<list list-type="bullet">
<list-item><p><italic>Mean CV accuracy:</italic><xref ref-type="fn" rid="fn0006"><sup>6</sup></xref> Computed by first taking the mean CV accuracies for each submodel over the 10 folds, followed by averaging over the <italic>n</italic> submodels.</p></list-item>
<list-item><p><italic>Mean CV F1 (macro):</italic> Identical with F1 in place of accuracy.</p></list-item>
<list-item><p><italic>Mean test balanced accuracy</italic>: The balanced accuracy on test data averaged across the <italic>n</italic> submodels.</p></list-item>
<list-item><p><italic>Mean test F1 (macro)</italic>: Identical for F1.</p></list-item>
<list-item><p><italic>Mean test recall</italic>: Identical for recall.</p></list-item>
<list-item><p>The corresponding SDs of the above three test metrics.</p></list-item>
</list>
<p><bold>Ensemble 1: Random sample batching:</bold></p>
<p>The <italic>random sample batching</italic> ensemble was trained using <italic>n</italic> &#x0003D; 20 batches, each consisting of 246 randomly sampled (without replacement) posts from the <italic>no burnout</italic> training samples concatenated with 246 burnout training samples to create BNB-balanced datasets.</p>
<p><bold>Ensemble 2: Batching by subreddit:</bold></p>
<p>The <italic>subreddit batching</italic> ensemble was trained by creating a balanced dataset corresponding to each of the subreddits appearing in the <italic>no burnout</italic> training data for which at least 246 samples had been collected. There were <italic>n</italic> &#x0003D; 17 such subreddits in total.</p>
<p>The effect of changing the voting threshold <italic>p</italic> on the ensemble performance was also tested. Values of <italic>p</italic> &#x0003D; 0.4, 0.5, 0.6, 0.7, 0.8, and 0.9 were evaluated.</p>
</sec>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>3. Results</title>
<sec>
<title>3.1. Single Classifier Models</title>
<sec>
<title>3.1.1. Burnout vs. No Burnout</title>
<p>The results of the single classifier experiments on Dataset 1 (Burnout vs. No Burnout BNB) are displayed in <xref ref-type="table" rid="T2">Table 2</xref>. The <italic>Baseline</italic> row corresponds to a model that predicts the label <italic>no burnout</italic> for all samples. Such a model achieves 97% accuracy due to the class imbalance in Dataset 1 (BNB). Indeed, accuracy is a misleading measure in such a situation: all classifiers in this experiment demonstrated an accuracy of approximately 97% despite large differences in performance. For this reason, balanced accuracy provides a more meaningful metric for model performance.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Results for Burnout vs. No Burnout &#x02013; Dataset 1 (BNB) (no. test samples = 4071).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="left"><bold>Mean CV Bal. Acc</bold>.</th>
<th valign="top" align="left"><bold>Mean CV F1</bold></th>
<th valign="top" align="left"><bold>Test Bal. Acc</bold>.</th>
<th valign="top" align="left"><bold>Test F1</bold></th>
<th valign="top" align="left"><bold>Test Recall</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Logistic Regression</td>
<td valign="top" align="left">0.72</td>
<td valign="top" align="left">0.48</td>
<td valign="top" align="left">0.75</td>
<td valign="top" align="left">0.49</td>
<td valign="top" align="left">0.50</td>
</tr>
<tr>
<td valign="top" align="left">SVM Linear</td>
<td valign="top" align="left">0.72</td>
<td valign="top" align="left">0.40</td>
<td valign="top" align="left">0.75</td>
<td valign="top" align="left">0.45</td>
<td valign="top" align="left">0.51</td>
</tr>
<tr>
<td valign="top" align="left">SVM RBF</td>
<td valign="top" align="left">0.51</td>
<td valign="top" align="left">0.04</td>
<td valign="top" align="left">0.51</td>
<td valign="top" align="left">0.03</td>
<td valign="top" align="left">0.01</td>
</tr>
<tr>
<td valign="top" align="left">SVM Poly Degree 3</td>
<td valign="top" align="left">0.55</td>
<td valign="top" align="left">0.16</td>
<td valign="top" align="left">0.56</td>
<td valign="top" align="left">0.18</td>
<td valign="top" align="left">0.12</td>
</tr>
<tr>
<td valign="top" align="left">SVM Sigmoid</td>
<td valign="top" align="left">0.57</td>
<td valign="top" align="left">0.23</td>
<td valign="top" align="left">0.56</td>
<td valign="top" align="left">0.21</td>
<td valign="top" align="left">0.12</td>
</tr>
<tr>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="left">0.50</td>
<td valign="top" align="left">0.02</td>
<td valign="top" align="left">0.51</td>
<td valign="top" align="left">0.04</td>
<td valign="top" align="left">0.02</td>
</tr>
<tr>
<td valign="top" align="left">Baseline</td>
<td valign="top" align="left">0.50</td>
<td valign="top" align="left">0.0</td>
<td valign="top" align="left">0.50</td>
<td valign="top" align="left">0.0</td>
<td valign="top" align="left">0.0</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Baseline refers to a model predicting no burnout for all samples. The mean CV statistics are computed by taking an average of overall 10 folds cross-validation (CV) during training</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>Only logistic regression and SVM linear demonstrated significant improvement over the baseline, although roughly 50% of burnout samples were incorrectly classified as <italic>no burnout</italic>.</p>
</sec>
<sec>
<title>3.1.2. Burnout vs. No Burnout (Balanced)</title>
<p>Here, the results of classifiers trained using Dataset 2 (BNB-balanced) are presented. Aside from the SVM poly degree 3 classifier, the models in <xref ref-type="table" rid="T3">Table 3</xref> appear to demonstrate good performance.<xref ref-type="fn" rid="fn0007"><sup>7</sup></xref> It was noted that these results are dependent on the random sample of <italic>no burnout</italic> data points that are used to construct Dataset 2 (BNB-balanced). While random forest classifiers demonstrated the best performance in this instance, there were also cases in which logistic regression performed best. For the best models, accuracies and F1 scores approximately distributed between 0.90 and 0.97 were observed.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Results for Burnout vs. No Burnout (Balanced)&#x02014;Dataset 2 (BNB-balanced) (no. test samples = 234).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>Mean CV Accuracy</bold></th>
<th valign="top" align="center"><bold>Mean CV F1</bold></th>
<th valign="top" align="center"><bold>Test Accuracy</bold></th>
<th valign="top" align="center"><bold>Test F1</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Logistic regression</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">0.88</td>
</tr>
<tr>
<td valign="top" align="left">SVM Linear</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.85</td>
</tr>
<tr>
<td valign="top" align="left">SVM RBF</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">0.89</td>
</tr>
<tr>
<td valign="top" align="left">SVM Poly degree 3</td>
<td valign="top" align="center">0.60</td>
<td valign="top" align="center">0.35</td>
<td valign="top" align="center">0.60</td>
<td valign="top" align="center">0.41</td>
</tr>
<tr>
<td valign="top" align="left">SVM Sigmoid</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.82</td>
</tr>
<tr>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.89</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The mean CV statistics are computed by taking an average of overall 10 folds CV during training</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>3.1.3. Burnout vs. No Burnout (No Keywords) (BNB-No-Keywords)</title>
<p>The data collection process applied in this work explicitly searches for burnout-related keywords. It is, therefore, possible that trained models identify the presence of such keywords as a key defining feature for posts belonging to the <italic>burnout</italic> class. The effect of the presence of such keywords was measured, and it was tested whether they provided a significant basis for the models&#x00027; predictions. Therefore, all keywords related to burnout were removed from Dataset 2 (BNB-balanced) to obtain Dataset 3 (BNB-no-keywords) and the experiment was repeated. The corresponding results are displayed in <xref ref-type="table" rid="T4">Table 4</xref>. As one might expect, the removal of keywords resulted in decreased model performance. However, the decrease was not very important, providing evidence that the presence of keywords is not an overly important factor in any of our other experiments.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Results for Burnout vs. No Burnout (no keywords)&#x02014;Dataset 3 (BNB-no-key-words) (no. test samples = 234).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>Mean CV Accuracy</bold></th>
<th valign="top" align="center"><bold>Mean CV F1</bold></th>
<th valign="top" align="center"><bold>Test Accuracy</bold></th>
<th valign="top" align="center"><bold>Test F1</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Logistic regression</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.87</td>
</tr>
<tr>
<td valign="top" align="left">SVM Linear</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.83</td>
</tr>
<tr>
<td valign="top" align="left">SVM RBF</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.83</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.86</td>
</tr>
<tr>
<td valign="top" align="left">SVM Poly degree 3</td>
<td valign="top" align="center">0.59</td>
<td valign="top" align="center">0.34</td>
<td valign="top" align="center">0.59</td>
<td valign="top" align="center">0.40</td>
</tr>
<tr>
<td valign="top" align="left">SVM Sigmoid</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.82</td>
</tr>
<tr>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.88</td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">0.88</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The mean CV statistics are computed by taking an average of overall 10 folds CV during training</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>3.1.4. Burnout vs. Depression (BD)</title>
<p>In this experiment, as shown in <xref ref-type="table" rid="T5">Table 5</xref>, the Burnout vs. Depression dataset (Dataset 4, BD) was classified by using the models described previously. Again, it was found that logistic regression and SVM linear performed best, with random forest following closely. The datasets are balanced, and the random baseline for both accuracy and F1 score is set at 50%.</p>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Results for Burnout vs. Depression&#x02014;Dataset 4 (BD).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>Mean CV Accuracy</bold></th>
<th valign="top" align="center"><bold>Mean CV F1</bold></th>
<th valign="top" align="center"><bold>Test Accuracy</bold></th>
<th valign="top" align="center"><bold>Test F1</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Logistic regression</td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">0.87</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.82</td>
</tr>
<tr>
<td valign="top" align="left">SVM Linear</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.78</td>
</tr>
<tr>
<td valign="top" align="left">SVM RBF</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.77</td>
</tr>
<tr>
<td valign="top" align="left">SVM Poly degree 3</td>
<td valign="top" align="center">0.59</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.65</td>
<td valign="top" align="center">0.43</td>
</tr>
<tr>
<td valign="top" align="left">SVM Sigmoid</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.76</td>
</tr>
<tr>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.80</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The mean CV statistics are computed by taking an average of overall 10 folds CV during training</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>Although these models perform well, an across-the-board decrease of roughly 0.04 points is observed compared to the results listed in <xref ref-type="table" rid="T3">Table 3</xref>.</p>
</sec>
</sec>
<sec>
<title>3.2. Ensemble Models</title>
<p><xref ref-type="table" rid="T6">Tables 6</xref>, <xref ref-type="table" rid="T7">7</xref> record metrics and statistics that pertain exclusively to the submodels and not to the overall ensembles. They are meant as a means to compare the performance of the individual submodels to that of the ensembles (<xref ref-type="table" rid="T8">Table 8</xref>).</p>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>Submodel CV averages.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>Mean CV Accuracy</bold></th>
<th valign="top" align="center"><bold>Mean CV F1</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Random Batching</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.90</td>
</tr>
<tr>
<td valign="top" align="left">Subreddit Batching</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.96</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T7">
<label>Table 7</label>
<caption><p>Submodel test statistics (no. test samples = 4,071).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>Mean Test Bal. Accuracy</bold></th>
<th valign="top" align="center"><bold>Std. Dev. Test Bal. Acc</bold>.</th>
<th valign="top" align="center"><bold>Mean Test F1</bold></th>
<th valign="top" align="center"><bold>Std. Dev. Test F1</bold></th>
<th valign="top" align="center"><bold>Mean Test Recall</bold></th>
<th valign="top" align="center"><bold>Std. Dev. Test Recall</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Random Batching</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">0.35</td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.02</td>
</tr>
<tr>
<td valign="top" align="left">Subreddit Batching</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.08</td>
<td valign="top" align="center">0.13</td>
<td valign="top" align="center">0.04</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.02</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap position="float" id="T8">
<label>Table 8</label>
<caption><p>Ensemble vs. baseline performance (Threshold <italic>p</italic> &#x0003D; 80%, no. test samples = 4,071).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>Test Bal. Acc</bold>.</th>
<th valign="top" align="center"><bold>Test F1</bold></th>
<th valign="top" align="center"><bold>Test Recall</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Random Batching Ensemble</td>
<td valign="top" align="center">0.91</td>
<td valign="top" align="center">0.56</td>
<td valign="top" align="center">0.84</td>
</tr>
<tr>
<td valign="top" align="left">Subreddit Batching Ensemble</td>
<td valign="top" align="center">0.93</td>
<td valign="top" align="center">0.34</td>
<td valign="top" align="center">0.95</td>
</tr>
<tr>
<td valign="top" align="left">Baseline 1: Unbalanced LR</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.49</td>
<td valign="top" align="center">0.50</td>
</tr>
<tr>
<td valign="top" align="left">Baseline 2: Random Undersampling LR</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">0.91</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="T6">Table 6</xref> records the average CV performance metrics over the submodels within each of the ensemble classifiers. Recall that each of these submodels is a logistic regression classifier trained on a balanced dataset. The <italic>Mean CV Accuracy</italic> and <italic>Mean CV F1</italic> columns in <xref ref-type="table" rid="T6">Table 6</xref> are thus comparable to the corresponding columns in <xref ref-type="table" rid="T3">Table 3</xref>.</p>
<p><xref ref-type="table" rid="T7">Table 7</xref> records the average test statistics for the submodels. The <italic>Mean test Bal. Acc</italic>. and <italic>Mean test F1</italic> columns refer to the average performance of the submodels on the unbalanced test set consisting of 4,071 samples, of which 106 belong to the <italic>burnout</italic> class. It also provides the corresponding SDs. The random batching submodels were much more consistent than the subreddit batching submodels, the latter of which demonstrated greater variance and lower average balanced accuracy and F1 score while achieving higher recall scores.</p>
<p>The test results reveal the limitations of the previously presented non-ensemble models trained on balanced data. Those models appeared to demonstrate very good performance on unseen test data (<xref ref-type="table" rid="T3">Table 3</xref>), but were tested on small balanced test sets consisting of only 234 posts. In comparison, the mean test F1 scores in <xref ref-type="table" rid="T7">Table 7</xref> are relatively low, which shows that the high test performance observed in <xref ref-type="table" rid="T3">Table 3</xref> does not imply similar performance on the unbalanced dataset of 4,071 samples. The high test recall in <xref ref-type="table" rid="T7">Table 7</xref> indicates that the test F1 scores are primarily reduced due to low precision, i.e., a relatively large number of false positives.</p>
<p>The mean test metrics in <xref ref-type="table" rid="T7">Table 7</xref> corresponding to random batching give an indication of how the logistic regression model trained on Dataset 2 (BNB-balanced) (<xref ref-type="table" rid="T3">Table 3</xref>) would perform on the large unbalanced test dataset used in our ensemble experiments. Note that the performance of the balanced data model depends on the random sample of 352 <italic>no burnout</italic> posts used to construct Dataset 2 (BNB-balanced), and significant fluctuations in performance were observed depending on the sample, encapsulated in the SDs recorded in <xref ref-type="table" rid="T7">Table 7</xref>. Indeed, the pursuit of ensemble approaches presented in this work was driven partially by the desire for a model with more stable performance. Effectively, <xref ref-type="table" rid="T7">Table 7</xref> portrays the average performance of logistic regression models trained on balanced No Burnout vs. Burnout data over <italic>n</italic> &#x0003D; 20 disjoint random samples of <italic>no burnout</italic> data. It was observed that the submodels trained <italic>via</italic> subreddit batching demonstrated lower performance on the test data than those trained <italic>via</italic> random batching.</p>
<p><xref ref-type="table" rid="T8">Table 8</xref> shows the test results of the two ensemble models. The logistic regression (LR) model trained on Dataset 1 (BNB) was used as a first baseline, which demonstrated the best overall performance on the unbalanced test data among the single-model classifiers. As a second baseline, a single logistic regression classifier trained on a balanced dataset obtained by randomly undersampling from the <italic>no burnout</italic> class was considered, as was done to construct Dataset 2 (BNB-balanced). The model demonstrated performance similar to the averages recorded in <xref ref-type="table" rid="T7">Table 7</xref>.</p>
<p>The ensemble models demonstrated substantially improved balanced accuracy and recall relative to the baseline unbalanced LR model. However, the unbalanced LR model achieved the second-highest F1 score. Both ensembles and the baseline random undersampling LR demonstrated similar performance, with the random batching ensemble exhibiting a trade-off between F1 and recall. In comparing <xref ref-type="table" rid="T7">Tables 7</xref>, <xref ref-type="table" rid="T8">8</xref>, one can see that the random batching ensemble demonstrates a relatively modest performance improvement over the submodels composing it. On the other hand, the subreddit batching ensemble performs markedly better than its component submodels.</p>
<p>The confusion matrices in <xref ref-type="fig" rid="F4">Figures 4</xref>, <xref ref-type="fig" rid="F5">5</xref> describe the distribution of the ensemble models&#x00027; test predictions. Both the submodels and ensembles had test recall scores near 1. However, the high recall of the submodels came at the cost of a large number of false positives. It was observed that each submodel identified approximately 400&#x02013;500 (random batching) to 1,000&#x02013;3,000 (subreddit batching) test samples as belonging to the <italic>burnout</italic> class, whereas the correct number was 106. In contrast, with a majority vote threshold of <italic>p</italic> &#x0003D; 80%, the ensemble models placed 216 (random batching). There were 486 (subreddit batching) test samples in the burnout class while maintaining a recall score close to 1. The majority vote ensemble rule is thus an effective method for eliminating false positives while preserving true positives.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Confusion matrix for random batching ensemble, <italic>p</italic> &#x0003D; 0.8.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-05-863100-g0004.tif"/>
</fig>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Confusion matrix for subreddit batching ensemble, <italic>p</italic> &#x0003D; 0.8.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-05-863100-g0005.tif"/>
</fig>
<p>The effect of modifying the voting threshold on performance was also tested. The results are depicted in <xref ref-type="fig" rid="F6">Figure 6</xref>.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Ensemble performance vs. voting threshold.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fdata-05-863100-g0006.tif"/>
</fig>
<p>With random batching, a trade-off between recall/balanced accuracy and F1 score was experienced; while subreddit batching demonstrated a trade-off between recall and F1 score, balanced accuracy, and F1 score could be simultaneously improved, with increased voting threshold.</p>
<p>In practice, such approaches are interested in capturing as many burnout samples as possible while maintaining a manageable number of false positives. The threshold can be modified accordingly, for example, aiming to maximize the F1 score under the condition that recall is greater than 0.9. The subreddit batching ensemble with <italic>p</italic> &#x0003D; 0.85 and the random batching ensemble with <italic>p</italic> &#x0003D; 0.5 both demonstrated performance close to such an optimum (as shown in <xref ref-type="table" rid="T9">Table 9</xref>). Both of these ensembles achieve better results than either of the baseline models.</p>
<table-wrap position="float" id="T9">
<label>Table 9</label>
<caption><p>Optimal ensembles (no. test samples = 4,071).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>Test Bal. Acc</bold>.</th>
<th valign="top" align="center"><bold>Test F1</bold></th>
<th valign="top" align="center"><bold>Test Recall</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Subreddit Batching (<italic>p</italic> &#x0003D; 0.85)</td>
<td valign="top" align="center">0.93</td>
<td valign="top" align="center">0.43</td>
<td valign="top" align="center">0.93</td>
</tr>
<tr>
<td valign="top" align="left">Random Batching (<italic>p</italic> &#x0003D; 0.5)</td>
<td valign="top" align="center">0.93</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">0.93</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Finally, a qualitative analysis of test samples incorrectly classified as burnout by the ensemble models revealed posts from the no-burnout dataset that contained topics similar to burnout posts, e.g., work-related, stress, depression, and anxiety. This indicates that the classifiers presented in this work are indeed identifying features related to burnout. It even appears that, in some cases, it may be the labels rather than the predictions that are incorrect, i.e., posts from scraped sub-breddits where users write about experience with burnout.</p>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>4. Discussion</title>
<p>The work presented in this article makes the following contributions:</p>
<list list-type="bullet">
<list-item><p>It demonstrates that NLP methods applied to free text are an effective means to detect indicators for burnout, measured against both a control group of general text and a group composed of text samples related to depression.</p></list-item>
<list-item><p>A machine learning ensemble classifier trained on data from Reddit posts to detect burnout indicators with a promising accuracy is presented.</p></list-item>
<list-item><p>A range of machine learning classifiers trained to detect burnout indicators are compared, showing in particular that the presented ensemble classifier outperforms two single classifier baselines: logistic regression classifiers trained on either a large unbalanced dataset or an undersampled balanced dataset. The best-performing model attained a balanced accuracy of 93%, F1 score of 0.43, and recall of 93% on unbalanced test data.</p></list-item>
</list>
<p>These findings have a large potential to be further developed with an interdisciplinary approach toward a new generation of smart tools for clinical psychology, eventually supporting a wider array of conditions and mental health diagnoses in the future.</p>
<sec>
<title>4.1. Burnout Detection for a Clinical Setting</title>
<p>Extracting data from social media is one of the most commonly used methods in research in this area (e.g., Shen and Rudzicz, <xref ref-type="bibr" rid="B37">2017</xref>; Thorstad and Wolff, <xref ref-type="bibr" rid="B40">2019</xref>). The research presented in this article also relies primarily on data extracted from the social media website Reddit, particularly because it was easy to obtain a large quantity of data to train our model. Nonetheless, clinical data are a more reliable source for detecting burnout due to the certainty of labeling. Clinical data also have the advantage of more closely resembling the data such models are expected to be applied to in the future. A first attempt of working with clinical data to detect burnout has shown promising results. By presenting a dataset from real-world burnout patient data, Nath and Kurpicz-Briki (<xref ref-type="bibr" rid="B30">2021</xref>) managed to go beyond typical burnout detection approaches, which usually includes the use of inventories with scaling questions and worked on applying NLP to mental health. The dataset consisted of data extracted from German-language interviews with burnout patients, a control group, and experts. The authors proceeded to train an SVM classifier on the dataset and ended up achieving accuracy greater than their original baseline.</p>
</sec>
<sec>
<title>4.2. Burnout vs. Depression</title>
<p>A poorer classifier performance on Dataset 4 than on Datasets 2 or 3 was observed. This is likely due to the fact that depression- and burnout-related texts share many similar characteristics. Indeed, depression and burnout are not disjoint categories, and some degree of classification ambiguity is inevitable. This overlap is a significant object of scientific investigation, e.g., by Schonfeld and Bianchi (<xref ref-type="bibr" rid="B34">2016</xref>). The work in this article provides evidence of the non-trivial nature of differentiating burnout and depression. Ongoing work of the authors aims to more closely analyze the markers that indicate and differentiate depression and burnout in free text first-person accounts.</p>
</sec>
<sec>
<title>4.3. Methods for Dealing With Unbalanced Data</title>
<p>Class imbalance is a natural phenomenon in many real-world applications (e.g., fraud detection, tumor detection, software defect prediction). It is well-documented in machine learning literature that unbalanced training data impairs the classification performance of many machine learning models (e.g., Chawla et al., <xref ref-type="bibr" rid="B8">2004</xref>; Garc&#x000ED;a et al., <xref ref-type="bibr" rid="B15">2010</xref>). For example, in cases of extreme class imbalance, models can tend toward placing all samples in the majority class. For a detailed survey on the unbalanced data problem, refer to He and Garcia (<xref ref-type="bibr" rid="B18">2009</xref>). Class imbalance is considered to be intrinsic to the task of burnout detection from real-world (clinical) data, rather than being an artifact of the data collection methods used in this article, and, therefore, it was aimed to address the problem in this work.</p>
<p>Common solutions involve oversampling the minority class or undersampling the majority class to achieve class balance or using cost-sensitive methods that apply a higher penalty to the incorrect classification of samples from the minority class. A number of ensemble methods use oversampling and/or undersampling to train separate models and aggregate their predictions. Successful ensemble methods for unbalanced learning include EasyEnsemble (Liu et al., <xref ref-type="bibr" rid="B25">2008</xref>), SMOTE-Boost (Chawla et al., <xref ref-type="bibr" rid="B9">2003</xref>), UnderBagging (Barandela et al., <xref ref-type="bibr" rid="B4">2003</xref>), and Cluster/SplitBal (Sun et al., <xref ref-type="bibr" rid="B38">2015</xref>).</p>
<p>Sun et al. (<xref ref-type="bibr" rid="B38">2015</xref>) argue that most existing methods might suffer from the loss of potentially useful information and/or overfitting by altering the original data distribution. Of the ensemble methods explored in this work, only UnderBagging, ClusterBal, and SplitBal do not discard data or change the data distribution. These three methods differ mainly in how balanced data batches are constructed and how the predictions of the submodels are aggregated. The method presented in this article is most similar to that of UnderBagging, which was chosen for the ease of implementation in the given setting and the fact that (Sun et al., <xref ref-type="bibr" rid="B38">2015</xref>) found that it performs well across several classifier types. The method presented in this article differs only in that different voting thresholds are considered, not all of the majority class samples are exhausted, and balanced batches based on subgroupings inherent in the presented dataset (subreddits) are constructed.</p>
<p>The single model experiments reflect some of the problems of class imbalance. The best classifiers trained on Dataset 1 reached lower benchmark metrics but demonstrated more consistent performance between training and test data. This is consistent with the expectation that larger training datasets generalize better. Many of the classifiers trained on the unbalanced Dataset 1 performed very poorly, essentially predicting only the majority class. On the other hand, classifiers trained on the balanced Dataset 2 attained a high benchmark performance on relatively small balanced data batches, but that performance dropped considerably (as measured by F1 scores) when applied to highly unbalanced test data. The balanced data models use undersampling and demonstrate the drawbacks of throwing out data points: much of the variance in the <italic>no burnout</italic> dataset is not accounted for, and the undersampling-based models incorrectly classified a relatively large number of more general <italic>no burnout</italic> data. As one would expect, this effect is most pronounced in the models trained using a single subreddit, where a very specialized sample of <italic>no burnout</italic> data were used for training.</p>
<p>Overall, the presented results provide evidence that both undersampling&#x02014;as long as attention is paid to maintaining the variance in the majority class data&#x02014;and ensemble methods are viable approaches to handling the unbalanced data problem in this context. The single logistic regression classifiers trained on undersampled, balanced data performed at a level similar to the ensembles, although the subreddit batching ensemble with <italic>p</italic> &#x0003D; 0.85 and random batching ensemble with <italic>p</italic> &#x0003D; 0.5 both outperformed the single random batching classifier in all three metrics. Undersampling does have the advantage of requiring many fewer training data with both faster training and inference, although this speed difference can be erased by running ensemble submodels in parallel. However, better performance was achieved with ensembles. The ensemble methods provide additional advantages: the voting threshold hyperparameter allows to easily fine-tune the ensemble model according to the relative importance placed on recall and F1 score; in addition, the performance of the ensemble model is more stable, i.e., immune to fluctuations according to the subsample of <italic>no burnout</italic> data used for training.</p>
</sec>
<sec>
<title>4.4. Random vs. Subreddit Batching</title>
<p>As similar performance with both methods for creating balanced data batches was achieved, the experiments do not indicate which, if either, of the two procedures is preferable. However, it was noted that several differences between the two methods exist. Perhaps the most important difference is that the subreddit batching ensembles required fewer training data to achieve the same performance. In addition, as <xref ref-type="table" rid="T6">Table 6</xref> shows that the submodels in the subreddit batching ensemble achieved higher accuracy and F1 score during CV, which might result from the relative ease of distinguishing between burnout-related posts and a single specialized topic with little relation to the condition. This results in overfitting, as reflected in the gap between CV and test results recorded in <xref ref-type="table" rid="T6">Tables 6</xref>, <xref ref-type="table" rid="T7">7</xref>. <xref ref-type="table" rid="T7">Table 7</xref> also shows that the performance of the individual submodels in the subreddit batching ensemble varied much more than for random batching; a comparison with <xref ref-type="table" rid="T8">Table 8</xref> also shows that the relative gain achieved by using ensemble methods over single classifiers was much greater in the case of subreddit batching. This is consistent with expectations and findings in the literature, which suggest that ensembles are an effective method for combining weak learners with considerable variance in their predictions into a strong learner (Schapire, <xref ref-type="bibr" rid="B33">1990</xref>). It is also possible that the subreddits that were excluded from the ensemble due to an insufficient number of posts are over-represented among the misclassified samples and that performance could be improved by including more subreddits. Experiments in this direction are suggested for future work.</p>
</sec>
<sec>
<title>4.5. Recall as an Evaluation Metric</title>
<p>The use of recall as an evaluation metric was chosen because it is assumed that recall is of great importance in real-world applications. In the case of burnout detection, it is better to capture most or all of the true positives at the cost of a manageable number of false positives than to miss positive cases. In practice, marking individuals who are potentially experiencing burnout should help mental health professionals decide which cases should be subjected to further analysis. For this reason, even though the Baseline 1: Unbalanced LR model attained an F1 score better than or on par with the other models (<xref ref-type="table" rid="T8">Table 8</xref>), the significantly lower recall score makes this model unequivocally the least desirable. A tool that misses half of the patients demonstrating potential burnout is not useful.</p>
</sec>
<sec>
<title>4.6. Limitations</title>
<p>In this work, data procured from Reddit posts were used largely because of the ease in obtaining large quantities of data for use in model training. It is expected in the future to apply these methods to the verbal responses of patients in clinical interventions in order to train models to their destined target application. Therefore, the data origin is a limitation of this study. Obtaining a sufficient quantity (and in different local languages) of data for machine learning-based methods poses a significant challenge and will be addressed in future work by other data collection methods, involving also clinical institutions. The authors intend to collaborate with researchers and practitioners in psychology for data collection and to aid in developing a beneficial, easy-to-use clinical tool as well as expanding their work toward other areas of mental health. Another limitation of this work is the diversity in the available data. Being completely anonymous data from online forums, no information about gender, origin, socio-economic background, or similar is available. Therefore, the classifiers presented in this work may not work with the same efficiency for different groups of society. In future work, and before implementing such methods into a product, further validations and potentially additional training data will be required.</p>
</sec>
<sec>
<title>4.7. Future Work</title>
<p>In future work, the authors would like to experiment with more sophisticated ensemble methods, such as those outlined in Sun et al. (<xref ref-type="bibr" rid="B38">2015</xref>), where the general superiority of ensembles over other methods for addressing the class imbalance in several experiments was demonstrated. Since undersampling also showed promising results, more sophisticated methods for undersampling should be explored, such as clustering-based methods (Lin et al., <xref ref-type="bibr" rid="B24">2017</xref>). However, the low variance observed among the random batching submodels may delineate the limits of undersampling-based methods. Furthermore, the use of classifier types beyond logistic regression could be explored, perhaps by incorporating neural network-based models and using other methods for creating balanced data batches for submodel training. Mixing different types of classifiers within an ensemble could be a means to capture <italic>burnout</italic> samples that are otherwise overlooked by logistic regression. It should also be considered to experiment with other vectorization methods in the future, particularly the use of word embeddings learned from deep learning-based language models, such as Word2Vec (Mikolov et al., <xref ref-type="bibr" rid="B27">2017</xref>), GLoVe (Pennington et al., <xref ref-type="bibr" rid="B31">2014</xref>), BERT (Devlin et al., <xref ref-type="bibr" rid="B13">2019</xref>), and fastText (Joulin et al., <xref ref-type="bibr" rid="B21">2016</xref>).</p>
</sec>
</sec>
<sec sec-type="data-availability" id="s5">
<title>Data Availability Statement</title>
<p>The datasets presented in this article are not readily available because the collected online data can be subject to change (e.g., deletions) over time. A similar dataset can be created following the instructions in the article.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>MK-B is the principal investigator of the project and brought the original idea. SN was the main contributor for data collection and the first experiments, in particular, the single classifier models. AP and GM contributed by extending the single classifier models and by developing the ensemble models in collaboration with MK-B. All authors contributed to idea development, experimental setup, and paper writing/editing. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec sec-type="funding-information" id="s7">
<title>Funding</title>
<p>The authors gratefully acknowledge the funding for this work by the Swiss National Science Foundation (SNSF) through an SNSF Spark grant.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s8">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>ADP</surname> <given-names>A. D. P.</given-names></name></person-group> (<year>2018</year>). <source>The Workforce View in Europe 2018.</source> (accessed December 29, 2021).</citation>
</ref>
<ref id="B2">
<citation citation-type="web"><person-group person-group-type="author"><collab>American Psychological Association</collab></person-group> (<year>2019</year>). <source>Stress in America: Stress and Current Events. Stress in America&#x02122; Survey</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.apa.org/news/press/releases/stress/2019/stress-america-2019.pdf">https://www.apa.org/news/press/releases/stress/2019/stress-america-2019.pdf</ext-link></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Azoulay</surname> <given-names>E.</given-names></name> <name><surname>De Waele</surname> <given-names>J.</given-names></name> <name><surname>Ferrer</surname> <given-names>R.</given-names></name> <name><surname>Staudinger</surname> <given-names>T.</given-names></name> <name><surname>Borkowska</surname> <given-names>M.</given-names></name> <name><surname>Povoa</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Symptoms of burnout in intensive care unit specialists facing the covid-19 outbreak</article-title>. <source>Ann. Intensive Care</source> <volume>10</volume>, <fpage>1</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1186/s13613-020-00722-3</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barandela</surname> <given-names>R.</given-names></name> <name><surname>Valdovinos</surname> <given-names>R.</given-names></name> <name><surname>S&#x000E1;nchez</surname> <given-names>J.</given-names></name></person-group> (<year>2003</year>). <article-title>New applications of ensembles of classifiers</article-title>. <source>Pattern Anal. Appl.</source> <volume>6</volume>, <fpage>245</fpage>&#x02013;<lpage>256</lpage>. <pub-id pub-id-type="doi">10.1007/s10044-003-0192-z</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Boe</surname> <given-names>B.</given-names></name></person-group> (<year>2011</year>). <source>PRAW the Python Reddit Api Wrapper.</source> (accessed January 14, 2022).</citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brul&#x000E9;</surname> <given-names>G.</given-names></name> <name><surname>Veenhoven</surname> <given-names>R.</given-names></name></person-group> (<year>2017</year>). <article-title>The &#x02018;10 excess&#x00027; phenomenon in responses to survey questions on happiness</article-title>. <source>Soc. Indicators Res.</source> <volume>131</volume>, <fpage>853</fpage>&#x02013;<lpage>870</lpage>. <pub-id pub-id-type="doi">10.1007/s11205-016-1265-x</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Burisch</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <source>Das Burnout-Syndrom</source>. <publisher-loc>Berlin; Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>.</citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chawla</surname> <given-names>N. V.</given-names></name> <name><surname>Japkowicz</surname> <given-names>N.</given-names></name> <name><surname>Kotcz</surname> <given-names>A.</given-names></name></person-group> (<year>2004</year>). <article-title>Special issue on learning from imbalanced data sets</article-title>. <source>ACM SIGKDD Explor. Newslett.</source> <volume>6</volume>, <fpage>1</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1145/1007730.1007733</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chawla</surname> <given-names>N. V.</given-names></name> <name><surname>Lazarevic</surname> <given-names>A.</given-names></name> <name><surname>Hall</surname> <given-names>L. O.</given-names></name> <name><surname>Bowyer</surname> <given-names>K. W.</given-names></name></person-group> (<year>2003</year>). <article-title>Smoteboost: improving prediction of the minority class in boosting</article-title>, in <source>European Conference on Principles of Data Mining and Knowledge Discovery</source> (<publisher-loc>Cavtat</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>107</fpage>&#x02013;<lpage>119</lpage>.</citation>
</ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Coppersmith</surname> <given-names>G.</given-names></name> <name><surname>Dredze</surname> <given-names>M.</given-names></name> <name><surname>Harman</surname> <given-names>C.</given-names></name></person-group> (<year>2014</year>). <article-title>Quantifying mental health signals in twitter</article-title>, in <source>Proceedings of the Workshop on Computational Linguistics and Clinical Psychology: From Linguistic Signal to Clinical Reality</source> (<publisher-loc>Baltimore, MD</publisher-loc>), <fpage>51</fpage>&#x02013;<lpage>60</lpage>.</citation>
</ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>De Choudhury</surname> <given-names>M.</given-names></name> <name><surname>Counts</surname> <given-names>S.</given-names></name> <name><surname>Horvitz</surname> <given-names>E.</given-names></name></person-group> (<year>2013</year>). <article-title>Predicting postpartum changes in emotion and behavior via social media</article-title>, in <source>Proceedings of the SIGCHI Conference on Human Factors in Computing Systems</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>3267</fpage>&#x02013;<lpage>3276</lpage>.</citation>
</ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>De Choudhury</surname> <given-names>M.</given-names></name> <name><surname>Counts</surname> <given-names>S.</given-names></name> <name><surname>Horvitz</surname> <given-names>E. J.</given-names></name> <name><surname>Hoff</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Characterizing and predicting postpartum depression from shared facebook data</article-title>, in <source>Proceedings of the 17th ACM Conference on Computer Supported Cooperative Work &#x00026; Social Computing</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>626</fpage>&#x02013;<lpage>638</lpage>.</citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Devlin</surname> <given-names>J.</given-names></name> <name><surname>Chang</surname> <given-names>M.-W.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>Toutanova</surname> <given-names>K.</given-names></name></person-group> (<year>2019</year>). <article-title>Bert: pre-training of deep bidirectional transformers for language understanding</article-title>, in <source>Proceedings of NAACL-HLT (1)</source> (<publisher-loc>Minneapolis, MN</publisher-loc>), <fpage>4171</fpage>&#x02013;<lpage>4186</lpage>.</citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Diener</surname> <given-names>E.</given-names></name> <name><surname>Sandvik</surname> <given-names>E.</given-names></name> <name><surname>Pavot</surname> <given-names>W.</given-names></name> <name><surname>Gallagher</surname> <given-names>D.</given-names></name></person-group> (<year>1991</year>). <article-title>Response artifacts in the measurement of subjective well-being</article-title>. <source>Soc. Indicators Res.</source> <volume>24</volume>, <fpage>35</fpage>&#x02013;<lpage>56</lpage>.</citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Garc&#x000ED;a</surname> <given-names>V.</given-names></name> <name><surname>Mollineda</surname> <given-names>R. A.</given-names></name> <name><surname>S&#x000E1;nchez</surname> <given-names>J. S.</given-names></name></person-group> (<year>2010</year>). <article-title>Theoretical analysis of a performance measure for imbalanced data</article-title>, in <source>2010 20th International Conference on Pattern Recognition</source> (<publisher-loc>Istanbul</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>617</fpage>&#x02013;<lpage>620</lpage>.</citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Greenleaf</surname> <given-names>E. A.</given-names></name></person-group> (<year>1992</year>). <article-title>Measuring extreme response style</article-title>. <source>Publ. Opin. Q.</source> <volume>56</volume>, <fpage>328</fpage>&#x02013;<lpage>351</lpage>.</citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Havigerov&#x000E1;</surname> <given-names>J. M.</given-names></name> <name><surname>Haviger</surname> <given-names>J.</given-names></name> <name><surname>Ku&#x0010D;era</surname> <given-names>D.</given-names></name> <name><surname>Hoffmannov&#x000E1;</surname> <given-names>P.</given-names></name></person-group> (<year>2019</year>). <article-title>Text-based detection of the risk of depression</article-title>. <source>Front. Psychol.</source> <volume>10</volume>, <fpage>513</fpage>. <pub-id pub-id-type="doi">10.3389/fpsyg.2019.00513</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>He</surname> <given-names>H.</given-names></name> <name><surname>Garcia</surname> <given-names>E. A.</given-names></name></person-group> (<year>2009</year>). <article-title>Learning from imbalanced data</article-title>. <source>IEEE Trans. Knowl. Data Eng.</source> <volume>21</volume>, <fpage>1263</fpage>&#x02013;<lpage>1284</lpage>. <pub-id pub-id-type="doi">10.1109/TKDE.2008.239</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Holden</surname> <given-names>R. R.</given-names></name></person-group> (<year>2007</year>). <article-title>Socially desirable responding does moderate personality scale validity both in experimental and in nonexperimental contexts</article-title>. <source>Can. J. Behav. Sci./Revue canadienne des sciences du comportement</source> <volume>39</volume>, <fpage>184</fpage>. <pub-id pub-id-type="doi">10.1037/cjbs2007015</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jaggi</surname> <given-names>F.</given-names></name></person-group> (<year>2019</year>). <source>Burnout Praxisnah</source>. <publisher-name>Lehmanns Media</publisher-name>.</citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Joulin</surname> <given-names>A.</given-names></name> <name><surname>Grave</surname> <given-names>E.</given-names></name> <name><surname>Bojanowski</surname> <given-names>P.</given-names></name> <name><surname>Douze</surname> <given-names>M.</given-names></name> <name><surname>J&#x000E9;gou</surname> <given-names>H.</given-names></name> <name><surname>Mikolov</surname> <given-names>T.</given-names></name></person-group> (<year>2016</year>). <article-title>Fasttext. zip: Compressing text classification models</article-title>. <source>arXiv preprint</source> arXiv:1612.03651.</citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Just</surname> <given-names>M. A.</given-names></name> <name><surname>Pan</surname> <given-names>L.</given-names></name> <name><surname>Cherkassky</surname> <given-names>V. L.</given-names></name> <name><surname>McMakin</surname> <given-names>D. L.</given-names></name> <name><surname>Cha</surname> <given-names>C.</given-names></name> <name><surname>Nock</surname> <given-names>M. K.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Machine learning of neural representations of suicide and emotion concepts identifies suicidal youth</article-title>. <source>Nat. Hum. Behav.</source> <volume>1</volume>, <fpage>911</fpage>&#x02013;<lpage>919</lpage>. <pub-id pub-id-type="doi">10.1038/s41562-017-0234-y</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lambert</surname> <given-names>C. E.</given-names></name></person-group> (<year>2013</year>). <source>Identifying Faking on Self-Report Personality Inventories: Relative Merits of Traditional Lie Scales, New Lie Scales, Response Patterns, and Response Times</source> (<publisher-loc>Kingston, ON</publisher-loc>: <publisher-name>Queen&#x00027;s University</publisher-name>).</citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>W.-C.</given-names></name> <name><surname>Tsai</surname> <given-names>C.-F.</given-names></name> <name><surname>Hu</surname> <given-names>Y.-H.</given-names></name> <name><surname>Jhang</surname> <given-names>J.-S.</given-names></name></person-group> (<year>2017</year>). <article-title>Clustering-based undersampling in class-imbalanced data</article-title>. <source>Inf. Sci.</source> <volume>409</volume>, <fpage>17</fpage>&#x02013;<lpage>26</lpage>. <pub-id pub-id-type="doi">10.1016/j.ins.2017.05.008</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.-Y.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name> <name><surname>Zhou</surname> <given-names>Z.-H.</given-names></name></person-group> (<year>2008</year>). <article-title>Exploratory undersampling for class-imbalance learning</article-title>. <source>IEEE Trans. Syst. Man Cybern. B (Cybern.)</source> <volume>39</volume>, <fpage>539</fpage>&#x02013;<lpage>550</lpage>. <pub-id pub-id-type="doi">10.1109/TSMCB.2008.2007853</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>M&#x000E4;ntyl&#x000E4;</surname> <given-names>M.</given-names></name> <name><surname>Adams</surname> <given-names>B.</given-names></name> <name><surname>Destefanis</surname> <given-names>G.</given-names></name> <name><surname>Graziotin</surname> <given-names>D.</given-names></name> <name><surname>Ortu</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Mining valence, arousal, and dominance - possibilities for detecting burnout and productivity?</article-title> <source>CoRR</source>, abs/1603.04287.</citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mikolov</surname> <given-names>T.</given-names></name> <name><surname>Grave</surname> <given-names>E.</given-names></name> <name><surname>Bojanowski</surname> <given-names>P.</given-names></name> <name><surname>Puhrsch</surname> <given-names>C.</given-names></name> <name><surname>Joulin</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). <article-title>Advances in pre-training distributed word representations</article-title>. <source>arXiv preprint</source> arXiv:1712.09405.</citation>
</ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Morales</surname> <given-names>M.</given-names></name> <name><surname>Dey</surname> <given-names>P.</given-names></name> <name><surname>Theisen</surname> <given-names>T.</given-names></name> <name><surname>Belitz</surname> <given-names>D.</given-names></name> <name><surname>Chernova</surname> <given-names>N.</given-names></name></person-group> (<year>2019</year>). <article-title>An investigation of deep learning systems for suicide risk assessment</article-title>, in <source>Proceedings of the Sixth Workshop on Computational Linguistics and Clinical Psychology</source> (<publisher-loc>Minneapolis, MN</publisher-loc>), <fpage>177</fpage>&#x02013;<lpage>181</lpage>.</citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moreno</surname> <given-names>M. A.</given-names></name> <name><surname>Jelenchick</surname> <given-names>L. A.</given-names></name> <name><surname>Egan</surname> <given-names>K. G.</given-names></name> <name><surname>Cox</surname> <given-names>E.</given-names></name> <name><surname>Young</surname> <given-names>H.</given-names></name> <name><surname>Gannon</surname> <given-names>K. E.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Feeling bad on facebook: depression disclosures by college students on a social networking site</article-title>. <source>Depress. Anxiety</source> <volume>28</volume>, <fpage>447</fpage>&#x02013;<lpage>455</lpage>. <pub-id pub-id-type="doi">10.1002/da.20805</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nath</surname> <given-names>S.</given-names></name> <name><surname>Kurpicz-Briki</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>Burnoutwords - detecting burnout for a clinical setting</article-title>, in <source>Proceedings of the 10th International Conference on Soft Computing, Artificial Intelligence and Applications (SCAI 2021), CS &#x00026; IT Conference Proceedings</source> (<publisher-loc>Zurich</publisher-loc>), vol. <volume>11</volume>, <fpage>177</fpage>&#x02013;<lpage>191</lpage>.</citation>
</ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pennington</surname> <given-names>J.</given-names></name> <name><surname>Socher</surname> <given-names>R.</given-names></name> <name><surname>Manning</surname> <given-names>C. D.</given-names></name></person-group> (<year>2014</year>). <article-title>Glove: global vectors for word representation</article-title>, in <source>Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source> (<publisher-loc>Doha</publisher-loc>), <fpage>1532</fpage>&#x02013;<lpage>1543</lpage>.</citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rui</surname> <given-names>Y.</given-names></name></person-group> (<year>2017</year>). <article-title>From artificial intelligence to augmented intelligence</article-title>. <source>IEEE MultiMedia</source> <volume>24</volume>, <fpage>4</fpage>&#x02013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1109/MMUL.2017.8</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schapire</surname> <given-names>R. E.</given-names></name></person-group> (<year>1990</year>). <article-title>The strength of weak learnability</article-title>. <source>Mach. Learn.</source> <volume>5</volume>, <fpage>197</fpage>&#x02013;<lpage>227</lpage>.</citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schonfeld</surname> <given-names>I. S.</given-names></name> <name><surname>Bianchi</surname> <given-names>R.</given-names></name></person-group> (<year>2016</year>). <article-title>Burnout and depression: two entities or one?</article-title> <source>J. Clin. Psychol.</source> <volume>72</volume>, <fpage>22</fpage>&#x02013;<lpage>37</lpage>. <pub-id pub-id-type="doi">10.1002/jclp.22229</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Schwartz</surname> <given-names>H. A.</given-names></name> <name><surname>Eichstaedt</surname> <given-names>J.</given-names></name> <name><surname>Kern</surname> <given-names>M.</given-names></name> <name><surname>Park</surname> <given-names>G.</given-names></name> <name><surname>Sap</surname> <given-names>M.</given-names></name> <name><surname>Stillwell</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Towards assessing changes in degree of depression through facebook</article-title>, in <source>Proceedings of the Workshop on Computational Linguistics and Clinical Psychology: From linguistic Signal to Clinical Reality</source> (<publisher-loc>Baltimore, MD</publisher-loc>), <fpage>118</fpage>&#x02013;<lpage>125</lpage>.</citation>
</ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><collab>SECO</collab></person-group> (<year>2015</year>). <source>The Sixth European Working Conditions Survey (EWCS).</source> (accessed December 29, 2021).</citation>
</ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Shen</surname> <given-names>J. H.</given-names></name> <name><surname>Rudzicz</surname> <given-names>F.</given-names></name></person-group> (<year>2017</year>). <article-title>Detecting anxiety through reddit</article-title>, in <source>Proceedings of the Fourth Workshop on Computational Linguistics and Clinical Psychology-From Linguistic Signal to Clinical Reality</source> (<publisher-loc>Vancouver, BC</publisher-loc>), <fpage>58</fpage>&#x02013;<lpage>65</lpage>.</citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>Z.</given-names></name> <name><surname>Song</surname> <given-names>Q.</given-names></name> <name><surname>Zhu</surname> <given-names>X.</given-names></name> <name><surname>Sun</surname> <given-names>H.</given-names></name> <name><surname>Xu</surname> <given-names>B.</given-names></name> <name><surname>Zhou</surname> <given-names>Y.</given-names></name></person-group> (<year>2015</year>). <article-title>A novel ensemble method for classifying imbalanced data</article-title>. <source>Pattern Recogn.</source> <volume>48</volume>, <fpage>1623</fpage>&#x02013;<lpage>1637</lpage>. <pub-id pub-id-type="doi">10.1155/2017/1827016</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tadesse</surname> <given-names>M. M.</given-names></name> <name><surname>Lin</surname> <given-names>H.</given-names></name> <name><surname>Xu</surname> <given-names>B.</given-names></name> <name><surname>Yang</surname> <given-names>L.</given-names></name></person-group> (<year>2019</year>). <article-title>Detection of depression-related posts in reddit social media forum</article-title>. <source>IEEE Access</source> <volume>7</volume>, <fpage>44883</fpage>&#x02013;<lpage>44893</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2019.2909180</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thorstad</surname> <given-names>R.</given-names></name> <name><surname>Wolff</surname> <given-names>P.</given-names></name></person-group> (<year>2019</year>). <article-title>Predicting future mental illness from social media: a big-data approach</article-title>. <source>Behav. Res. Meth.</source> <volume>51</volume>, <fpage>1586</fpage>&#x02013;<lpage>1600</lpage>. <pub-id pub-id-type="doi">10.3758/s13428-019-01235-z</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Warriner</surname> <given-names>A. B.</given-names></name> <name><surname>Kuperman</surname> <given-names>V.</given-names></name> <name><surname>Brysbaert</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>Norms of valence, arousal, and dominance for 13,915 english lemmas</article-title>. <source>Behav. Res. Meth.</source> <volume>45</volume>, <fpage>1191</fpage>&#x02013;<lpage>1207</lpage>. <pub-id pub-id-type="doi">10.3758/s13428-012-0314-x</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Williams</surname> <given-names>M. M.</given-names></name> <name><surname>Rogers</surname> <given-names>R.</given-names></name> <name><surname>Sharf</surname> <given-names>A. J.</given-names></name> <name><surname>Ross</surname> <given-names>C. A.</given-names></name></person-group> (<year>2019</year>). <article-title>Faking good: an investigation of social desirability and defensiveness in an inpatient sample with personality disorder traits</article-title>. <source>J. Pers. Assess.</source> <volume>101</volume>, <fpage>253</fpage>&#x02013;<lpage>263</lpage>. <pub-id pub-id-type="doi">10.1080/00223891.2018.1455691</pub-id></citation>
</ref>
</ref-list>
<fn-group>
<fn id="fn0001"><p><sup>1</sup><ext-link ext-link-type="uri" xlink:href="https://icd.who.int/browse11/l-m/en&#x00023;/http://id.who.int/icd/entity/129180281">https://icd.who.int/browse11/l-m/en&#x00023;/http://id.who.int/icd/entity/129180281</ext-link></p></fn>
<fn id="fn0002"><p><sup>2</sup><ext-link ext-link-type="uri" xlink:href="https://spacy.io/">https://spacy.io/</ext-link></p></fn>
<fn id="fn0003"><p><sup>3</sup>For the unbalanced dataset, a 70&#x02013;30% split was taken on each class separately in order to ensure that the training and test sets had roughly the same class distribution.</p></fn>
<fn id="fn0004"><p><sup>4</sup>In our experiments, 10 &#x02264; <italic>n</italic> &#x02264; 20.</p></fn>
<fn id="fn0005"><p><sup>5</sup>Here, <italic>p</italic> is a threshold that may vary. Values of 0.4 &#x02264; <italic>p</italic> &#x02264; 0.99 were used.</p></fn>
<fn id="fn0006"><p><sup>6</sup>Here, we do not need to take the balanced accuracy because the submodels are trained on balanced datasets.</p></fn>
<fn id="fn0007"><p><sup>7</sup>The ensemble experiments revealed that logistic regression classifiers trained on Datasets 2 and 3 do not perform as well on a highly unbalanced data sampled from Dataset 1. This will be further discussed in Section 3.2.</p></fn>
</fn-group>
</back>
</article>