<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frai.2023.1230649</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Supervised machine learning models for depression sentiment analysis</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Obagbuwa</surname> <given-names>Ibidun Christiana</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2326835/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Danster</surname> <given-names>Samantha</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Chibaya</surname> <given-names>Onil Colin</given-names></name>
</contrib>
</contrib-group>
<aff><institution>Department of Computer Science and Information Technology, School of Natural and Applied Sciences, Sol Plaatje University</institution>, <addr-line>Kimberley</addr-line>, <country>South Africa</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Ayodele Adebiyi, Landmark University, Nigeria</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Bassey Isong, North-West University, South Africa; Pius Adewale Owolawi, Tshwane University of Technology, South Africa</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Ibidun Christiana Obagbuwa <email>Ibidun.obagbuwa&#x00040;spu.ac.za</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>19</day>
<month>07</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>6</volume>
<elocation-id>1230649</elocation-id>
<history>
<date date-type="received">
<day>29</day>
<month>05</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>29</day>
<month>06</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2023 Obagbuwa, Danster and Chibaya.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Obagbuwa, Danster and Chibaya</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract>
<sec>
<title>Introduction</title>
<p>Globally, the prevalence of mental health problems, especially depression, is at an all-time high. The objective of this study is to utilize machine learning models and sentiment analysis techniques to predict the level of depression earlier in social media users&#x00027; posts.</p></sec>
<sec>
<title>Methods</title>
<p>The datasets used in this research were obtained from Twitter posts. Four machine learning models, namely extreme gradient boost (XGB) Classifier, Random Forest, Logistic Regression, and support vector machine (SVM), were employed for the prediction task.</p></sec>
<sec>
<title>Results</title>
<p>The SVM and Logistic Regression models yielded the most accurate results when applied to the provided datasets. However, the Logistic Regression model exhibited a slightly higher level of accuracy compared to SVM. Importantly, the logistic regression model demonstrated the advantage of requiring less execution time.</p></sec>
<sec>
<title>Discussion</title>
<p>The findings of this study highlight the potential of utilizing machine learning models and sentiment analysis techniques for early detection of depression in social media users. The effectiveness of SVM and Logistic Regression models, with Logistic Regression being more efficient in terms of execution time, suggests their suitability for practical implementation in real-world scenarios.</p></sec></abstract>
<kwd-group>
<kwd>Twitter</kwd>
<kwd>depression</kwd>
<kwd>sentiment analysis</kwd>
<kwd>text pre-processing</kwd>
<kwd>machine learning techniques</kwd>
<kwd>social media</kwd>
<kwd>natural language processing</kwd>
<kwd>mental health</kwd>
</kwd-group>
<counts>
<fig-count count="15"/>
<table-count count="2"/>
<equation-count count="0"/>
<ref-count count="32"/>
<page-count count="12"/>
<word-count count="6510"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Medicine and Public Health</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1. Introduction</title>
<p>It is critical to understand people&#x00027;s emotions and daily online activities. Many researchers are interested in this topic because depression is a major cause of mental health problems that manifest themselves through social media posts. Twitter is one of the most popular social media platforms, with many people using it for person-to-person communication and sharing common interests based on their perspectives on real-life events (Ricard et al., <xref ref-type="bibr" rid="B19">2018</xref>; Sood et al., <xref ref-type="bibr" rid="B25">2018</xref>). Sentiment analysis can be used to monitor various social media sites in real-time. In short, Twitter will be used to classify the sentiment polarity of a tweet as positive, negative, or neutral (Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>), because tweets and written text appear to be incomplete and unstructured in nature.</p>
<p>This study uses a machine learning approach to create models that will help identify depressed social media users or persons earlier and help before it is too late. To accomplish this, data pre-processing, which included data cleaning, tokenization, stop words removal, stemming, lemmatization, bigram creation, sentiment classification, duplicate removal, and URL and number removal to improve tweet content was carried out. Machine learning classifiers: XGB Classifier, Random Forest, Logistic Regression, and support vector machine was used to build depression sentiment models, and the four classifiers performed excellently on the datasets.</p>
<p>The remainder of the paper is organized as follows: Section 2 introduces existing relevant work in the literature. Section 3 explains how the model is developed methodologically and practically. Section 4 displays the outcomes of our proposed approach&#x00027;s performance. Section 5 highlights the main points of the research. Finally, Section 6 presents the main conclusions of this paper.</p>
</sec>
<sec id="s2">
<title>2. Literature review</title>
<sec>
<title>2.1. Depression</title>
<p>Depression is a major public health issue that affects people psychologically all over the world. It is defined as a collection of mixed impairment symptoms and disturbance in one&#x00027;s cognition and behavior (Orabi et al., <xref ref-type="bibr" rid="B15">2018</xref>). During the 2019&#x02013;2021 COVID era, there was a rapid increase in mental health issues and suicidal cases (Zulfiker et al., <xref ref-type="bibr" rid="B32">2021</xref>). The World Health Organization states that more than 300 million people worldwide suffer from depression, prompting many researchers to focus on this topic (Priya et al., <xref ref-type="bibr" rid="B17">2020</xref>). Chronic diseases can also be caused by depression. <xref ref-type="fig" rid="F1">Figure 1</xref> illustrates the symptoms of depression in a detailed manner.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Depression symptoms (Source: The University of Queensland, <xref ref-type="bibr" rid="B27">2022</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1230649-g0001.tif"/>
</fig>
<p>Depression affects men and women differently, in such a way that women experience depression more severely than men (Seney et al., <xref ref-type="bibr" rid="B22">2018</xref>). Because women are more likely to have anxiety disorders. Women are more involved in social gatherings, which exposes them to intimacy and emotional disclosure. Mental health issues can also be exacerbated by modern lifestyles influenced by social media and the pressure to live up to a certain standard that requires the approval of others (Kumar et al., <xref ref-type="bibr" rid="B12">2020</xref>). Furthermore, studies revealed that younger individuals are disproportionately affected by mental health issues such as anxiety, depression, and obsessive-compulsive behavior, which has led many to commit suicide or consider it (Orsolini et al., <xref ref-type="bibr" rid="B16">2021</xref>).</p>
</sec>
<sec>
<title>2.2. Social media posts linked with mental health issues on social media platforms</title>
<p>People use social media to share and communicate their ideas, and emotional states of being across many social platforms. Imagery is also a popular method of self-expression on social media platforms such as instagram (Mun and Kim, <xref ref-type="bibr" rid="B14">2021</xref>). More evidence of depression can be found in Instagram photos (Smith and Anderson, <xref ref-type="bibr" rid="B24">2018</xref>). As a result, this method effectively elicits deeper psychological consciousness by allowing emotional expressions that cannot be expressed in writing.</p>
<p>Another popular way to deal with mental health issues is through expressive writing or texting. Users of these social media platforms tend to document and narrate their lives through these platforms making it easy for us to understand their personal lives (Orabi et al., <xref ref-type="bibr" rid="B15">2018</xref>). Twitter and Facebook are also popular social media platforms and important for person-to-person communication where people share their perspectives on real-life events (Ricard et al., <xref ref-type="bibr" rid="B19">2018</xref>; Sood et al., <xref ref-type="bibr" rid="B25">2018</xref>). Thus, depression can be identified through word sentiment analysis (Seabrook et al., <xref ref-type="bibr" rid="B21">2018</xref>). According to Ricard et al. (<xref ref-type="bibr" rid="B19">2018</xref>), community-generated content responses can be used to identify levels of depression in people who have similar user-generated content. In retrospect, social media can also be beneficial to people&#x00027;s mental health.</p>
</sec>
<sec>
<title>2.3. Different methods of mining social media data</title>
<p>Before this paper can go into what social media mining entails, it is necessary to comprehend what this topic stems from. Social media mining stems from data mining (Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>), which could also be considered as a stem of machine learning in various fields (Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>). Data mining can be considered the process of finding useful sets of data from larger pools of accessible data (Jagadishwari et al., <xref ref-type="bibr" rid="B8">2021</xref>; Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>). In some studies, this is considered an automated process of uncovering knowledge, relationships, and patterns in interrelated sets of data (Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>). Data mining techniques are used in the process of social media mining for several reasons and several methods could be implemented in this process.</p>
<p>The growth and impact of social media have grown rapidly over the past few years. With this growth, more and more data has become available for both private and commercial use through these platforms. Noticeably, we are all constantly interacting with each other through these social media platforms, and one could say, our whole lives are now documented through this data. From whom we talk to, whom we know, and what we like or dislike, this information is accessible to almost anyone through these platforms, with this, the analysis of people or their social cues can be done through this data. Social media mining can be considered a process of gathering interrelated data from social media platforms to identify patterns and relationships in the data (Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>). There are several forms in which this data could be found. Text, image, and voice are three common forms of data that are mined during the social media mining process (Jagadishwari et al., <xref ref-type="bibr" rid="B8">2021</xref>).</p>
<sec>
<title>2.3.1. Text mining</title>
<p>This is the process of identifying relationships and patterns from large amounts of textual data to discover new knowledge or generate an understanding of some sort (Gaikwad et al., <xref ref-type="bibr" rid="B7">2014</xref>). This process makes use of data mining algorithms and techniques such as classification, clustering, and association rules to discover new information and relationships in textual sources (Gaikwad et al., <xref ref-type="bibr" rid="B7">2014</xref>; Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>). Several methods and techniques have been developed to solve the text mining problem according to the user&#x00027;s requirements, some of these methods include the term-based method (TBM) and the phrase-based method (PBM). The term-based method searches large amounts of text looking for similar terms, in this context term refers to text with a related language or logic (Gaikwad et al., <xref ref-type="bibr" rid="B7">2014</xref>). The term-based method has the advantage of efficient computational performance and mature theories for term weighing. However, this method suffers from the problems of polysemy and synonymy (Gaikwad et al., <xref ref-type="bibr" rid="B7">2014</xref>). Where polysemy refers to words with multiple meanings and synonymy refers to words with similar meanings. In the phrase-based method, the text is analyzed in phrases as these are less ambiguous and more discriminative in comparison to terms (Gaikwad et al., <xref ref-type="bibr" rid="B7">2014</xref>). However, these phrase-based methods perform relatively poorly in comparison to a term-based method, and this could be because how phrases have inferior statistical properties in comparison to terms, these phrases also have a low occurrence frequency and large numbers of noise and redundant phrases are located among them (Gaikwad et al., <xref ref-type="bibr" rid="B7">2014</xref>). <xref ref-type="fig" rid="F2">Figure 2</xref> depicts the process of text mining.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Text mining process (Source: Gaikwad et al., <xref ref-type="bibr" rid="B7">2014</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1230649-g0002.tif"/>
</fig>
</sec>
<sec>
<title>2.3.2. Image mining</title>
<p>Image mining is a subset of multimedia mining, which is used to extract informative and interesting graphical data (Shukla and Vala, <xref ref-type="bibr" rid="B23">2016</xref>). Image mining, unlike computer vision and image processing techniques, focuses on extracting patterns from large collections of images (Shukla and Vala, <xref ref-type="bibr" rid="B23">2016</xref>). Computer vision and image processing use a single image to understand or extract specific features (Shukla and Vala, <xref ref-type="bibr" rid="B23">2016</xref>). Some of the common techniques centered around image mining include object recognition, image retrieval, image indexing, classification, clustering, and accusation rule mining (Shukla and Vala, <xref ref-type="bibr" rid="B23">2016</xref>). The image mining process is shown in <xref ref-type="fig" rid="F3">Figure 3</xref> below.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>The image mining process (Source: Shukla and Vala, <xref ref-type="bibr" rid="B23">2016</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1230649-g0003.tif"/>
</fig>
</sec>
<sec>
<title>2.3.3. Voice mining</title>
<p>Voice mining is another subset of multimedia mining (Shukla and Vala, <xref ref-type="bibr" rid="B23">2016</xref>). The use of voice has become very popular on social media platforms as these vocal messages contain emotional information and this emotional information has become a new topic in data mining and social media analytics (Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>). While it is not as popular as text mining or image mining now, there is clear growth, and this may become a major subset of the data mining and social media analytics sector.</p>
<p>For this study, Text mining techniques were utilized, as well as data mining algorithms and techniques to detect the rate of depression found in posts made by individuals on the Twitter social media application. Some of the most popular text mining methods are term-based mining (TBM) and phrase-based mining (PBM). Making use of single words to identify depression would not work because the word could mean different things depending on how it is used but considering how people generally use slang when posting on social media and all that, the models would have to adapt to this slang in order to make sense of the sentences. As a result, crucial clues regarding the sentences the word is used in or the context in which it is used would have been overlooked if a single word were to be identified as a trigger for depression.</p>
</sec>
</sec>
<sec>
<title>2.4. Sentiment analysis for depression prediction</title>
<p>This is a growing topic that is used to understand people&#x00027;s sentiments about their everyday lives. This can be defined as a classification of text blocks, traditionally as either neutral, negative, or positive (Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>). However, this does not only depend on the polarity of the text, but the emotions associated with it, be it happy, sad angry, etc. Many studies have been conducted and commonly, various Natural Language Processing algorithms are used (Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>). Studies also show that Binary and Ternary classification techniques are regularly used, with multi-class classification providing more accurate results (Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>). Multi-class classification divides the data into multiple sub-classes and then works on these sub-classes separately based on the class&#x00027;s polarities (Gaikwad et al., <xref ref-type="bibr" rid="B7">2014</xref>; Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>). Deep learning techniques are also used for the classification process (Gaikwad et al., <xref ref-type="bibr" rid="B7">2014</xref>; Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>). However, the two common sentiment analysis techniques are rule-based sentiment analysis and machine learning-based sentiment analysis.</p>
<p>Rule-based sentiment analysis makes use of rules and word collections labeled by polarity to identify the opinion or context of the text (Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>). In this technique, sentiment value is made up of a combination of attributes to understand sarcasm, negation, or dependent clauses (Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>).</p>
<p>Machine learning-based sentiment analysis is focused on training a machine learning model using a sentiment-labeled training set (AlSagri and Ykhlef, <xref ref-type="bibr" rid="B3">2020</xref>; Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>), to train the model to understand the polarity of words given in a certain order.</p>
<p>Sentiment analysis can be broken down into four major processes, namely, (1) data collection, (2) text preparation (data preprocessing), (3) sentiment detection (feature extraction), and (4) sentiment classification and presentation as output (Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>). The data collection process is made to allow for the relevant data to be collected, in this case, from various social media platforms. In a paper by Samsari et al. (<xref ref-type="bibr" rid="B20">2022</xref>) this data collection process consisted of a dataset made up of tweets collected during the COVID pandemic. Similarly in an article by Lui (<xref ref-type="bibr" rid="B13">2020</xref>), the data used in the study was collected from the Facebook and Twitter social media platforms. AlSagri and Ykhlef (<xref ref-type="bibr" rid="B3">2020</xref>), also collected data from tweets, and like most studies that made use of the text mining method, emoticons and emojis were removed as part of the preprocessing process (Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>). The data preparation process is used to clean the data by removing unrelated data and removing noise and words with no analytical significance (Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>). The third is the sentiment detection process, here the text is analyzed to extract opinions, reviews, and feedback while removing any text related to facts or popular knowledge (Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>). Thereafter, the classification and presentation process are conducted (Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>). Samsari et al. (<xref ref-type="bibr" rid="B20">2022</xref>) made use of the Na&#x000EF;ve Bayes classifier to classify the data as either positive, negative, or neutral. Jagadishwari et al. (<xref ref-type="bibr" rid="B8">2021</xref>) made use of the Linear regression model as well as the Support Vector Machine (SMV) model, and these two models generated similar results. Ranganathan and Tzacheva (<xref ref-type="bibr" rid="B18">2019</xref>) used the Support Vector Machine LibLinear in an article titled Emotion Mining in Social Media Data. Tiwari et al. (<xref ref-type="bibr" rid="B28">2021</xref>) analyzed the performance of five different classification models and the most accurate results were generated by the Decision Tree classifier. The sentiment analysis process is shown in <xref ref-type="fig" rid="F4">Figure 4</xref> below.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Sentiment analysis process (Source: Babu and Kanaga, <xref ref-type="bibr" rid="B4">2022</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1230649-g0004.tif"/>
</fig>
<p>These studies suggest that the two most common sentiment analysis methods used are the Rule-based method and the Machine learning-based method (AlSagri and Ykhlef, <xref ref-type="bibr" rid="B3">2020</xref>). Regardless of the method used, there are four major processes in the sentiment analysis life cycle as mentioned, (1) data collection, (2) text preparation (data preprocessing), (3) sentiment detection (feature extraction), and finally (4) sentiment classification and preparation for output. Na&#x000EF;ve Bayes, Decision Trees, SVM&#x00027;s, and Linear regression are some of the most common classification models used. However, studies suggest that Decision Trees have shown more accurate results when trained correctly (AlSagri and Ykhlef, <xref ref-type="bibr" rid="B3">2020</xref>).</p>
</sec>
</sec>
<sec id="s3">
<title>3. Methodology</title>
<sec>
<title>3.1. Data collection and preparation</title>
<p>Four separate Twitter datasets were collected from Kaggle to narrow it down to three columns namely as one dataset: Tweet texts.</p>
<list list-type="simple">
<list-item><p>(i) Target (0, 1, 2)</p></list-item>
<list-item><p>(ii) The sentiment (positive, negative, and neutral)</p></list-item>
</list>
<p><xref ref-type="fig" rid="F5">Figure 5</xref> is an illustration of the before and after pre-processing of the datasets merged as one dataset.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Before and after pre-processing.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1230649-g0005.tif"/>
</fig>
<sec>
<title>3.1.1. Pre-processing</title>
<p>This process improves the quality of the dataset based on the tweets of the users. There are four datasets retrieved from the Kaggle website on depression. The datasets were cleaned the removing unnecessary features and merging the four datasets as one. For the text (Tweets) pre-processing, Natural Language Toolkit (NLTK), an open-source Python library for natural language processing techniques was employed to perform the following tasks:</p>
<list list-type="simple">
<list-item><p>i. Tokenization&#x02014;Users&#x00027; tweets are divided into several tokens, making stemming and word removal easier.</p></list-item>
<list-item><p>ii. Removal of Stop Words&#x02014;Eliminating stop words such as &#x0201C;on,&#x0201D; &#x0201C;at,&#x0201D; and &#x0201C;the&#x0201D; to improve algorithm processing time.</p></list-item>
<list-item><p>iii. Stemming&#x02014;Using stemming to identify the root of words in user tweets.</p></list-item>
<list-item><p>iv. Lemmatization&#x02014;The &#x0201C;text normalization technique&#x0201D; will be used to bring tweets or words to their dictionary form. This process is like stemming, but the root words have meaning.</p></list-item>
<list-item><p>v. Creation of bigrams/trigrams&#x02014;A bigram is two consecutive words in a sentence, while a trigram is three consecutive words in a sentence.</p></list-item>
</list>
<p>Furthermore, the tweets were classified as negative, positive, or neutral. Resulting in using negative reviews in relation to depression because texts or tweets about depression are perceived as negative. Duplicates were removed, sample description was carried out to determine how many ids and tweets are unique. Stop words were removed from the dataset, and special characters and links were replaced with blank spaces. URL links were removed from the corpus to improve tweet content. Numbers were removed because they are not useful for measuring sentiments. In addition, all the text was changed to lowercase.</p>
</sec>
<sec>
<title>3.1.2. Appling target value to the different sentiments</title>
<p>Positive Sentiment: Target = 0</p>
<p>Negative Sentiment: Target = 1</p>
<p>Neutral Sentiment: Target = 2</p>
</sec>
<sec>
<title>3.1.3. Baseline and evaluation</title>
<p>Two classification methods were used in this study. The first was a numerical classifier in which tweets were classified in a range of one to four, then the second was a three-way classifier which classified the tweets according to their polarity as either, negative, positive, or neutral. The numerical classifier was performed on all the datasets in order to generate a common target value. However, after the pre-processing stage, a single dataset containing a pre-existing sentiment column was used. The standard C-Method was then used in this research as a starting point technique and applied all six pre-processing methods, including removing URLs, removing stop words, removing numbers, reverting words that contain repeated letters to their original form, replacing negative mentions, and expanding acronyms to the original word. The accuracy and computational time are used to measure the overall classification process while the text pre-processing is measured by the loss or gain of accuracy.</p>
</sec>
<sec>
<title>3.1.4. Sentiment visualization</title>
<p>To determine the most prevalent words, this study used word clouds in our dataset according to each sentiment (positive, negative, and neutral). Word clouds visualize the most frequent words in large sizes and the less frequent words in smaller sizes.</p>
<p>The classification of a tweet&#x00027;s sentiment polarity is depicted in <xref ref-type="fig" rid="F6">Figures 6</xref>&#x02013;<xref ref-type="fig" rid="F8">8</xref>. Word clouds were used to visualize the Tweets&#x00027; Sentiment Polarity. <xref ref-type="fig" rid="F6">Figure 6</xref> depicts the most common words in the entire dataset, <xref ref-type="fig" rid="F7">Figure 7</xref> shows the most common positive words, and <xref ref-type="fig" rid="F8">Figure 8</xref> depicts the most negative/depressed words.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>The most common words in the entire dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1230649-g0006.tif"/>
</fig>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>The most common positive words.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1230649-g0007.tif"/>
</fig>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>The most negative/depressed words.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1230649-g0008.tif"/>
</fig>
</sec>
<sec>
<title>3.1.5. Datasets</title>
<p>The datasets used for this study were a collection of Twitter datasets related to depression and sentiment analysis from the Kaggle website. The pre-processing stage of the study was a little difficult due to the different structures of the datasets as some datasets contained target values pre-set sentiments while others did not. Similar columns that were required were the tweets column and the id column.</p>
<p>Out of the five datasets used, the &#x0201C;training.1600000. processed.noemoticon&#x0201D; dataset was the most useful. This dataset contained 1,599,999 rows &#x000D7; 6 columns. Of the six features, only two of the six features were concentrated on: the target column and the TextTweet. The Clean_TweetText column was then added that contained the cleaned tweets. <xref ref-type="fig" rid="F9">Figure 9</xref> shows the complete dataset before feature selection, and <xref ref-type="fig" rid="F10">Figure 10</xref> shows the dataset after feature selection was done.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p>Complete dataset before selection of columns.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1230649-g0009.tif"/>
</fig>
<fig id="F10" position="float">
<label>Figure 10</label>
<caption><p>Dataset after column selection.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1230649-g0010.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec id="s4">
<title>4. Performance evaluation metrics</title>
<p>In this study, the evaluation of four machine learning models, namely XGB Classifier, Random Forest Classifier, Logistic Regression, and Support Vector Machine C-Support Vector Classification Model, was performed using the confusion matrix. The effectiveness of the models&#x00027; predictions was assessed using metrics such as the accuracy score.</p>
<p>The confusion matrix shown in <xref ref-type="fig" rid="F11">Figure 11</xref> is organized into four categories:</p>
<list list-type="order">
<list-item><p>True Positives (TP): Instances where the model correctly predicts tweets expressing depression sentiment.</p></list-item>
<list-item><p>True Negatives (TN): Instances where the model correctly predicts tweets not expressing depression sentiment.</p></list-item>
<list-item><p>False Positives (FP): Instances where the model incorrectly predicts tweets as expressing depression sentiment when they do not (a Type I error).</p></list-item>
<list-item><p>False Negatives (FN): Instances where the model incorrectly predicts tweets as not expressing depression sentiment when they do (a Type II error).</p></list-item>
</list>
<fig id="F11" position="float">
<label>Figure 11</label>
<caption><p>Confusion matrix (Draelos, <xref ref-type="bibr" rid="B6">2019</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1230649-g0011.tif"/>
</fig>
<p>The confusion matrix allows us to calculate several evaluation metrics:</p>
<list list-type="order">
<list-item><p>Accuracy: It measures the overall correctness of the model and is calculated as (TP &#x0002B; TN)/(TP &#x0002B; TN &#x0002B; FP &#x0002B; FN). The accuracy values for different classifiers are given for comparison. The accuracy metric is commonly used to evaluate the performance of a classification model.</p></list-item>
<list-item><p>Precision: It indicates the proportion of correctly predicted positive instances out of the total instances predicted as positive. Precision is calculated as TP/(TP &#x0002B; FP).</p></list-item>
<list-item><p>Recall (Sensitivity or True Positive Rate): It represents the proportion of correctly predicted positive instances out of the total actual positive instances. Recall is calculated as TP/(TP &#x0002B; FN).</p></list-item>
<list-item><p>F1 Score: It is the harmonic mean of precision and recall, providing a balance between the two metrics. The F1 score is calculated as 2 <sup>&#x0002A;</sup> (Precision <sup>&#x0002A;</sup> Recall)/(Precision &#x0002B; Recall).</p></list-item>
</list>
</sec>
<sec id="s5">
<title>5. Experiment and results</title>
<p>This section looks at the results presented by the individual models before discussing which of the models was best and how this was concluded. The comparison process and rating criteria were based on two factors. The accuracy of the model and the time the model took to execute. Our final study looked at comparing four models on the same dataset.</p>
<sec>
<title>5.1. Machine learning classifier</title>
<p>The four models we looked at included python XGB Classifier, Random Forest Classifier, Logistic Regression, and Support Vector Machine C-Support Vector Classification Model. The four models are described in the Sections 5.1.1&#x02013;5.1.4.</p>
<sec>
<title>5.1.1. XGB classifier</title>
<p>XGB Classifier is a machine learning model popular for its speed and accuracy, and it is widely used in different industries for solving classification problems. This model is primarily designed to solve classification problems by creating a set of decision trees iteratively, hence uses a decision tree ensemble method called Gradient Boosting. Moreover, in each iteration, the model identifies the instances that were not classified correctly in the previous iteration and focuses on them to improve the accuracy of the model. <xref ref-type="fig" rid="F12">Figure 12</xref> illustrates the functioning process of the model.</p>
<fig id="F12" position="float">
<label>Figure 12</label>
<caption><p>Simplified XGB classifier (Wang et al., <xref ref-type="bibr" rid="B30">2020</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1230649-g0012.tif"/>
</fig>
</sec>
<sec>
<title>5.1.2. Random forest</title>
<p>Random Forest model is an ensemble learning method that combines the predictions of multiple decision trees. It is commonly used for both classification and regression tasks in various domains. Random Forests are known for their ability to handle high-dimensional data, reduce overfitting, and provide robust predictions. It randomly selects subsets of features and training data to build each tree independently. The predictions of the individual trees are then aggregated to make the final prediction.</p>
<p><xref ref-type="fig" rid="F13">Figure 13</xref> shows how a random forest model works.</p>
<fig id="F13" position="float">
<label>Figure 13</label>
<caption><p>Simplified random forest model (Wikipedia, <xref ref-type="bibr" rid="B31">2023</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1230649-g0013.tif"/>
</fig>
</sec>
<sec>
<title>5.1.3. Logistic regression</title>
<p>Logistic regression is a machine learning algorithm specifically designed for predicting categorical dependent variables with binary outcomes, such as yes or no, true or false, or 0 or 1. It models the relationship between the input features and the probability of the binary outcome using a logistic or sigmoid function. By estimating coefficients through training, the algorithm maximizes the likelihood of the observed data. The predicted probabilities can be transformed into binary predictions using a threshold value. Logistic regression is favored for its simplicity and interpretability, although it assumes a linear relationship between the features and may have limitations in complex scenarios. The model&#x00027;s performance is commonly evaluated using metrics like accuracy and precision. See the depiction of the Logistic Regression model in <xref ref-type="fig" rid="F14">Figure 14</xref>.</p>
<fig id="F14" position="float">
<label>Figure 14</label>
<caption><p>Logistic regression model (Torres et al., <xref ref-type="bibr" rid="B29">2019</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1230649-g0014.tif"/>
</fig>
</sec>
<sec>
<title>5.1.4. Support vector machine</title>
<p>Support Vector Machines (SVM) are machine learning models that are versatile and used for classification and regression tasks. They aim to find an optimal decision boundary that maximizes the margin between classes. SVM utilizes the kernel trick to handle non-linear data, and support vectors are crucial in defining the decision boundary. The model is trained by optimizing the boundary and balancing regularization parameters. SVM is effective in handling high-dimensional data and complex decision boundaries.</p>
<p><xref ref-type="fig" rid="F15">Figure 15</xref> is a depiction of SVM.</p>
<fig id="F15" position="float">
<label>Figure 15</label>
<caption><p>Support vector machine model (JavaTpoint, <xref ref-type="bibr" rid="B10">2011&#x02013;2021</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-06-1230649-g0015.tif"/>
</fig>
</sec>
</sec>
<sec>
<title>5.2. Results obtained from the models</title>
<p><xref ref-type="table" rid="T1">Table 1</xref> showcases the results of different models, namely XGB Classifier, Random Forest, Logistic Regression, and SVM/SVC. The table presents their performance based on accuracy scores and computation time in seconds. The accuracy scores range from 95.2 to 96.3%, while the computation time varies significantly across the models, with values ranging from 0.29 to 1,072.32 s. These results provide insights into the models&#x00027; predictive accuracy and computational efficiency, serving as a basis for further analysis and comparison.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Model results.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Rating criteria</bold></th>
<th valign="top" align="center"><bold>XGB classifier</bold></th>
<th valign="top" align="center"><bold>Random forest</bold></th>
<th valign="top" align="center"><bold>Logistic regression</bold></th>
<th valign="top" align="center"><bold>SVM</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Accuracy</td>
<td valign="top" align="center">96.1%</td>
<td valign="top" align="center">95.2%</td>
<td valign="top" align="center">96.3%</td>
<td valign="top" align="center">96.2%</td>
</tr>
<tr>
<td valign="top" align="left">Computation time (s)</td>
<td valign="top" align="center">6.75</td>
<td valign="top" align="center">1,072.32</td>
<td valign="top" align="center">0.29</td>
<td valign="top" align="center">29.92</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The individual results of these models are presented in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<p>The results presented suggest that the SVM model and Logistic Regression model produced the most accurate results, with Logistic Regression slightly outperforming the SVM model, while the logistic regression model computed in the shortest amount of time. A more detailed analysis of the results suggests that the accuracy of the results was relatively similar for all four models. The lowest of the four models was the Random Forest Model with an accuracy of 95.2%, surprising as this is an Ensemble method, and it also had the longest computational time of 1,072.32 s. The second was the XGB Classifier with an accuracy of 96.1% and a computational time of 6.75 s. The third is the SVM Model with an accuracy of 96.2% and a computational time of 29.92 s and again, the fastest was the Logistic Regression model with an accuracy of 96.3% and a computational time of 0.29 s.</p>
</sec>
</sec>
<sec id="s6">
<title>6. Discussion of results</title>
<p>The choice of model in the analysis process depends on the specific objectives of the study. In this case, the goal is to identify signs of depression in tweets. The effectiveness of the models can be evaluated based on two key factors: speed and accuracy. If the primary objective is to identify depression in tweets at a fast pace, models with faster computation times would be more suitable. These models may sacrifice some accuracy for speed. On the other hand, if accuracy is of utmost importance, models with higher accuracy rates should be prioritized, even if they have longer computation times. Considering the focus of this research on identifying depression signs in real-time as tweets come in, it is crucial to have a model that can classify them quickly and accurately. Therefore, it is necessary to analyze both the computational time and accuracy of the models to make an informed decision.</p>
<p>Based on the results shown in <xref ref-type="table" rid="T1">Table 1</xref>, the Logistic Regression model stands out as the most effective option. It achieved the highest accuracy rate of 96.3% while maintaining a relatively low computational time of 0.29 s. This combination of high accuracy and fast computation makes it a strong contender for solving the depression identification problem in real-time tweet analysis. Looking at this, this paper can clearly state the Logistic Regression model emerges as the most suitable choice. It balances both accuracy and computational time, making it an effective tool for identifying signs of depression in tweets.</p>
<sec>
<title>6.1. Comparison with existing studies</title>
<p>When comparing the results with previous studies, several insights emerge. Previous studies, however, suggest that ensemble methods should be more effective in sentiment analysis. In the comparison of the accuracy scores presented in <xref ref-type="table" rid="T2">Table 2</xref>, the random forest classifier&#x00027;s results are quite alarming, considering the higher expectations for ensemble methods. Jain et al. (<xref ref-type="bibr" rid="B9">2022</xref>) conducted a similar study and confirmed the effectiveness of the SVM classifier, which outperformed logistic regression and random forest in three out of the represented categories. This aligns with the findings of Jianqiang and Xiaolin (<xref ref-type="bibr" rid="B11">2017</xref>), who also highlighted the superior performance of SVM compared to logistic regression and random forest in their study.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Presents a comparison of the accuracy scores of the four models with previous studies, highlighting their performance.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>ML models</bold></th>
<th valign="top" align="center"><bold>Jain et al., <xref ref-type="bibr" rid="B9">2022</xref></bold></th>
<th valign="top" align="center"><bold>Aliman et al., <xref ref-type="bibr" rid="B1">2022</xref></bold></th>
<th valign="top" align="center"><bold>Dave, <xref ref-type="bibr" rid="B5">2023</xref></bold></th>
<th valign="top" align="center"><bold>Sujithra et al., <xref ref-type="bibr" rid="B26">2023</xref></bold></th>
<th valign="top" align="center"><bold>Aljabri et al., <xref ref-type="bibr" rid="B2">2022</xref></bold></th>
<th valign="top" align="center"><bold>This study</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Logistic regression</td>
<td valign="top" align="center">79% (highest)</td>
<td valign="top" align="center">81% (highest)</td>
<td valign="top" align="center">83.62%</td>
<td valign="top" align="center">74.78%</td>
<td valign="top" align="center">N/A</td>
<td valign="top" align="center">96.3% (highest)</td>
</tr> <tr>
<td valign="top" align="left">SVM/SVC</td>
<td valign="top" align="center">77.12%</td>
<td valign="top" align="center">69%</td>
<td valign="top" align="center">86.95%</td>
<td valign="top" align="center">N/A</td>
<td valign="top" align="center">88%</td>
<td valign="top" align="center">96.2%</td>
</tr> <tr>
<td valign="top" align="left">XGB classifier</td>
<td valign="top" align="center">N/A</td>
<td valign="top" align="center">N/A</td>
<td valign="top" align="center">86.76%</td>
<td valign="top" align="center">74.22%</td>
<td valign="top" align="center">90% (highest)</td>
<td valign="top" align="center">96.1%</td>
</tr>
<tr>
<td valign="top" align="left">Random forest</td>
<td valign="top" align="center">77.298%</td>
<td valign="top" align="center">N/A</td>
<td valign="top" align="center">88.38% (highest)</td>
<td valign="top" align="center">75.12% (highest)</td>
<td valign="top" align="center">78%</td>
<td valign="top" align="center">95.2%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Interestingly, Dave (<xref ref-type="bibr" rid="B5">2023</xref>) reported a relatively higher accuracy score for logistic regression compared to other models, reaching 83.62%. In addition, this study presents an even higher accuracy score for logistic regression, at 96.3%, indicating its effectiveness in accurately predicting the presence of depression sentiment in tweets. Hence the results obtained in this study indicate that logistic regression emerged as the most suitable model for the paper.</p>
<p>The SVM model, although not consistently outperforming the other models across previous studies, still demonstrates competitive accuracy scores. For instance, in the study by Aljabri et al. (<xref ref-type="bibr" rid="B2">2022</xref>), SVM achieved an accuracy of 88%. Similarly, in the current study, SVM/SVC performed well with an accuracy score of 96.2%.</p>
<p>The XGB Classifier achieved an accuracy of 96.1% in this study, indicating its strong performance in detecting depression sentiment in tweets. Compared to other models in the study, the XGB Classifier had the third highest accuracy score. Additionally, one previous study by Aljabri et al. (<xref ref-type="bibr" rid="B2">2022</xref>) reported a high accuracy of 90% for the XGB Classifier, further highlighting its effectiveness. The XGB Classifier&#x00027;s ability to capture complex patterns and interactions in the data likely contributed to its successful performance. Overall, the XGB Classifier shows promise as a reliable model for depression detection in tweets.</p>
<p>On the other hand, the random forest model presents mixed results. While it achieved the highest accuracy score in the study by Aliman et al. (<xref ref-type="bibr" rid="B1">2022</xref>) at 88.38%, it obtained a relatively lower score in the current study, with 95.2%. These variations could be attributed to different datasets or other factors.</p>
<p>Overall, considering the consistently high accuracy scores and the specific requirements of the paper, logistic regression emerged as the best model choice. However, the inclusion of other models such as SVM and XGB Classifier allows for a comprehensive comparison and exploration of their performance in sentiment analysis.</p>
</sec>
</sec>
<sec id="s7">
<title>7. Conclusion</title>
<p>The paper aimed to identify depression using user tweets more reliably early. As a result, this research proposed a tool based on four classifiers, NLP, and sentiment analysis techniques to improve performance in the early detection of depression. A series of experiments were carried out to evaluate the accuracy and efficacy of the four classification models (XGBClassifier, Random Forest, Logistic Regression, and SVM) that were used on the four datasets combined as one. The results show that the Logistic Regression and SVM models were the most accurate, with Logistic Regression outperforming the SVM model slightly. However, the Logistic regression model was the fastest in terms of computational time of the depressive tweets. Future research should investigate ways to reduce computational time, while also improving model accuracy during the predictive process. Furthermore, in the extension of this work, we are interested in testing the model on new datasets to detect depression.</p>
</sec>
<sec sec-type="data-availability" id="s8">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec sec-type="author-contributions" id="s9">
<title>Author contributions</title>
<p>IO, SD, and OC: study conception and design, analysis and interpretation of results, and draft manuscript preparation. SD and OC: data collection. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
</body>
<back>
<ack><p>The authors gladly recognize the infrastructure support offered by Sol Plaatje University for this study.</p>
</ack>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aliman</surname> <given-names>G.</given-names></name> <name><surname>Nivera</surname> <given-names>T.</given-names></name> <name><surname>Olazo</surname> <given-names>J.</given-names></name> <name><surname>Ramos</surname> <given-names>D. J.</given-names></name> <name><surname>Sanchez</surname> <given-names>C.</given-names></name> <name><surname>Amado</surname> <given-names>T.</given-names></name> <name><surname>Valenzuela</surname> <given-names>I. C.</given-names></name></person-group> (<year>2022</year>). <article-title>Sentiment analysis using logistic regression</article-title>. <source>J. Comp. Innovat. Eng. Appl.</source> <fpage>35</fpage>&#x02013;<lpage>40</lpage>.</citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aljabri</surname> <given-names>M.</given-names></name> <name><surname>Aljameel</surname> <given-names>S. S.</given-names></name> <name><surname>Khan</surname> <given-names>I. U.</given-names></name> <name><surname>Aslam</surname> <given-names>N.</given-names></name> <name><surname>Charouf</surname> <given-names>S. M.</given-names></name> <name><surname>Alzahrani</surname> <given-names>N.</given-names></name></person-group> (<year>2022</year>). <article-title>Machine learning model for sentiment analysis of COVID-19 tweets</article-title>. <source>Int. J. Adv. Sci. Eng. Inf. Technol.</source> <volume>12</volume>, <fpage>1206</fpage>&#x02013;<lpage>1214</lpage>. <pub-id pub-id-type="doi">10.18517/ijaseit.12.3.14724</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>AlSagri</surname> <given-names>H. S.</given-names></name> <name><surname>Ykhlef</surname> <given-names>M.</given-names></name></person-group> (<year>2020</year>). <article-title>Machine learning-based approach for depression detection in twitter using content and activity features</article-title>. <source>IEICE Transact. Inf. Syst.</source> <volume>103</volume>, <fpage>1825</fpage>&#x02013;<lpage>1832</lpage>. <pub-id pub-id-type="doi">10.1587/transinf.2020EDP7023</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Babu</surname> <given-names>N. V.</given-names></name> <name><surname>Kanaga</surname> <given-names>E. G. M.</given-names></name></person-group> (<year>2022</year>). <article-title>Sentiment analysis in social media data for depression detection using artificial intelligence: a review</article-title>. <source>SN Comp. Sci.</source> <volume>3</volume>, <fpage>1</fpage>&#x02013;<lpage>20</lpage>. <pub-id pub-id-type="doi">10.1007/s42979-021-00958-1</pub-id><pub-id pub-id-type="pmid">34816124</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Dave</surname> <given-names>G. H.</given-names></name></person-group> (<year>2023</year>). <article-title>Leveraging big data for early detection of depression: developing a machine learning model using tweets</article-title>. <source>Vidhyayana Int. Multidiscipl. Peer Rev. Eur J.</source> <volume>8</volume>, <fpage>777</fpage>&#x02013;<lpage>784</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://vidhyayanaejournal.org/journal/article/view/784">https://vidhyayanaejournal.org/journal/article/view/784</ext-link></citation>
</ref>
<ref id="B6">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Draelos</surname> <given-names>R.</given-names></name></person-group> (<year>2019</year>). <source>GLASS BOX - Machine Learning and Medicine</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://glassboxmedicine.com/2019/02/17/measuring-performance-the-confusion-matrix/">https://glassboxmedicine.com/2019/02/17/measuring-performance-the-confusion-matrix/</ext-link> (accessed October 30, 2022).</citation>
</ref>
<ref id="B7">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Gaikwad</surname> <given-names>S.</given-names></name> <name><surname>Khairnar</surname> <given-names>U.</given-names></name> <name><surname>Deshpande</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Text mining process: Techniques and tools</article-title>. <source>Int. J. Adv. Res. Comput. Engg. Technol</source>. <volume>3</volume>, <fpage>413</fpage>&#x02013;<lpage>418</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://csjournals.com/IJITKM/PDF%203-1/86.pdf">http://csjournals.com/IJITKM/PDF%203-1/86.pdf</ext-link></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jagadishwari</surname> <given-names>V.</given-names></name> <name><surname>Indulekha</surname> <given-names>A.</given-names></name> <name><surname>Kiran</surname> <given-names>R.aghu, Harshini, P.</given-names></name></person-group> (<year>2021</year>). <article-title>Sentiment analysis of social media text-emoticon post with machine learning models contribution title</article-title>. <source>J. Phys.</source> <volume>2070</volume>, <fpage>012079</fpage>. <pub-id pub-id-type="doi">10.1088/1742-6596/2070/1/012079</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jain</surname> <given-names>P.</given-names></name> <name><surname>Srinivas</surname> <given-names>K.</given-names></name> <name><surname>Vichare</surname> <given-names>A.</given-names></name></person-group> (<year>2022</year>). <article-title>Depression and suicide analysis using machine learning and NLP</article-title>. <source>J. Phys.</source> <volume>2161</volume>, <fpage>012034</fpage>. <pub-id pub-id-type="doi">10.1088/1742-6596/2161/1/012034</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="web"><person-group person-group-type="author"><collab>JavaTpoint</collab></person-group> (<year>2011&#x02013;2021</year>). <source>JavaTpoint</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.javatpoint.com/machine-learning-support-vector-machine-algorithm">https://www.javatpoint.com/machine-learning-support-vector-machine-algorithm</ext-link> (accessed June 14, 2023).</citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jianqiang</surname> <given-names>Z.</given-names></name> <name><surname>Xiaolin</surname> <given-names>G.</given-names></name></person-group> (<year>2017</year>). <article-title>Comparison research on text pre-processing methods on twitter sentiment analysis</article-title>. <source>IEEE Access</source> <volume>5</volume>, <fpage>2870</fpage>&#x02013;<lpage>2879</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2017.2672677</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kumar</surname> <given-names>P.</given-names></name> <name><surname>Garg</surname> <given-names>S.</given-names></name> <name><surname>Garg</surname> <given-names>A.</given-names></name></person-group> (<year>2020</year>). <article-title>Assessment of anxiety, depression and stress using learning models</article-title>. <source>Proced. Comput. Sci</source>. <volume>171</volume>, <fpage>1989</fpage>&#x02013;<lpage>1998</lpage>. <pub-id pub-id-type="doi">10.1016/j.procs.2020.04.213</pub-id><pub-id pub-id-type="pmid">35735574</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lui</surname> <given-names>C.</given-names></name></person-group> (<year>2020</year>). <article-title>Social media and public opinion during COVID-19 pandemic: A cross-country analysis</article-title>. <source>Comput. Hum. Behav</source>. 110, 106380. <pub-id pub-id-type="doi">10.1016/j.chb.2020.106380</pub-id><pub-id pub-id-type="pmid">32292239</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mun</surname> <given-names>B.</given-names></name> <name><surname>Kim</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>Influence of false self-presentation on mental health and deleting behavior on instagram: the mediating role of perceived popularity</article-title>. <source>Front. Psychol</source>. 12, 660484. <pub-id pub-id-type="doi">10.3389/fpsyg.2021.660484</pub-id><pub-id pub-id-type="pmid">33912119</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Orabi</surname> <given-names>A.</given-names></name> <name><surname>Buddhitha</surname> <given-names>P.</given-names></name> <name><surname>Orabi</surname> <given-names>M.</given-names></name> <name><surname>Inkpen</surname> <given-names>D.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Deep learning for depression detection of twitter users,&#x0201D;</article-title> in <source>Proceedings of the Fifth Workshop on Computational Linguistics and Clinical Psychology: From Keyboard to Clinic</source> (<publisher-loc>New Orleans, LA</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>88</fpage>&#x02013;<lpage>97</lpage>.</citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Orsolini</surname> <given-names>L.</given-names></name> <name><surname>Pompili</surname> <given-names>S.</given-names></name> <name><surname>Salvi</surname> <given-names>V.</given-names></name> <name><surname>Volpe</surname> <given-names>U.</given-names></name></person-group> (<year>2021</year>). <article-title>A systematic review on telemental health in youth mental health: Focus on anxiety, depression and obsessive-compulsive disorder (Medicina: MDPI).</article-title> 57, 793. <pub-id pub-id-type="doi">10.3390/medicina57080793</pub-id><pub-id pub-id-type="pmid">34440999</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Priya</surname> <given-names>A.</given-names></name> <name><surname>Garg</surname> <given-names>S.</given-names></name> <name><surname>Tigga</surname> <given-names>N.</given-names></name></person-group> (<year>2020</year>). <article-title>Predicting anxiety, depression, and stress in modern life using machine learning algorithms</article-title>. <source>Proced. Comput. Sci.</source> <volume>167</volume>, <fpage>1258</fpage>&#x02013;<lpage>1267</lpage>. <pub-id pub-id-type="doi">10.1016/j.procs.2020.03.442</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ranganathan</surname> <given-names>J.</given-names></name> <name><surname>Tzacheva</surname> <given-names>A.</given-names></name></person-group> (<year>2019</year>). <article-title>Emotion mining in social media data</article-title>. <source>Proc. Comp. Sci.</source> <volume>159</volume>, <fpage>58</fpage>&#x02013;<lpage>66</lpage>. <pub-id pub-id-type="doi">10.1016/j.procs.2019.09.160</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ricard</surname> <given-names>B.</given-names></name> <name><surname>Marsch</surname> <given-names>L.</given-names></name> <name><surname>Crosier</surname> <given-names>B.</given-names></name> <name><surname>Hassanpour</surname> <given-names>S.</given-names></name></person-group> (<year>2018</year>). <article-title>Exploring the utility of community-generated social media content for detecting depression: an analytical study on instagram</article-title>. <source>J. Med. Int. Res.</source> <volume>20</volume>, <fpage>e118</fpage>. <pub-id pub-id-type="doi">10.2196/11817</pub-id><pub-id pub-id-type="pmid">30522991</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Samsari</surname> <given-names>N. S.</given-names></name> <name><surname>Mohamad</surname> <given-names>M.</given-names></name> <name><surname>Selamati</surname> <given-names>A.</given-names></name></person-group> (<year>2022</year>). <article-title>Sentiment analysis on students&#x00027; stress and depression due to online distance learning during the COVID-19 pandemic</article-title>. <source>Math. Sci. Inf. J</source>. <volume>3</volume>, <fpage>66</fpage>&#x02013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.24191/mij.v3i1.18273</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Seabrook</surname> <given-names>E.</given-names></name> <name><surname>Kern</surname> <given-names>M. L.</given-names></name> <name><surname>Fulcher</surname> <given-names>B. D.</given-names></name> <name><surname>Rickard</surname> <given-names>N. S.</given-names></name></person-group> (<year>2018</year>). <article-title>Predicting depression from language-based emotion dynamics: a longitudinal analysis of Facebook and Twitter status updates</article-title>. <source>J. Med. Int. Res.</source> <volume>20</volume>, <fpage>e168</fpage>. <pub-id pub-id-type="doi">10.2196/jmir.9267</pub-id><pub-id pub-id-type="pmid">29739736</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Seney</surname> <given-names>M.</given-names></name> <name><surname>Zhiguang</surname> <given-names>H.</given-names></name> <name><surname>Cahill</surname> <given-names>K. L. F.</given-names></name> <name><surname>Puralewski</surname> <given-names>R.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Opposite molecular signatures of depression in men and women</article-title>. <source>Biol. Psychiatry.</source> <volume>84</volume>, <fpage>8</fpage>&#x02013;<lpage>27</lpage>. <pub-id pub-id-type="doi">10.1016/j.biopsych.2018.01.017</pub-id><pub-id pub-id-type="pmid">29548746</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shukla</surname> <given-names>V. S.</given-names></name> <name><surname>Vala</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>A survey on image mining, its techniques, and application</article-title>. <source>Int. J. Comp. Appl.</source> <volume>133</volume>, <fpage>12</fpage>&#x02013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.5120/ijca2016907978</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Smith</surname> <given-names>R.</given-names></name> <name><surname>Anderson</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>Instagram photos reveal predictive markers of depression</article-title>. <source>EPJ Data Sci</source>. 7, 15. <pub-id pub-id-type="doi">10.1140/epjds/s13688-018-0140-6</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sood</surname> <given-names>A.</given-names></name> <name><surname>Hooda</surname> <given-names>M.</given-names></name> <name><surname>Dhir</surname> <given-names>S.</given-names></name> <name><surname>Bhatia</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>An Initiative To Identify Depression Using Sentiment Analysis: A Machine Learning Approach</article-title>. <source>Indian J. Sci. Technol.</source> <volume>11</volume>, <fpage>1</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.17485/ijst/2018/v11i4/119594</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sujithra</surname> <given-names>M.</given-names></name> <name><surname>Rathika</surname> <given-names>J.</given-names></name> <name><surname>Velvadivu</surname> <given-names>P.</given-names></name> <name><surname>Marimuthu</surname> <given-names>M.</given-names></name></person-group> (<year>2023</year>). <article-title>An intellectual decision system for classification of mental health illness on social media using computational intelligence approach</article-title>. <source>J. Ubiquit. Comp. Commun. Technol.</source> <volume>5</volume>, <fpage>23</fpage>&#x02013;<lpage>35</lpage>. <pub-id pub-id-type="doi">10.36548/jucct.2023.1.002</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="web"><person-group person-group-type="author"><collab>The University of Queensland</collab></person-group> (<year>2022</year>). <source>University of Queensland</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://qbi.uq.edu.au/brain/brain-diseases/depression">https://qbi.uq.edu.au/brain/brain-diseases/depression</ext-link> (accessed August 04, 2022).</citation>
</ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tiwari</surname> <given-names>A.</given-names></name> <name><surname>Mishra</surname> <given-names>A.</given-names></name> <name><surname>Rath</surname> <given-names>S. K.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Emotion mining in social media data,&#x0201D;</article-title> in <source>Intelligent Computing Techniques for Cyber Security</source> (<publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>227</fpage>&#x02013;<lpage>238</lpage>.</citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Torres</surname> <given-names>R.</given-names></name> <name><surname>Ohashi</surname> <given-names>O.</given-names></name> <name><surname>Pessin</surname> <given-names>G.</given-names></name></person-group> (<year>2019</year>). <article-title>A machine-learning approach to distinguish passengers and drivers reading while driving</article-title>. <source>Sensors (Basel)</source>. <volume>19</volume>:<fpage>3174</fpage>. <pub-id pub-id-type="doi">10.3390/s19143174</pub-id><pub-id pub-id-type="pmid">31330929</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Chakraborty</surname> <given-names>G.</given-names></name> <name><surname>Chakraborty</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). <source>Predicting the Risk of Chronic Kidney Disease (CKD) Using Machine Learning Algorithm. ResearchGate</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.researchgate.net/figure/Simplified-structure-of-XGBoost_fig2_348025909">https://www.researchgate.net/figure/Simplified-structure-of-XGBoost_fig2_348025909</ext-link> (accessed June 14, 2023).</citation>
</ref>
<ref id="B31">
<citation citation-type="web"><person-group person-group-type="author"><collab>Wikipedia</collab></person-group> (<year>2023</year>). <source>Wikipedia Organisation.</source> Available online at: <ext-link ext-link-type="uri" xlink:href="https://en.wikipedia.org/wiki/Random_forest">https://en.wikipedia.org/wiki/Random_forest</ext-link> (accessed June 14, 2023).</citation>
</ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zulfiker</surname> <given-names>M.</given-names></name> <name><surname>Kabir</surname> <given-names>N.</given-names></name> <name><surname>Biswas</surname> <given-names>A.</given-names></name> <name><surname>Nazneen</surname> <given-names>T.</given-names></name> <name><surname>Uddin</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <source>An In-Depth Analysis of Machine Learning Approaches to Predict Depression</source>. <publisher-name>Elsevier</publisher-name>, <fpage>38</fpage>&#x02013;<lpage>50</lpage>.<pub-id pub-id-type="pmid">35749157</pub-id></citation></ref>
</ref-list> 
</back>
</article> 