<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frai.2021.783778</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Augmenting Semantic Lexicons Using Word Embeddings and Transfer Learning</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Alshaabi</surname> <given-names>Thayer</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1494075/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Van Oort</surname> <given-names>Colin M.</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Fudolig</surname> <given-names>Mikaela Irene</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1520204/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Arnold</surname> <given-names>Michael V.</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Danforth</surname> <given-names>Christopher M.</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1011454/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Dodds</surname> <given-names>Peter Sheridan</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="aff" rid="aff5"><sup>5</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Advanced Bioimaging Center, University of California, Berkeley</institution>, <addr-line>Berkeley, CA</addr-line>, <country>United States</country></aff>
<aff id="aff2"><sup>2</sup><institution>Vermont Complex Systems Center, University of Vermont</institution>, <addr-line>Burlington, VT</addr-line>, <country>United States</country></aff>
<aff id="aff3"><sup>3</sup><institution>The MITRE Corporation</institution>, <addr-line>McLean, VA</addr-line>, <country>United States</country></aff>
<aff id="aff4"><sup>4</sup><institution>Department of Mathematics &#x00026; Statistics, University of Vermont</institution>, <addr-line>Burlington, VT</addr-line>, <country>United States</country></aff>
<aff id="aff5"><sup>5</sup><institution>Department of Computer Science, University of Vermont</institution>, <addr-line>Burlington, VT</addr-line>, <country>United States</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Preslav Nakov, Qatar Computing Research Institute, Qatar</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Svetlana Kiritchenko, National Research Council Canada (NRC-CNRC), Canada; Erik Cambria, Nanyang Technological University, Singapore</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Thayer Alshaabi <email>thayeralshaabi&#x00040;berkeley.edu</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Language and Computation, a section of the journal Frontiers in Artificial Intelligence</p></fn></author-notes>
<pub-date pub-type="epub">
<day>24</day>
<month>01</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>4</volume>
<elocation-id>783778</elocation-id>
<history>
<date date-type="received">
<day>26</day>
<month>09</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>12</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Alshaabi, Van Oort, Fudolig, Arnold, Danforth and Dodds.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Alshaabi, Van Oort, Fudolig, Arnold, Danforth and Dodds</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> 
</permissions>
<abstract>
<p>Sentiment-aware intelligent systems are essential to a wide array of applications. These systems are driven by language models which broadly fall into two paradigms: Lexicon-based and contextual. Although recent contextual models are increasingly dominant, we still see demand for lexicon-based models because of their interpretability and ease of use. For example, lexicon-based models allow researchers to readily determine which words and phrases contribute most to a change in measured sentiment. A challenge for any lexicon-based approach is that the lexicon needs to be routinely expanded with new words and expressions. Here, we propose two models for automatic lexicon expansion. Our first model establishes a baseline employing a simple and shallow neural network initialized with pre-trained word embeddings using a non-contextual approach. Our second model improves upon our baseline, featuring a deep Transformer-based network that brings to bear word definitions to estimate their lexical polarity. Our evaluation shows that both models are able to score new words with a similar accuracy to reviewers from Amazon Mechanical Turk, but at a fraction of the cost.</p></abstract>
<kwd-group>
<kwd>sentiment analysis</kwd>
<kwd>semantic lexicons</kwd>
<kwd>transformers</kwd>
<kwd>BERT</kwd>
<kwd>FastText</kwd>
<kwd>word embedding</kwd>
<kwd>labMT</kwd>
</kwd-group>
<contract-num rid="cn001">OAC-1827314</contract-num>
<contract-sponsor id="cn001">National Science Foundation<named-content content-type="fundref-id">10.13039/100000001</named-content></contract-sponsor>
<contract-sponsor id="cn002">MassMutual Financial Group<named-content content-type="fundref-id">10.13039/100004769</named-content></contract-sponsor>
<contract-sponsor id="cn003">Google<named-content content-type="fundref-id">10.13039/100006785</named-content></contract-sponsor>
<counts>
<fig-count count="8"/>
<table-count count="1"/>
<equation-count count="0"/>
<ref-count count="117"/>
<page-count count="16"/>
<word-count count="11323"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>In computational linguistics and natural language processing (NLP), sentiment analysis involves extracting emotion and opinion from text data. There is an increasing demand for sentiment-aware intelligent systems. The growth of sentiment-aware frameworks in online services can be seen across a vast, multidisciplinary set of applications (Nasukawa and Yi, <xref ref-type="bibr" rid="B72">2003</xref>; Medhat et al., <xref ref-type="bibr" rid="B67">2014</xref>; Bakshi et al., <xref ref-type="bibr" rid="B11">2016</xref>).</p>
<p>With the modern volume of text data&#x02014;which has long rendered human annotation infeasible&#x02014;automated sentiment analysis is used, for example, by businesses in evaluating customer feedback to make informed decisions regarding product development and risk management (Turney, <xref ref-type="bibr" rid="B105">2002</xref>; Cabral and Hortacsu, <xref ref-type="bibr" rid="B18">2010</xref>). Combined with recommender systems, sentiment analysis has also been used with the intent to improve consumer experience through aggregated and curated feedback from other consumers, particularly in retail (Kumar and Lee, <xref ref-type="bibr" rid="B58">2006</xref>; Tang et al., <xref ref-type="bibr" rid="B97">2009</xref>; Yu et al., <xref ref-type="bibr" rid="B115">2013</xref>), e-commerce (Bhatt et al., <xref ref-type="bibr" rid="B15">2015</xref>; Haque et al., <xref ref-type="bibr" rid="B42">2018</xref>), and entertainment (Terveen et al., <xref ref-type="bibr" rid="B99">1997</xref>; Pang et al., <xref ref-type="bibr" rid="B76">2002</xref>).</p>
<p>Beyond applications in industry, sentiment analysis has been widely applied in academic research, particularly in the social and political sciences (Chen et al., <xref ref-type="bibr" rid="B20">2021</xref>). Public opinion, e.g., support for or opposition to policies, can be potentially gauged from online political discourse, giving policymakers an important window into public awareness and attitude (Laver et al., <xref ref-type="bibr" rid="B60">2003</xref>; Thomas et al., <xref ref-type="bibr" rid="B102">2006</xref>). Sentiment analysis tools have shown mixed results in forecasting elections (Tumasjan et al., <xref ref-type="bibr" rid="B104">2010</xref>) and monitoring inflammatory discourse on social media, with vital relevance to national security (Pang and Lee, <xref ref-type="bibr" rid="B75">2008</xref>). Sentiment analysis has also been used in the public health domain (Coppersmith et al., <xref ref-type="bibr" rid="B24">2014</xref>; Yadollahi et al., <xref ref-type="bibr" rid="B112">2017</xref>; Gohil et al., <xref ref-type="bibr" rid="B39">2018</xref>), with recent studies analyzing social media discourse surrounding mental health (Bathina et al., <xref ref-type="bibr" rid="B12">2021</xref>; Stupinski et al., <xref ref-type="bibr" rid="B94">2021</xref>), disaster response and emergency management (Beigi et al., <xref ref-type="bibr" rid="B13">2016</xref>).</p>
<p>The growing number of applications of sentiment-aware systems has led the NLP community in the past decade to develop end-to-end models to examine short- and medium-length text documents (Wilson et al., <xref ref-type="bibr" rid="B109">2005</xref>; Feldman, <xref ref-type="bibr" rid="B35">2013</xref>), particularly for social media (Pak and Paroubek, <xref ref-type="bibr" rid="B74">2010</xref>; Agarwal et al., <xref ref-type="bibr" rid="B2">2011</xref>; Korkontzelos et al., <xref ref-type="bibr" rid="B55">2016</xref>). Some researchers have considered the many social and political implications of using AI for sentiment detection across media (Crawford, <xref ref-type="bibr" rid="B25">2019</xref>; Crawford and Paglen, <xref ref-type="bibr" rid="B26">2021</xref>). Recent studies highlight some of the implicit hazards of crowdsourcing text data (Shmueli et al., <xref ref-type="bibr" rid="B89">2021</xref>), especially in light of the latest advances in NLP and emerging ethical concerns (Conway and O&#x00027;Connor, <xref ref-type="bibr" rid="B23">2016</xref>; Hovy and Spruit, <xref ref-type="bibr" rid="B49">2016</xref>). Identifying potential racial and gender disparity in NLP models is essential to develop better models (Tatman, <xref ref-type="bibr" rid="B98">2017</xref>).</p>
<p>Sentiment analysis tools fall into one of two groups, depending on their definition of sentiment and their model for its estimation. One of the more popular paradigms is discrete classification, where sentiment is divided into several classes (e.g., positive, negative) and pieces of text are associated with each class. However, sometimes a continuous measure is desired, requiring a spectrum of sentiment scores rather than sentiment classes (Thelwall et al., <xref ref-type="bibr" rid="B101">2010</xref>). This more nuanced sentiment scoring paradigm has been widely adopted for e-commerce, movies, and restaurant reviews (Snyder and Barzilay, <xref ref-type="bibr" rid="B90">2007</xref>).</p>
<p>Sentiment analysis models largely derive from two major paradigms: 1. Lexicon-based models and 2. Contextual models. Lexicon-based models compute sentiment scores based on sentiment dictionaries (sentiment lexicons) typically constructed by human annotators (Taboada et al., <xref ref-type="bibr" rid="B95">2011</xref>; Dodds et al., <xref ref-type="bibr" rid="B30">2015</xref>; Augustyniak et al., <xref ref-type="bibr" rid="B8">2016</xref>). A sentiment lexicon contains not only terms that express a particular sentiment/emotion, but also terms that are associated with a particular sentiment/emotion (denotation vs. connotation). Contextual models, on the other hand, extrapolate semantics by converting words to vectors in an embedding space, and learning from large-scale annotated datasets to predict sentiment based on co-occurrence relationships between words (Wilson et al., <xref ref-type="bibr" rid="B109">2005</xref>; Pak and Paroubek, <xref ref-type="bibr" rid="B74">2010</xref>; Agarwal et al., <xref ref-type="bibr" rid="B2">2011</xref>; Feldman, <xref ref-type="bibr" rid="B35">2013</xref>; Socher et al., <xref ref-type="bibr" rid="B92">2013b</xref>). Contextual models have the advantage in differentiating multiple meanings, as in the case of &#x0201C;The dog is <italic>lying</italic> on the beach&#x0201D; vs. &#x0201C;I never said that&#x02014;you are <italic>lying</italic>,&#x0201D; while lexicon-based models usually have a single score for each word, regardless of usage. Despite the flexibility of contextual models, their results can be difficult to interpret, as the high-dimensional latent space in which they are embedded renders explanation difficult. The ease of use and transparent comprehension of lexicon-based models help explain their continued popularity (Pang and Lee, <xref ref-type="bibr" rid="B75">2008</xref>; Taboada et al., <xref ref-type="bibr" rid="B95">2011</xref>; Dodds et al., <xref ref-type="bibr" rid="B30">2015</xref>). For example, while the linguistic mechanisms leading to change in sentiment may be hard to explain with word embeddings, one can straightforwardly use lexicon scores to reveal the words contributing to shifted sentiment (Dodds et al., <xref ref-type="bibr" rid="B32">2011</xref>; Reagan et al., <xref ref-type="bibr" rid="B82">2017</xref>; Gallagher et al., <xref ref-type="bibr" rid="B38">2021</xref>).</p>
<p>A major challenge for the simpler and more interpretable lexicon-based models, however, is the time and financial investment associated with maintaining them. Sentiment lexicons must be updated regularly to mitigate the out-of-vocabulary (OOV) problem&#x02014;words and phrases that were either not considered or did not exist when the dictionaries were originally constructed (Riloff, <xref ref-type="bibr" rid="B84">1996</xref>). While researchers show general sentiment trends are observable unless the lexicon does not have enough words, having a versatile dictionary with specialized and rarely used words improves the signal (Dodds and Danforth, <xref ref-type="bibr" rid="B31">2010</xref>; Reagan et al., <xref ref-type="bibr" rid="B82">2017</xref>). Notably, language is an evolving sociotechnical phenomenon. New words and phrases are created constantly, especially on social media (Alshaabi et al., <xref ref-type="bibr" rid="B4">2021a</xref>). Word usage changes over time. New words are created, old words lose popularity, and the meaning of words can change. For example, the word &#x0201C;covid&#x0201D; grew to be the most narratively trending <italic>n</italic>-gram in reference to the global Coronavirus outbreak during February and March 2020 (Alshaabi et al., <xref ref-type="bibr" rid="B5">2021b</xref>).</p>
<p>Sentiment analysis applications are often developed to investigate bipolar relationships (e.g., positive&#x02013;negative, happy&#x02013;sad, excited&#x02013;bored). These bipolar relationships are conveniently handled by binary classification systems, however, such a formalization leads to multiple varieties of neutral sentiment (Colhon et al., <xref ref-type="bibr" rid="B22">2017</xref>). Many sentiment analysis applications avoid, ignore, or remove text with neutral sentiment. Excluding neutral sentiment text during training can have significant impacts on trained models, which are often confused by or uncertain of neutral sentiment text (Koppel and Schler, <xref ref-type="bibr" rid="B54">2006</xref>). For classification-based applications, explicitly representing neutral sentiment as a third class can improve model performance (Ribeiro et al., <xref ref-type="bibr" rid="B83">2016</xref>). Humans process emotionally charged words differently than neutral words, thus sentiment analysis model may find success via similar processes Kissler and Herbert (<xref ref-type="bibr" rid="B53">2013</xref>).</p>
<p>In this work, we propose an automated framework extending sentiment for semantic lexicons to OOV words, reducing the need for crowdsourcing scores from human annotators, a process that can be time-consuming and expensive. Although our framework can be used in a more general sense, we focus on predicting <italic>happiness scores</italic> based on the labMT dataset (Dodds et al., <xref ref-type="bibr" rid="B30">2015</xref>). This dataset was constructed from human ratings of the &#x0201C;happiness&#x0201D; of words on a continuous scale, averaging scores from multiple annotators for more than 10,000 words. We discuss this dataset in detail in section 3.1. In section 2, we discuss recent developments using deep learning in NLP, and how they relate to our work. We introduce two models, demonstrating accuracy on par with human performance (see section 3 for technical details). We first introduce a baseline model&#x02014;a neural network initialized with pre-trained word embeddings&#x02014;to gauge happiness scores. Second, we present a deep Transformer-based model that uses word definitions to estimate their sentiment scores. We will refer to our models as the &#x0201C;Token&#x0201D; and &#x0201C;Dictionary&#x0201D; models, respectively. We present our results and model evaluation in section 4, highlighting how the models perform compared with reviewers from Amazon&#x00027;s Mechanical Turk. Finally, we highlight key limitations of our approach, and outline some potential future developments in concluding remarks.</p>
</sec>
<sec id="s2">
<title>2. Related Work</title>
<p>Word embeddings are abstract numerical representations of the relationships between words, derived from statistics on individual corpora, and encoding language patterns so that concepts with similar semantics have similar representations (Bengio et al., <xref ref-type="bibr" rid="B14">2003</xref>). Researchers have shown that efficient representations of words can both express meanings and preserve context (Maas et al., <xref ref-type="bibr" rid="B65">2011</xref>; Hollis and Westbury, <xref ref-type="bibr" rid="B47">2016</xref>; Hollis et al., <xref ref-type="bibr" rid="B48">2017</xref>; Li et al., <xref ref-type="bibr" rid="B62">2017</xref>). While there are many ways to construct word embedding models (e.g., matrix factorization), we often use the term to refer to a specific class of word embeddings that are learnable via neural networks.</p>
<p>Word2Vec is one of the key breakthroughs in NLP, introducing an efficient way for learning word embeddings from a given text corpus (Mikolov et al., <xref ref-type="bibr" rid="B68">2013a</xref>,<xref ref-type="bibr" rid="B69">b</xref>). At its core, it builds off of a simple idea borrowed from linguistics and formally known as the &#x0201C;distributional hypothesis&#x0201D;&#x02014;words that are semantically similar are also used in similar ways, and likely to appear with similar context words (Harris, <xref ref-type="bibr" rid="B43">1954</xref>).</p>
<p>Starting from a fixed vocabulary, we can learn a vector representation for each word via a shallow network with a single hidden layer trained in one of two fashions (Mikolov et al., <xref ref-type="bibr" rid="B68">2013a</xref>,<xref ref-type="bibr" rid="B69">b</xref>). Both approaches formalize the task as a unsupervised prediction problem, whereby an embedding is learned jointly with a network that is trained to either predict an anchor word given the words around it (i.e., continuous bag-of-words (CBOW)), or by predicting context words for an anchor word (i.e., skip-gram) (Mikolov et al., <xref ref-type="bibr" rid="B68">2013a</xref>). Both approaches, however, are limited to local context bounded by the size of the context window. Global Vectors (GloVe) addresses that problem by capturing corpus global statistics with a word co-occurrence probability matrix (Pennington et al., <xref ref-type="bibr" rid="B77">2014</xref>).</p>
<p>While Word2Vec and GloVe offer substantial improvements over previous methods, they both fail to encode unfamiliar words&#x02014;tokens that were not processed in the training corpora. FastText refines word embeddings by supplementing the learned embedding matrix with subwords to overcome the challenge of OOV tokens (Bojanowski et al., <xref ref-type="bibr" rid="B16">2017</xref>; Joulin et al., <xref ref-type="bibr" rid="B50">2017</xref>). This is achieved by training the network with character-level <italic>n</italic>-grams (<italic>n</italic> &#x02208; {3,4,5,6}), then taking the sum of all subwords to construct a vector representation for any given word. Although the idea behind FastText is rather simple, it presents an elegant solution to account for rare words, allowing the model to learn more general word representations.</p>
<p>A major shortcoming of the earlier models is their inability to capture contextual descriptions of words as they all produce a fixed vector representation for each word. In building context-aware models, researchers often use fundamental building blocks such as recurrent neural networks (RNN) (Rumelhart et al., <xref ref-type="bibr" rid="B85">1986</xref>)&#x02014;particularly long short-term memory (LSTM) (Hochreiter and Schmidhuber, <xref ref-type="bibr" rid="B45">1997</xref>)&#x02014;that are designed to process sequential data. Many methods have provided incremental improvements over time (Chen et al., <xref ref-type="bibr" rid="B21">2017</xref>; Lee et al., <xref ref-type="bibr" rid="B61">2017</xref>; Peters et al., <xref ref-type="bibr" rid="B78">2017</xref>). ELMo is one of the key milestones toward efficient contextualized models, using deep bi-directional LSTM language representations (Peters et al., <xref ref-type="bibr" rid="B79">2018</xref>).</p>
<p>In late 2017, the advent of deep attention-based models, dubbed transformers, rapidly changed the landscape in the NLP community (Vaswani et al., <xref ref-type="bibr" rid="B107">2017</xref>). The encoder-decoder framework, powered by attention blocks, enables faster processing of the input sequence while also preserving context (Vaswani et al., <xref ref-type="bibr" rid="B107">2017</xref>). Recent adaptations of the building blocks of Transformers continue to break records, improving the state-of-the-art across all NLP benchmarks with recent applications to computer vision and pattern recognition (Dosovitskiy et al., <xref ref-type="bibr" rid="B33">2021</xref>).</p>
<p>Exploiting the versatile nature of Transformers, we observe the emergence of a new family of language models widely known as &#x0201C;self-supervised&#x0201D; including as bidirectional encoders (e.g., BERT) (Devlin et al., <xref ref-type="bibr" rid="B29">2019</xref>), and left-to-right decoders (e.g., GPT) (Radford et al., <xref ref-type="bibr" rid="B81">2018</xref>). Self-supervised language models are pre-trained by masking random tokens in the unlabeled input data and training the model to predict these tokens. Researchers leverage recent subword tokenization techniques, such as WordPiece (Wu et al., <xref ref-type="bibr" rid="B111">2016</xref>), SentencePiece (Kudo and Richardson, <xref ref-type="bibr" rid="B57">2018</xref>), and Byte Pair Encoding (BPE) (Sennrich et al., <xref ref-type="bibr" rid="B88">2016</xref>), to overcome the challenge of rare and OOV words. Subtle contextualized representations of words can be learned by predicting whether sentence B follows sentence A (Devlin et al., <xref ref-type="bibr" rid="B29">2019</xref>). Pre-trained language models can then be fine-tuned using labeled data for downstream NLP tasks, such as named entity recognition, question answering, text summarization, and sentiment analysis (Radford et al., <xref ref-type="bibr" rid="B81">2018</xref>; Devlin et al., <xref ref-type="bibr" rid="B29">2019</xref>).</p>
<p>Recent advances in NLP continue to improve the language facility of Transformer-based models. The introduction of XLNet (Yang et al., <xref ref-type="bibr" rid="B114">2019</xref>) is another remarkable breakthrough that combines the bi-directionality of BERT (Devlin et al., <xref ref-type="bibr" rid="B29">2019</xref>) and the autoregressive pre-training scheme from Transformer-XL (Dai et al., <xref ref-type="bibr" rid="B27">2019</xref>). While the current trend of making ever-larger and deeper language models shows an impressive track record, it is arguably unfruitful to maintain unreasonably large models that only giant corporations can afford to use due to hardware limitations (Thompson et al., <xref ref-type="bibr" rid="B103">2020</xref>). Vitally, less expensive language models need to be both computationally efficient and exhibit performance on par with larger models. Addressing that challenge, researchers proposed clever techniques of leveraging knowledge distillation (Hinton et al., <xref ref-type="bibr" rid="B44">2015</xref>) to train smaller and faster models [e.g., DistilBERT (Sanh et al., <xref ref-type="bibr" rid="B87">2019</xref>)]. Similarly, efficient parameterization strategies via sharing weights across layers can also reduce the size of the model while maintaining state-of-the-art results [e.g., ALBERT (Lan et al., <xref ref-type="bibr" rid="B59">2020</xref>)].</p>
<p>Previous work on automatic sentiment lexicon generation (ASLG) has used a variety of heuristics to assign sentiment scores to OOV words. Most ASLG methods start with a seed lexicon containing words of known sentiment, then use a distance function to propagate sentiment scores from known words to unknown words. Word co-occurrence frequencies (Turney and Littman, <xref ref-type="bibr" rid="B106">2003</xref>; Kiritchenko et al., <xref ref-type="bibr" rid="B52">2014</xref>) and shortest path distances within a semantic word graph (Qiu et al., <xref ref-type="bibr" rid="B80">2009</xref>; Baccianella et al., <xref ref-type="bibr" rid="B9">2010</xref>; San Vicente et al., <xref ref-type="bibr" rid="B86">2014</xref>) [such as WordNet (Fellbaum, <xref ref-type="bibr" rid="B36">1998</xref>)] were common distance functions in earlier work. More recently, distance functions based on learned word embeddings have gained popularity (Tang et al., <xref ref-type="bibr" rid="B96">2014</xref>; Wang et al., <xref ref-type="bibr" rid="B108">2016</xref>; Ljube&#x00161;i&#x00107; et al., <xref ref-type="bibr" rid="B64">2018</xref>; Thavareesan and Mahesan, <xref ref-type="bibr" rid="B100">2020</xref>). The outputs of word embedding models usually need to be projected into a lower dimension before they can be used for ASLG. This can be done using a variety of machine learning models, though linear models are likely one of the most popular options (Qiu et al., <xref ref-type="bibr" rid="B80">2009</xref>; Amir et al., <xref ref-type="bibr" rid="B7">2015</xref>; Wang et al., <xref ref-type="bibr" rid="B108">2016</xref>; Li et al., <xref ref-type="bibr" rid="B62">2017</xref>; Alshari et al., <xref ref-type="bibr" rid="B6">2018</xref>; Ljube&#x00161;i&#x00107; et al., <xref ref-type="bibr" rid="B64">2018</xref>; Thavareesan and Mahesan, <xref ref-type="bibr" rid="B100">2020</xref>). Amir et al. (<xref ref-type="bibr" rid="B7">2015</xref>) proposed the use of a support vector regressor (SVR) trained with CBOW (Mikolov et al., <xref ref-type="bibr" rid="B68">2013a</xref>) or GloVe (Pennington et al., <xref ref-type="bibr" rid="B77">2014</xref>) word embeddings, finding that the SVR model out performed various linear models [e.g., Lasso (Yuan and Lin, <xref ref-type="bibr" rid="B116">2006</xref>), Ridge (Hoerl and Kennard, <xref ref-type="bibr" rid="B46">1970</xref>), ElasticNet (Zou and Hastie, <xref ref-type="bibr" rid="B117">2005</xref>) regressors] on the labMT lexicon. However, their models only predicted a binary sentiment polarity (&#x00177; &#x02208; [0, 1]), rather than continuous scores. Li et al. (<xref ref-type="bibr" rid="B62">2017</xref>) extended their work, proposing a class of linear regression models trained with word embeddings to predict affective meanings in several sentiment lexicons such as ANEW (Bradley and Lang, <xref ref-type="bibr" rid="B17">1999</xref>), VAD (Mohammad, <xref ref-type="bibr" rid="B71">2018</xref>). Darwich et al. (<xref ref-type="bibr" rid="B28">2019</xref>) present an excellent review of ASLG.</p>
<p>Many of the human engineered heuristics used in previous work on ASLG can be largely automated via clever application of new machine learning techniques. Sentiment analysis knowledge bases can be constructed using graph-mining and multi-dimensional scaling techniques (Bajpai et al., <xref ref-type="bibr" rid="B10">2016</xref>). Once constructed, these knowledge bases allow for the application of a host of additional methods. Neural tensor networks can be used for knowledge base completion, inferring relationships that were missed during construction (Socher et al., <xref ref-type="bibr" rid="B91">2013a</xref>). Graph neural networks can create rich features from the relationships captured in knowledge bases, allowing sentiment analysis models to handle complex context-based problems (Dowlagar and Mamidi, <xref ref-type="bibr" rid="B34">2021</xref>; Liao et al., <xref ref-type="bibr" rid="B63">2021</xref>; Yang et al., <xref ref-type="bibr" rid="B113">2021</xref>). Ensembles of symbolic and sub-symbolic AI can be used to cover the individual weaknesses of each method (Cambria et al., <xref ref-type="bibr" rid="B19">2020</xref>).</p>
<p>Building on the many of the models discussed above, we develop a framework for augmenting semantic lexicons using word embeddings and pre-trained large language models. Our models output continuous valued sentiment scores that can represent degrees of negative, neutral, and positive sentiment. Our tool reduces the need for crowdsourcing scores from human annotators while still providing similar, and often better, results compared with random reviewers from Amazon Mechanical Turk at a fraction of the cost.</p>
</sec>
<sec sec-type="materials and methods" id="s3">
<title>3. Materials and Methods</title>
<p>We propose two models for predicting happiness scores for the labMT lexicon (Dodds et al., <xref ref-type="bibr" rid="B30">2015</xref>)&#x02014;a general-purpose sentiment lexicon used to measure happiness in text corpora (see section 3.1 for more details).</p>
<p>Our first model is a neural network initialized with pre-trained FastText word embeddings. The model uses fixed word representations to gauge the happiness score for a given expression, enabling us to augment the labMT dataset at a low cost. For simplicity, we will refer to this model as the Token model.</p>
<p>Bridging the link between lexicon-based and contextualized models, we also propose a deep Transformer-based model that uses word definitions to estimate their happiness scores&#x02014;namely, the Dictionary model. The contextualized nature of the input data allows our model to accurately estimate the expressed happiness score for a given word based on its lexical meaning.</p>
<p>We implement our models using Tensorflow (Abadi et al., <xref ref-type="bibr" rid="B1">2016</xref>) and Transformers (Wolf et al., <xref ref-type="bibr" rid="B110">2020</xref>). See section 3.2 and section 3.3 for additional details of our Token and Dictionary models, respectively. Our source code, along with pre-trained models, are publicly available via our GitLab repository (<ext-link ext-link-type="uri" xlink:href="https://gitlab.com/compstorylab/sentiment-analysis">https://gitlab.com/compstorylab/sentiment-analysis</ext-link>).</p>
<sec>
<title>3.1. Data</title>
<p>In this study, we use the labMT dataset as an example sentiment lexicon to test and evaluate our models (Dodds et al., <xref ref-type="bibr" rid="B30">2015</xref>). The labMT lexicon contains roughly ten thousand unique words&#x02014;combining the five thousand most frequently used words from New York Times articles, Google Books, Twitter messages, and music lyrics (Dodds et al., <xref ref-type="bibr" rid="B30">2015</xref>). It is a lexicon designed to gauge changes in the happiness (i.e., valence or hedonic tone) of text corpora. Happiness is defined on a continuous scale <italic>h</italic> &#x02208; {1 &#x02192; 9}, where 1 bounds the most negative (sad) side of the spectrum, and 9 is the most positive (happy). Ratings for each word are crowdsourced via Amazon Mechanical Turk (AMT), taking the average score <italic>h</italic><sub><italic>avg</italic></sub> from 50 reviewers to set a happiness score for any given word. For example, the words &#x0201C;suicide,&#x0201D; &#x0201C;terrorist,&#x0201D; and &#x0201C;coronavirus&#x0201D; have the lowest happiness scores, while the words &#x0201C;laughter,&#x0201D; &#x0201C;happiness,&#x0201D; and &#x0201C;love&#x0201D; have the highest scores. Function and stop words along with numbers and names tend to have neutral scores (<italic>h</italic><sub><italic>avg</italic></sub> &#x02248; 5), such as &#x0201C;the,&#x0201D; &#x0201C;fourth,&#x0201D; &#x0201C;where,&#x0201D; and &#x0201C;per.&#x0201D;</p>
<p>The labMT dataset also powers the Hedonometer, an instrument quantifying daily happiness on Twitter (Dodds et al., <xref ref-type="bibr" rid="B32">2011</xref>). Over the past few years, the labMT lexicon was updated to include new words that were not found in the original survey [e.g, terms related to the COVID19 pandemic (Alshaabi et al., <xref ref-type="bibr" rid="B5">2021b</xref>)].</p>
<p>We are particularly interested in this dataset because it also provides the standard deviation of human ratings for each word, which we use to evaluate our models. In this work, we propose two models to estimate <italic>h</italic><sub><italic>avg</italic></sub> using word embeddings, and thus provide an automated tool to augment the labMT dataset both reliably and efficiently.</p>
<p>In <xref ref-type="fig" rid="F1">Figure 1</xref>, we display a 2D histogram of the human rated happiness scores in the labMT dataset. The figure highlights the degree of uncertainty in human ratings of the emotional valence of words. For example, the word &#x0201C;the&#x0201D; has an average happiness score of <italic>h</italic><sub><italic>avg</italic></sub> &#x0003D; 4.98, with standard deviation of &#x003C3; &#x0003D; 0.91, while the word &#x0201C;hahaha&#x0201D; has a happier score with <italic>h</italic><sub><italic>avg</italic></sub> &#x0003D; 7.94 and &#x003C3; &#x0003D; 1.56. Some words also have a relatively large standard deviation such as &#x0201C;church&#x0201D; (<italic>h</italic><sub><italic>avg</italic></sub> &#x0003D; 5.48, &#x003C3; &#x0003D; 1.85), and &#x0201C;cigarettes&#x0201D; (<italic>h</italic><sub><italic>avg</italic></sub> &#x0003D; 3.31, &#x003C3; &#x0003D; 2.6).</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Emotional valence of words and uncertainty in human ratings of lexical polarity. A 2D histogram of happiness <italic>h</italic><sub><italic>avg</italic></sub> and standard deviation of human ratings for each word in the labMT dataset. Happiness is defined on a continuous scale from 1 to 9, where 1 is the least happy and 9 is the most. Words with a score between 4 and 6 are considered neutral. While the vast majority of words are neutral, there is a positive bias in human language (Dodds et al., <xref ref-type="bibr" rid="B30">2015</xref>). The average standard deviation of human ratings for estimating the emotional valence of words in the labMT dataset is 1.38.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-04-783778-g0001.tif"/>
</fig>
<p>While the majority of words are neutral, with a score between 4 and 6, we still observe a human positivity bias in the English language (Dodds et al., <xref ref-type="bibr" rid="B30">2015</xref>; Aithal and Tan, <xref ref-type="bibr" rid="B3">2021</xref>). On average, the standard deviation of human ratings is 1.38. In our evaluation (section 4), we show how our models perform relative to the uncertainty observed in human ratings.</p>
</sec>
<sec>
<title>3.2. Token Model</title>
<p>Our first model uses a neural network that learns to map words from the labMT lexicon to their corresponding sentiment scores. While still being able to learn a non-linear mapping between the words and their happiness scores, the model only considers the individual words as input&#x02014;enriching its internal utility function with subword representations to estimate the happiness score.</p>
<p>The input word is first processed into a token embedding&#x02014;sequentially breaking each word into its equivalent character-level <italic>n</italic>-grams whereby <italic>n</italic> &#x02208; {3,4,5} (see <xref ref-type="fig" rid="F2">Figure 2</xref> for an illustration). English words have an average length of 5 characters (Miller et al., <xref ref-type="bibr" rid="B70">1958</xref>; Mayzner and Tresselt, <xref ref-type="bibr" rid="B66">1965</xref>), which would yield 6 unique character-level <italic>n</italic>-grams given our tokenization scheme. While we did try shorter and longer sequences, we fix the length of the input sequence to a size of 50 and pad shorter sequences to ensure a universal input size. We choose a longer sequence length to allow us to encode longer <italic>n</italic>-grams and rare words.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Input sequence embeddings. We use two encoding schemes to prepare input sequences for our models: token embeddings (blue) and dictionary embeddings (orange) for our Token and Dictionary models, respectively. Given an input word (e.g., &#x0201C;coronavirus,&#x0201D;) we first break the input token into character-level <italic>n</italic>-grams (<italic>n</italic> &#x02208; {3, 4, 5}). The resulting sequence of <italic>n</italic>-grams along with the original word at the beginning of the embeddings are used in our Token model. Sequences shorter than a specified length are appended with PAD, a padding token ensuring a universal input size. For our Dictionary model, we first look up a dictionary definition for the given input. We then process the input word along with its definition into subwords using WordPiece (Wu et al., <xref ref-type="bibr" rid="B111">2016</xref>). Uncommon and novel words are broken into subwords, with double hashtags indicating that the given token is not a full word.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-04-783778-g0002.tif"/>
</fig>
<p>We then pass the token embeddings to a 300-dimensional embedding layer. We initialize the embedding layer with weights trained with subword information on Common Crawl and Wikipedia using FastText (Bojanowski et al., <xref ref-type="bibr" rid="B16">2017</xref>). In particular, we use weights from a pre-trained model using CBOW with character-level <italic>n</italic>-grams of length 5 and a window size of 5 and 10 (<ext-link ext-link-type="uri" xlink:href="https://fasttext.cc/docs/en/english-vectors.html">https://fasttext.cc/docs/en/english-vectors.html</ext-link>).</p>
<p>The output of the embedding layer is pooled down and passed to a sequence of three dense layers of decreasing sizes: 128, 64, and 32, respectively. We use a rectified linear activation function (ReLU) for all dense layers. We also add a dropout layer after each dense layer, with a 50% dropout rate to add stochasticity to the model, allowing for a simple estimate of uncertainty using the standard deviation of the network&#x00027;s predictions (Srivastava et al., <xref ref-type="bibr" rid="B93">2014</xref>).</p>
<p>We experimented with a few different layout configurations, finding that making the network either wider or deeper has minimal effect on the network performance. Therefore, we choose to keep our model rather simple with roughly 10 million trainable parameters. The output of the last dense layer is finally passed over to a single output layer with a linear activation function to regress a sentiment score between 1 and 9. See <xref ref-type="fig" rid="F3">Figure 3</xref> for a simple diagram of the model architecture.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Model architectures. Our first model is a neural network initialized with pre-trained word embeddings to estimate happiness scores. Our second model, is a deep Transformer-based model that uses word definitions to estimate their sentiment scores. See section 3.2 and 3.3 for further technical details of each model, respectively. Note the Token model is considerably smaller with roughly 10 million trainable parameters compared with the Dictionary model that has a little over 66 million parameters.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-04-783778-g0003.tif"/>
</fig>
</sec>
<sec>
<title>3.3. Dictionary Model</title>
<p>Historically, lexicon-based models have only considered simple statistical methods to estimate the emotional valence of words. Here, we try to bridge the connection between the conventional techniques among the community and recent advances in NLP.</p>
<p>For our second model, we use a contextualized Transformer-based language model to estimate the sentiment score for a given word based on its dictionary definition. While still predicting scores for individual words, we now do so by augmenting each word with its expressed meaning(s) from a general dictionary. Given an input word, we look up its definition via a free online dictionary API available at <ext-link ext-link-type="uri" xlink:href="https://dictionaryapi.dev">https://dictionaryapi.dev</ext-link>.</p>
<p>The average length of definitions for the words found in labMT is roughly 38 words. We choose a maximum definition length of 50 words&#x02014;which covers the 75th percentile of that distribution&#x02014;to ensure that words with multiple definitions are adequately represented. While increasing the sequence length beyond 50 did not improve our accuracy, it increases the model complexity slowing our training and inference time substantially. Therefore, we fix the length of word definitions to a maximum of 50 words. We pad shorter sequences, and truncate words 51 and beyond to ensure a fixed input size.</p>
<p>We estimate the sentiment of each labMT word as follows. The word, along with its definition, is processed into dictionary embeddings by breaking each word into subwords based on their frequency of usage using WordPiece (Wu et al., <xref ref-type="bibr" rid="B111">2016</xref>). This is a widely adopted tokenization technique that breaks uncommon and novel words into subwords, which reduces the vocabulary size of language models and enables them to handle OOV tokens. Other tokenization models will give similar results (Kudo and Richardson, <xref ref-type="bibr" rid="B57">2018</xref>). We only use the word as input to our model for terms without definitions.</p>
<p>In principle, the dictionary embeddings can be passed to a vanilla Transformer model [e.g., BERT (Devlin et al., <xref ref-type="bibr" rid="B29">2019</xref>), XLNet (Yang et al., <xref ref-type="bibr" rid="B114">2019</xref>)]. However, we prefer more manageable models (i.e., smaller and faster) due to their efficiency while maintaining state-of-the-art results. We tried both ALBERT (Lan et al., <xref ref-type="bibr" rid="B59">2020</xref>) and DistilBERT (Sanh et al., <xref ref-type="bibr" rid="B87">2019</xref>). Both models have equivalent performance on our task. The output of the model&#x00027;s pooling layer is passed to a sequence of three dense layers of decreasing sizes with dropout applied after each layer&#x02014;similar to our approach in the Token model. Finally, the output of the last dense layer is projected down to a single output value that servers as the sentiment score prediction.</p>
<p>The Token model is considerably lighter in terms of memory usage, and faster in terms of training and inference time than the Dictionary model. Our current configuration of the Token model results in roughly 10 million trainable parameters compared with the Dictionary model that has over 66 million parameters.</p>
</sec>
</sec>
<sec sec-type="results" id="s4">
<title>4. Results</title>
<sec>
<title>4.1. Ensemble Learning and <italic>k</italic>-Fold Cross-Validation</title>
<p>In the deep learning community, particularly in the NLP domain, it is common to scale up the number of parameters in successful models to eke out additional performance gains. The effectiveness of this approach tends to be correlated with the amount of training data available (i.e., larger models are more effective when trained on larger data sets). With the limited size of our training set, we needed alternative techniques to increase the performance of our models. Ensemble learning is a widely known and adopted family of methods in which the average performance of an ensemble is shown to be both less biased and better than the individual models (Hansen and Salamon, <xref ref-type="bibr" rid="B41">1990</xref>; Krogh and Vedelsby, <xref ref-type="bibr" rid="B56">1994</xref>).</p>
<p>First, we randomly subsample our dataset, taking a 20% subset as our holdout set for testing. Using a 5-fold cross-validation strategy, we break the remaining samples into 5 distinct subsets using a 80/20 split for training/validation. We train one model per fold for a maximum of 500 epochs each, and combine the 5 trained models to form an ensemble. While there are many gradient descent optimization algorithms, we use Adam (Kingma and Ba, <xref ref-type="bibr" rid="B51">2015</xref>) as a popular and well-established optimizer, keeping its default configuration and setting our initial learning rate to 0.001. In <xref ref-type="fig" rid="F4">Figure 4</xref>, we show a breakdown of our ensemble pipeline whereby the blue squares highlight the validation subset for each fold. Note, the holdout set is removed before training the ensemble and is only used for testing a complete ensemble.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Ensemble learning and <italic>k</italic>-fold cross-validation. Using an 80/20 split for training/validation, we train our models for a maximum of 500 epochs per fold for a total of 5 folds. We use the model trained from each fold to build an ensemble because the average performance of an ensemble is less biased and better than the individual models.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-04-783778-g0004.tif"/>
</fig>
<p>To estimate the happiness score for a given word, we take a Monte Carlo approach by sampling 100 predictions per model in the ensemble. We use the training setting for the dropout layers in each model, rather than the test time averaging that is commonly used, so that these predictions are heterogeneous. The mean over these predictions becomes the proposed happiness score, while the standard deviation serves as an estimate of model uncertainty (Gal and Ghahramani, <xref ref-type="bibr" rid="B37">2016</xref>). Providing a point estimate along with an uncertainty band allows us to compare and contrast the level of model uncertainty in our ensembles with the uncertainty observed between human annotators.</p>
</sec>
<sec>
<title>4.2. Comparison With Other Methods and Human Annotators</title>
<p>Although both of our proposed strategies&#x02014;namely using character-level <italic>n</italic>-grams and word definitions&#x02014;performed well, the Dictionary model outperforms the Token model. To evaluate our models we train 10 replicates each and then investigate error distributions obtained using the test set. We report the mean absolute error (MAE) as an estimate of overall performance, along with a selection of percentiles to compare tail behavior across models. Each of these statistics are averaged over the 10 replicates. This process provides us with a strong estimate of the generalization performance for our proposed models.</p>
<p><xref ref-type="table" rid="T1">Table 1</xref> summarizes the results of this evaluation process for our proposed models and ensembles. We provide baseline comparisons to models from previous work (Amir et al., <xref ref-type="bibr" rid="B7">2015</xref>; Li et al., <xref ref-type="bibr" rid="B62">2017</xref>), including popular linear models, random forests, and support vector machines trained with three different flavors of word embeddings: Word2Vec (Mikolov et al., <xref ref-type="bibr" rid="B68">2013a</xref>), GloVe (Pennington et al., <xref ref-type="bibr" rid="B77">2014</xref>), and FastText (Bojanowski et al., <xref ref-type="bibr" rid="B16">2017</xref>). These results indicate that our Token model outperforms all prior baselines, our Dictionary model outperforms our Token model, and both of our proposed models benefited from ensemble learning. Though the ensembles outperformed the individual models in both cases, it is interesting to note that they also had longer tails for their error distributions.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Summary statistics of the testing subset comparing our models to the annotated ratings reported in labMT.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left" style="border-bottom: thin solid #000000;" colspan="2"><bold>Mean absolute error (MAE)</bold></th>
<th valign="top" align="center" style="border-bottom: thin solid #000000;" colspan="5"><bold>Percentiles</bold></th>
</tr>
<tr>
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>Average</bold></th>
<th valign="top" align="center"><bold>25</bold><italic><bold><sup>th</sup></bold></italic></th>
<th valign="top" align="center"><bold>50</bold><italic><bold><sup>th</sup></bold></italic></th>
<th valign="top" align="center"><bold>75</bold><italic><bold><sup>th</sup></bold></italic></th>
<th valign="top" align="center"><bold>85</bold><italic><bold><sup>th</sup></bold></italic></th>
<th valign="top" align="center"><bold>95</bold><italic><bold><sup>th</sup></bold></italic></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><bold>Linear models</bold></td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">ElasticNet &#x0002B; <italic>Word2Vec</italic></td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.83</td>
</tr>
<tr>
<td valign="top" align="left">ElasticNet &#x0002B; <italic>GloVe</italic></td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.82</td>
</tr>
<tr>
<td valign="top" align="left">ElasticNet &#x0002B; <italic>FastText</italic></td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.82</td>
</tr>
<tr>
<td valign="top" align="left">LASSO &#x0002B; <italic>Word2Vec</italic></td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.83</td>
</tr>
<tr>
<td valign="top" align="left">LASSO &#x0002B; <italic>GloVe</italic></td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.82</td>
</tr>
<tr>
<td valign="top" align="left">LASSO &#x0002B; <italic>FastText</italic></td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.82</td>
</tr>
<tr>
<td valign="top" align="left">Ridge &#x0002B; <italic>Word2Vec</italic></td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.75</td>
</tr>
<tr>
<td valign="top" align="left">Ridge &#x0002B; <italic>GloVe</italic></td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.79</td>
</tr>
<tr>
<td valign="top" align="left">Ridge &#x0002B; <italic>FastText</italic></td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.73</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.74</td>
</tr>
<tr>
<td valign="top" align="left"><bold>Random forest (RF) models</bold></td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">RF &#x0002B; <italic>Word2Vec</italic></td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.71</td>
<td valign="top" align="center">0.78</td>
</tr>
<tr>
<td valign="top" align="left">RF &#x0002B; <italic>GloVe</italic></td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.70</td>
<td valign="top" align="center">0.71</td>
<td valign="top" align="center">0.71</td>
<td valign="top" align="center">0.71</td>
</tr>
<tr>
<td valign="top" align="left">RF &#x0002B; <italic>FastText</italic></td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.69</td>
</tr>
<tr>
<td valign="top" align="left"><bold>Support vector regressor (SVR) models</bold></td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">SVR &#x0002B; <italic>Word2Vec</italic></td>
<td valign="top" align="center">0.65</td>
<td valign="top" align="center">0.65</td>
<td valign="top" align="center">0.65</td>
<td valign="top" align="center">0.66</td>
<td valign="top" align="center">0.66</td>
<td valign="top" align="center">0.67</td>
</tr>
<tr>
<td valign="top" align="left">SVR &#x0002B; <italic>GloVe</italic></td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.67</td>
<td valign="top" align="center">0.66</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.69</td>
</tr>
<tr>
<td valign="top" align="left">SVR &#x0002B; <italic>FastText</italic></td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center">0.65</td>
<td valign="top" align="center">0.66</td>
<td valign="top" align="center">0.66</td>
</tr>
<tr>
<td valign="top" align="left"><bold>Proposed models</bold></td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">Token model (single)</td>
<td valign="top" align="center">0.62</td>
<td valign="top" align="center">0.60</td>
<td valign="top" align="center">0.61</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center">0.65</td>
<td valign="top" align="center">0.66</td>
</tr>
<tr>
<td valign="top" align="left">Token model (ensemble)</td>
<td valign="top" align="center">0.57</td>
<td valign="top" align="center">0.29</td>
<td valign="top" align="center">0.44</td>
<td valign="top" align="center">0.66</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.77</td>
</tr>
<tr>
<td valign="top" align="left">Dictionary model (single)</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.49</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.51</td>
<td valign="top" align="center"><bold>0.51</bold></td>
<td valign="top" align="center"><bold>0.52</bold></td>
</tr>
<tr>
<td valign="top" align="left">Dictionary model (ensemble)</td>
<td valign="top" align="center"><bold>0.45</bold></td>
<td valign="top" align="center"><bold>0.15</bold></td>
<td valign="top" align="center"><bold>0.31</bold></td>
<td valign="top" align="center"><bold>0.40</bold></td>
<td valign="top" align="center">0.52</td>
<td valign="top" align="center">0.59</td>
</tr>
<tr>
<td valign="top" align="left">Human ratings (standard deviation &#x003C3;)</td>
<td valign="top" align="center">1.38</td>
<td valign="top" align="center">1.18</td>
<td valign="top" align="center">1.36</td>
<td valign="top" align="center">1.56</td>
<td valign="top" align="center">1.69</td>
<td valign="top" align="center">1.90</td>
</tr>
<tr>
<td valign="top" align="left">Human ratings (variance &#x003C3;<sup>2</sup>)</td>
<td valign="top" align="center">1.99</td>
<td valign="top" align="center">1.39</td>
<td valign="top" align="center">1.85</td>
<td valign="top" align="center">2.43</td>
<td valign="top" align="center">2.86</td>
<td valign="top" align="center">3.61</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Each word in the labMT lexicon is scored by 50 distinct individuals and the final happiness score is the unweighted mean of their scores (Dodds et al., <xref ref-type="bibr" rid="B30">2015</xref>). We report the standard deviation and variance of the ratings as a baseline to assess the human&#x00027;s confidence in the reported scores. Comparing our predictions with the annotations crowdsourced via AMT, our mean absolute errors are on par with the variance observe in the human annotated labMT scores. In addition to our proposed models, we also evaluate three groups of baseline models based on linear regression, random forests, and support vector machines. We trained and evaluated each baseline model over 10 trials using one of three pre-trained embeddings as the primary input: word2vec-google-news-300 (Mikolov et al., <xref ref-type="bibr" rid="B68">2013a</xref>), glove-wiki-gigaword-300 (Pennington et al., <xref ref-type="bibr" rid="B77">2014</xref>), fasttext-wiki-news-subwords-300 (Bojanowski et al., <xref ref-type="bibr" rid="B16">2017</xref>). Our baseline models are similar to those seen in Amir et al. (<xref ref-type="bibr" rid="B7">2015</xref>) and Li et al. (<xref ref-type="bibr" rid="B62">2017</xref>). Best models are highlighted in bold</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>We further examine the error distributions to investigate if the models have a bias toward high or low happiness scores. In <xref ref-type="fig" rid="F5">Figures 5</xref>, <xref ref-type="fig" rid="F6">6</xref>, we display a breakdown of our MAE distributions for the Token and Dictionary models, respectively. For ease of interpretation and visualization, we categorize the happiness scores into three groups: negative (<italic>h</italic><sub><italic>avg</italic></sub> &#x02208; [1,4)), neutral (<italic>h</italic><sub><italic>avg</italic></sub> &#x02208; [4,6]), and positive (<italic>h</italic><sub><italic>avg</italic></sub> &#x02208; (6,9]). While the distributions show our models operate well on all words, particularly neutral expressions, we note a relatively higher MAE for negative words, whereby our predictions to these terms are more positive than the annotations.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Error distributions for the Token model. We display mean absolute errors for predictions using the Token model on all words in labMT. We arrange the happiness scores into three groups: negative (<italic>h</italic><sub><italic>avg</italic></sub> &#x02208; [1,4), orange), neutral (<italic>h</italic><sub><italic>avg</italic></sub> &#x02208; [4,6], gray), and positive (<italic>h</italic><sub><italic>avg</italic></sub> &#x02208; (6,9], green). Most words have an MAE less than 1 with the exception of a few outliers. We see a relatively higher MAE for negative and positive terms compared to neutral expressions.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-04-783778-g0005.tif"/>
</fig>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Error distributions for the Dictionary model. We display mean absolute errors for predictions using the Dictionary model on all words in labMT. Again, we categorize the happiness scores into three groups: negative (<italic>h</italic><sub><italic>avg</italic></sub> &#x02208; [1,4), orange), neutral (<italic>h</italic><sub><italic>avg</italic></sub> &#x02208; [4,6], gray), and positive (<italic>h</italic><sub><italic>avg</italic></sub> &#x02208; (6,9], green). Similar to the Token model, most words have an MAE less than 1 with the exception of a few outliers. While the Dictionary model outperforms the Token model, we still observe a higher MAE for negative and positive terms compared to neutral expressions.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-04-783778-g0006.tif"/>
</fig>
<p>We also compare our predictions to the ground-truth ratings, examining the degree to which the models either overshoot or undershoot the happiness scores crowdsourced via AMT. Words in the labMT lexicon were scored by taking the average happiness score of distinct evaluations from 50 different individuals (see Table S2, Dodds et al., <xref ref-type="bibr" rid="B30">2015</xref>). Since the variance of human ratings and our model MAEs are on the same scale, we can use the observed average variance of the ratings (1.17) as a baseline to assess rater confidence in the reported scores. Comparing our models to that baseline, we note that all models offer consistent predictions with similar expectations to a random and reliable reviewer from AMT. See <xref ref-type="table" rid="T1">Table 1</xref> for further statistical details.</p>
<p>In <xref ref-type="fig" rid="F7">Figures 7</xref>, <xref ref-type="fig" rid="F8">8</xref>, we display the top-50 words with the highest mean absolute error for the Token and Dictionary models, respectively. While the models always predict the right emotional attitude outlining each word based on its lexical polarity, they bias toward neutral by undershooting scores for happy words, and overshooting scores for sad expressions.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Token model: Top-50 words with the highest mean absolute error. Model predictions are shown in blue and the crowdsourced annotations are displayed in gray. While still maintaining relatively low MAE, most of our predictions are conservative&#x02014;marginally underestimating words with extremely high happiness scores, and overestimating words with low happiness scores.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-04-783778-g0007.tif"/>
</fig>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Dictionary model: Top-50 words with the highest mean absolute error. Model predictions are shown in blue and the crowdsourced annotations are displayed in gray. Note, the vast majority of words with relatively high MAE also have high standard deviations of AMT ratings. Words that have multiple definitions will have a neutral score (e.g., lying). A neutral happiness score is also often predicted for words because we are unable to obtain good definitions for them to use as input. Although we have definitions for most words in our dataset, we still have a little over 1,500 words with missing definitions. Most of these words are names (e.g., &#x0201C;&#x02018;Burke,&#x0201D;) and slang (e.g., &#x0201C;xmas,&#x0201D; and &#x0201C;ta.&#x0201D;)</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-04-783778-g0008.tif"/>
</fig>
<p>One possible explanation of this systematic behavior is the lack of words with extreme happiness scores in the labMT lexicon. It is possible to train models with a smaller but balanced subset of the dataset to overcome that challenge. Doing so, however, would reduce the size of training/validation samples substantially. Still, our margin of error is relatively low compared to human ratings. Future investigations may test and improve the models by examining larger sentiment lexicons.</p>
<p>Another key factor that plays a big role in our prediction error is obtaining good word definitions, or the lack thereof, to use as input for our Dictionary model. Surprisingly, outsourcing definitions from online dictionaries for a large set of words is rather challenging, especially if you opt-out of reliable but paid services. In our work, we choose not to use an urban dictionary or any services with paid APIs. We use a free online dictionary API that is available at <ext-link ext-link-type="uri" xlink:href="https://dictionaryapi.dev">https://dictionaryapi.dev</ext-link>.</p>
<p>While we do have definitions for most words in our dataset, a total of 1518 words have missing definitions. Most of these words are names, abbreviations, and slang terms (e.g., &#x0201C;xams,&#x0201D; &#x0201C;foto,&#x0201D; &#x0201C;nvm,&#x0201D; and &#x0201C;lmao&#x0201D;). Words with multiple definitions can also cancel each other&#x00027;s score (e.g., &#x0201C;lying&#x0201D;).</p>
<p>Notably, the vast majority of words with high MAE also have high AMT standard deviations. To further investigate prediction accuracy, we examine the overlap between the predictions and human ratings. In particular, we compute the intersection over union (IOU) between the predicted happiness score <inline-formula><mml:math id="M1"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>v</mml:mi><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x000B1;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x003C3;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, and the corresponding value from the annotated ratings <italic>h</italic><sub><italic>avg</italic></sub>&#x000B1;&#x003C3;.</p>
<p>The Token model underestimates the happiness score for &#x0201C;win&#x0201D;&#x02014;the only word with a prediction that falls outside the range of human annotated happiness scores. The remaining predicted happiness scores fall well within the range of scores crowdsourced via AMT. Similarly, the Dictionary model slightly underestimates the happiness scores for &#x0201C;mamma&#x0201D; while overestimating the scores for &#x0201C;lying,&#x0201D; and &#x0201C;coronavirus.&#x0201D;</p>
</sec>
</sec>
<sec sec-type="discussion" id="s5">
<title>5. Discussion</title>
<p>As the growing demand for sentiment-aware intelligent systems increases, we will continue to see improvements to both lexicon-based models and contextual language models. While contextualized models are suitable for a wide set of applications, lexicon-based models are used by computational linguists, journalists, and data scientists who are interested in studying how individual words contribute to sentiment trends.</p>
<p>Sentiment lexicons, however, have to be updated periodically to support new words and expressions that were not considered when the dictionaries were assembled. In this paper, we proposed two models for predicting sentiment scores to augment semantic dictionaries using word embeddings and pre-trained large language models. Our first model establishes a baseline using a neural network initialized with pre-trained word embeddings, while our second model features a deep Transformer-based network that brings into play word definitions to estimate their lexical polarity. Our results and evaluation of both models demonstrate human-level performance on a state-of-the-art human annotated list of words.</p>
<p>Although both models can predict scores for novel words, we acknowledge a few shortcomings. Our Token model relies on subword information to estimate a happiness score for any given word. For example, using subwords for &#x0201C;coronavirus&#x0201D; yields a good estimate given that it contains &#x0201C;virus.&#x0201D; By contrast, parsing character-level <italic>n</italic>-grams for other words (e.g., &#x0201C;covid&#x0201D;) may not reveal any further information. We can overcome that hurdle by using the word definition as input to our Dictionary model to gauge its happiness score. Words, however, often have different meanings based on context. Finding good definitions may be challenging, especially for slang, informal expressions, and abbreviations. We recommend using the Dictionary model whenever it is possible to outsource a good definition of the word.</p>
<p>A natural next step would be to develop similar models for other languages, for example by building a model for each language, or a multilingual model. Fortunately, FastText (Bojanowski et al., <xref ref-type="bibr" rid="B16">2017</xref>) provides pre-trained word embeddings for over 100 languages. Therefore, it is easy to upgrade the Token model to support other languages. Updating the Dictionary model is also a straightforward task by simply adopting a multilingual Transformer-based model pre-trained with several languages [e.g., Multilingual BERT (Devlin et al., <xref ref-type="bibr" rid="B29">2019</xref>)]. We caution against translating words and using the same English scores because most words do not have a one-to-one mapping into other languages, and are often used to express different meanings by the native speakers of any given language (Dodds et al., <xref ref-type="bibr" rid="B30">2015</xref>).</p>
<p>Another vast space of improvements would be to adopt our proposed strategies to develop prediction models for other semantic dictionaries. Researchers can further fine-tune these models to predict other sentiment scores. For example, the happiness scores in the labMT (Dodds et al., <xref ref-type="bibr" rid="B30">2015</xref>) dataset are closely aligned with the valence scores in the NRC-VAD lexicon (Mohammad, <xref ref-type="bibr" rid="B71">2018</xref>). We envision future work developing similar models to predict other semantic differentials such as arousal and dominance (Mohammad, <xref ref-type="bibr" rid="B71">2018</xref>), EPA (Osgood, <xref ref-type="bibr" rid="B73">1962</xref>), and SocialSent (Hamilton et al., <xref ref-type="bibr" rid="B40">2016</xref>). Our primary goal is to provide an easy and robust method to augment semantic dictionaries to empower researchers to maintain and expand them at a relatively low cost using today&#x00027;s state-of-the-art NLP methods.</p>
</sec>
<sec sec-type="data-availability" id="s6">
<title>Data Availability Statement</title>
<p>Our source code along with pre-trained models are publicly available on our Gitlab repository (<ext-link ext-link-type="uri" xlink:href="https://gitlab.com/compstorylab/sentiment-analysis">https://gitlab.com/compstorylab/sentiment-analysis</ext-link>).</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>TA designed and developed the methods. TA and CV verified and analyzed data. TA, CV, MF, MA, CD, and PD edited the manuscript. CD and PD supervised the project. All authors provided critical feedback and helped shape the research, analysis and manuscript.</p>
</sec>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>We are grateful for the computing resources provided by the Vermont Advanced Computing Core and financial support from Google and the Massachusetts Mutual Life Insurance Company. Computations were performed on the Vermont Advanced Computing Core supported in part by NSF award No. OAC-1827314.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec> 
</body>
<back>
<ack><p>We thank Anne Marie Stupinski, Julia Zimmerman and our colleagues at the Computational Story Lab for their insightful discussion and suggestions on this project.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Abadi</surname> <given-names>M.</given-names></name> <name><surname>Barham</surname> <given-names>P.</given-names></name> <name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Chen</surname> <given-names>Z.</given-names></name> <name><surname>Davis</surname> <given-names>A.</given-names></name> <name><surname>Dean</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Tensorflow: a system for large-scale machine learning</article-title>, in <source>Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation OSDI&#x00027;16</source> (<publisher-loc>Berkeley, CA</publisher-loc>: <publisher-name>USENIX Association</publisher-name>), <fpage>265</fpage>&#x02013;<lpage>283</lpage>.<pub-id pub-id-type="pmid">28708848</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Agarwal</surname> <given-names>A.</given-names></name> <name><surname>Xie</surname> <given-names>B.</given-names></name> <name><surname>Vovsha</surname> <given-names>I.</given-names></name> <name><surname>Rambow</surname> <given-names>O.</given-names></name> <name><surname>Passonneau</surname> <given-names>R.</given-names></name></person-group> (<year>2011</year>). <article-title>Sentiment analysis of Twitter data</article-title>, in <source>Proceedings of the Workshop on Language in Social Media (LSM 2011)</source> (<publisher-loc>Portland</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>30</fpage>&#x02013;<lpage>38</lpage>.</citation>
</ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Aithal</surname> <given-names>M.</given-names></name> <name><surname>Tan</surname> <given-names>C.</given-names></name></person-group> (<year>2021</year>). <article-title>On positivity bias in negative reviews</article-title>, in <source>Proceedings of the Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (ACL-IJCNLP 2021)</source> (<publisher-loc>Association for Computational Linguistics</publisher-loc>).</citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alshaabi</surname> <given-names>T.</given-names></name> <name><surname>Adams</surname> <given-names>J. L.</given-names></name> <name><surname>Arnold</surname> <given-names>M. V.</given-names></name> <name><surname>Minot</surname> <given-names>J. R.</given-names></name> <name><surname>Dewhurst</surname> <given-names>D. R.</given-names></name> <name><surname>Reagan</surname> <given-names>A. J.</given-names></name> <etal/></person-group>. (<year>2021a</year>). <article-title>Storywrangler: a massive exploratorium for sociolinguistic, cultural, socioeconomic, and political timelines using Twitter</article-title>. <source>Sci. Adv.</source> <volume>7</volume>:<fpage>eabe6534</fpage>. <pub-id pub-id-type="doi">10.1126/sciadv.abe6534</pub-id><pub-id pub-id-type="pmid">34272243</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alshaabi</surname> <given-names>T.</given-names></name> <name><surname>Arnold</surname> <given-names>M. V.</given-names></name> <name><surname>Minot</surname> <given-names>J. R.</given-names></name> <name><surname>Adams</surname> <given-names>J. L.</given-names></name> <name><surname>Dewhurst</surname> <given-names>D. R.</given-names></name> <name><surname>Reagan</surname> <given-names>A. J.</given-names></name> <etal/></person-group>. (<year>2021b</year>). <article-title>How the world&#x00027;s collective attention is being paid to a pandemic: COVID-19 related n-gram time series for 24 languages on Twitter</article-title>. <source>PLoS One</source> <volume>16</volume>:<fpage>e0244476</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0244476</pub-id><pub-id pub-id-type="pmid">33406101</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Alshari</surname> <given-names>E. M.</given-names></name> <name><surname>Azman</surname> <given-names>A.</given-names></name> <name><surname>Doraisamy</surname> <given-names>S.</given-names></name> <name><surname>Mustapha</surname> <given-names>N.</given-names></name> <name><surname>Alkeshr</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>Effective method for sentiment lexical dictionary enrichment based on Word2Vec for sentiment analysis</article-title>, in <source>2018 Fourth International Conference on Information Retrieval and Knowledge Management (CAMP)</source> (<publisher-loc>Kota Kinabalu</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>5</lpage>.</citation>
</ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Amir</surname> <given-names>S.</given-names></name> <name><surname>F. Astudillo</surname> <given-names>R.</given-names></name> <name><surname>Ling</surname> <given-names>W.</given-names></name> <name><surname>Martins</surname> <given-names>B.</given-names></name> <name><surname>Silva</surname> <given-names>M. J.</given-names></name> <name><surname>Trancoso</surname> <given-names>I.</given-names></name></person-group> (<year>2015</year>). <article-title>INESC-ID: a regression model for large scale Twitter sentiment lexicon induction</article-title>, in <source>Proceedings of the 9th International Workshop on Semantic Evaluation (SemEval 2015)</source> (<publisher-loc>Denver, CO</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>613</fpage>&#x02013;<lpage>618</lpage>.</citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Augustyniak</surname> <given-names>&#x00141;.</given-names></name> <name><surname>Szyma&#x00144;ski</surname> <given-names>P.</given-names></name> <name><surname>Kajdanowicz</surname> <given-names>T.</given-names></name> <name><surname>Tulig&#x00142;owicz</surname> <given-names>W.</given-names></name></person-group> (<year>2016</year>). <article-title>Comprehensive study on lexicon-based ensemble classification sentiment analysis</article-title>. <source>Entropy</source> <volume>18</volume>:<fpage>4</fpage>. <pub-id pub-id-type="doi">10.3390/e18010004</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Baccianella</surname> <given-names>S.</given-names></name> <name><surname>Esuli</surname> <given-names>A.</given-names></name> <name><surname>Sebastiani</surname> <given-names>F.</given-names></name></person-group> (<year>2010</year>). <article-title>SentiWordNet 3.0: An enhanced lexical resource for sentiment analysis and opinion mining</article-title>, in <source>Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC&#x00027;10)</source> (<publisher-loc>Valletta</publisher-loc>: <publisher-name>European Language Resources Association (ELRA</publisher-name>)).</citation>
</ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bajpai</surname> <given-names>R.</given-names></name> <name><surname>Ho</surname> <given-names>D.</given-names></name> <name><surname>Cambria</surname> <given-names>E.</given-names></name></person-group> (<year>2016</year>). <article-title>Developing a concept-level knowledge base for sentiment analysis in singlish</article-title>, in <source>International Conference on Intelligent Text Processing and Computational Linguistics</source> (<publisher-loc>Konya</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>347</fpage>&#x02013;<lpage>361</lpage>.</citation>
</ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bakshi</surname> <given-names>R. K.</given-names></name> <name><surname>Kaur</surname> <given-names>N.</given-names></name> <name><surname>Kaur</surname> <given-names>R.</given-names></name> <name><surname>Kaur</surname> <given-names>G.</given-names></name></person-group> (<year>2016</year>). <article-title>Opinion mining and sentiment analysis</article-title>, in <source>2016 3rd international Conference on Computing for Sustainable Global Development (INDIACom)</source> (<publisher-loc>New Delhi</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>452</fpage>&#x02013;<lpage>455</lpage>.<pub-id pub-id-type="pmid">34435102</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bathina</surname> <given-names>K. C.</given-names></name> <name><surname>Ten Thij</surname> <given-names>M.</given-names></name> <name><surname>Lorenzo-Luaces</surname> <given-names>L.</given-names></name> <name><surname>Rutter</surname> <given-names>L. A.</given-names></name> <name><surname>Bollen</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>Individuals with depression express more distorted thinking on social media</article-title>. <source>Nat. Human Behav.</source> <volume>5</volume>, <fpage>458</fpage>&#x02013;<lpage>466</lpage>. <pub-id pub-id-type="doi">10.1038/s41562-021-01050-7</pub-id><pub-id pub-id-type="pmid">33574604</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Beigi</surname> <given-names>G.</given-names></name> <name><surname>Hu</surname> <given-names>X.</given-names></name> <name><surname>Maciejewski</surname> <given-names>R.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name></person-group> (<year>2016</year>). <article-title>An overview of sentiment analysis in social media and its applications in disaster relief</article-title>, in <source>Sentiment Analysis and Ontology Engineering</source> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>313</fpage>&#x02013;<lpage>340</lpage>.</citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Ducharme</surname> <given-names>R.</given-names></name> <name><surname>Vincent</surname> <given-names>P.</given-names></name> <name><surname>Janvin</surname> <given-names>C.</given-names></name></person-group> (<year>2003</year>). <article-title>A neural probabilistic language model</article-title>. <source>J. Mach. Learn. Res.</source> <volume>3</volume>, <fpage>1137</fpage>&#x02013;<lpage>1155</lpage>. <pub-id pub-id-type="doi">10.5555/944919.944966</pub-id><pub-id pub-id-type="pmid">26215079</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bhatt</surname> <given-names>A.</given-names></name> <name><surname>Patel</surname> <given-names>A.</given-names></name> <name><surname>Chheda</surname> <given-names>H.</given-names></name> <name><surname>Gawande</surname> <given-names>K.</given-names></name></person-group> (<year>2015</year>). <article-title>Amazon review classification and sentiment analysis</article-title>. <source>Int. J. Comput. Sci. Inf. Technol.</source> <volume>6</volume>, <fpage>5107</fpage>&#x02013;<lpage>5110</lpage>.</citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bojanowski</surname> <given-names>P.</given-names></name> <name><surname>Grave</surname> <given-names>E.</given-names></name> <name><surname>Joulin</surname> <given-names>A.</given-names></name> <name><surname>Mikolov</surname> <given-names>T.</given-names></name></person-group> (<year>2017</year>). <article-title>Enriching word vectors with subword information</article-title>. <source>Trans. Assoc. Comput. Linguist.</source> <volume>5</volume>, <fpage>135</fpage>&#x02013;<lpage>146</lpage>. <pub-id pub-id-type="doi">10.1162/tacl_a_00051</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bradley</surname> <given-names>M. M.</given-names></name> <name><surname>Lang</surname> <given-names>P. J.</given-names></name></person-group> (<year>1999</year>). <source>Affective Norms for English Words (ANEW): Instruction Manual and Affective Ratings</source>. <publisher-loc>Technical Report, Technical Report C-1, Gainesville, FL</publisher-loc>: <publisher-name>Center for Research in Psychophysiology</publisher-name>.<pub-id pub-id-type="pmid">29322399</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cabral</surname> <given-names>L.</given-names></name> <name><surname>Hortacsu</surname> <given-names>A.</given-names></name></person-group> (<year>2010</year>). <article-title>The dynamics of seller reputation: evidence from eBay</article-title>. <source>J. Ind. Econ.</source> <volume>58</volume>, <fpage>54</fpage>&#x02013;<lpage>78</lpage>. <pub-id pub-id-type="doi">10.1111/j.1467-6451.2010.00405.x</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cambria</surname> <given-names>E.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Xing</surname> <given-names>F. Z.</given-names></name> <name><surname>Poria</surname> <given-names>S.</given-names></name> <name><surname>Kwok</surname> <given-names>K.</given-names></name></person-group> (<year>2020</year>). <article-title>Senticnet 6: ensemble application of symbolic and subsymbolic ai for sentiment analysis</article-title>, in <source>Proceedings of the 29th ACM international conference on information &#x00026; knowledge management</source> <fpage>105</fpage>&#x02013;<lpage>114</lpage>.</citation>
</ref>
<ref id="B20">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>H.</given-names></name> <name><surname>Yang</surname> <given-names>C.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Sun</surname> <given-names>M.</given-names></name> <name><surname>Jin</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <source>From Symbols to Embeddings: a Tale of Two Representations in Computational Social Science</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/2106.14198">https://arxiv.org/abs/2106.14198</ext-link></citation>
</ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>Q.</given-names></name> <name><surname>Zhu</surname> <given-names>X.</given-names></name> <name><surname>Ling</surname> <given-names>Z.-H.</given-names></name> <name><surname>Wei</surname> <given-names>S.</given-names></name> <name><surname>Jiang</surname> <given-names>H.</given-names></name> <name><surname>Inkpen</surname> <given-names>D.</given-names></name></person-group> (<year>2017</year>). <article-title>Enhanced LSTM for natural language inference</article-title>, in <source>Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source> (<publisher-loc>Vancouver, BC</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>1657</fpage>&#x02013;<lpage>1668</lpage>.</citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Colhon</surname> <given-names>M.</given-names></name> <name><surname>Vl&#x00103;du&#x00163;escu</surname> <given-names>&#x0015E;.</given-names></name> <name><surname>Negrea</surname> <given-names>X.</given-names></name></person-group> (<year>2017</year>). <article-title>How objective a neutral word is? a neutrosophic approach for the objectivity degrees of neutral words</article-title>. <source>Symmetry</source> <volume>9</volume>:<fpage>280</fpage>. <pub-id pub-id-type="doi">10.3390/sym9110280</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Conway</surname> <given-names>M.</given-names></name> <name><surname>O&#x00027;Connor</surname> <given-names>D.</given-names></name></person-group> (<year>2016</year>). <article-title>Social media, big data, and mental health: current advances and ethical implications</article-title>. <source>Curr. Opin. Psychol.</source> <volume>9</volume>, <fpage>77</fpage>&#x02013;<lpage>82</lpage>. <pub-id pub-id-type="doi">10.1016/j.copsyc.2016.01.004</pub-id><pub-id pub-id-type="pmid">27042689</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Coppersmith</surname> <given-names>G.</given-names></name> <name><surname>Dredze</surname> <given-names>M.</given-names></name> <name><surname>Harman</surname> <given-names>C.</given-names></name></person-group> (<year>2014</year>). <article-title>Quantifying mental health signals in Twitter</article-title>, in <source>Proceedings of the Workshop on Computational Linguistics and Clinical Psychology: From Linguistic Signal to Clinical Reality</source> (<publisher-loc>Baltimore, MD</publisher-loc>), <fpage>51</fpage>&#x02013;<lpage>60</lpage>.</citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Crawford</surname> <given-names>K.</given-names></name></person-group> (<year>2019</year>). <article-title>Halt the use of facial-recognition technology until it is regulated</article-title>. <source>Nature</source> <volume>572</volume>, <fpage>565</fpage>&#x02013;<lpage>566</lpage>. <pub-id pub-id-type="doi">10.1038/d41586-019-02514-7</pub-id><pub-id pub-id-type="pmid">31455918</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Crawford</surname> <given-names>K.</given-names></name> <name><surname>Paglen</surname> <given-names>T.</given-names></name></person-group> (<year>2021</year>). <article-title>Excavating ai: the politics of images in machine learning training sets</article-title>. <source>AI &#x00026; SOCIETY</source> <volume>9</volume>, <fpage>1</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1007/s00146-021-01162-8</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dai</surname> <given-names>Z.</given-names></name> <name><surname>Yang</surname> <given-names>Z.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Carbonell</surname> <given-names>J.</given-names></name> <name><surname>Le</surname> <given-names>Q.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R.</given-names></name></person-group> (<year>2019</year>). <article-title>Transformer-XL: attentive language models beyond a fixed-length context</article-title>, in <source>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</source> (<publisher-loc>Florence</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>2978</fpage>&#x02013;<lpage>2988</lpage>.</citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Darwich</surname> <given-names>M.</given-names></name> <name><surname>Mohd Noah</surname> <given-names>S. A.</given-names></name> <name><surname>Omar</surname> <given-names>N.</given-names></name> <name><surname>Osman</surname> <given-names>N. A.</given-names></name></person-group> (<year>2019</year>). <article-title>Corpus-based techniques for sentiment lexicon generation: a review</article-title>. <source>J. Digit. Inf. Manag.</source> <volume>17</volume>, <fpage>296</fpage>. <pub-id pub-id-type="doi">10.6025/jdim/2019/17/5/296-305</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Devlin</surname> <given-names>J.</given-names></name> <name><surname>Chang</surname> <given-names>M.-W.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>Toutanova</surname> <given-names>K.</given-names></name></person-group> (<year>2019</year>). <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>, in <source>Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)</source> (<publisher-loc>Minneapolis</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>4171</fpage>&#x02013;<lpage>4186</lpage>.</citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dodds</surname> <given-names>P. S.</given-names></name> <name><surname>Clark</surname> <given-names>E. M.</given-names></name> <name><surname>Desu</surname> <given-names>S.</given-names></name> <name><surname>Frank</surname> <given-names>M. R.</given-names></name> <name><surname>Reagan</surname> <given-names>A. J.</given-names></name> <name><surname>Williams</surname> <given-names>J. R.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Human language reveals a universal positivity bias</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>112</volume>, <fpage>2389</fpage>&#x02013;<lpage>2394</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1411678112</pub-id><pub-id pub-id-type="pmid">25675475</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dodds</surname> <given-names>P. S.</given-names></name> <name><surname>Danforth</surname> <given-names>C. M.</given-names></name></person-group> (<year>2010</year>). <article-title>Measuring the happiness of large-scale written expression: songs, blogs, and presidents</article-title>. <source>J. Happiness Stud.</source> <volume>11</volume>, <fpage>441</fpage>&#x02013;<lpage>456</lpage>. <pub-id pub-id-type="doi">10.1007/s10902-009-9150-9</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dodds</surname> <given-names>P. S.</given-names></name> <name><surname>Harris</surname> <given-names>K. D.</given-names></name> <name><surname>Kloumann</surname> <given-names>I. M.</given-names></name> <name><surname>Bliss</surname> <given-names>C. A.</given-names></name> <name><surname>Danforth</surname> <given-names>C. M.</given-names></name></person-group> (<year>2011</year>). <article-title>Temporal patterns of happiness and information in a global social network: hedonometrics and Twitter</article-title>. <source>PLoS One</source> <volume>6</volume>:<fpage>e26752</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0026752</pub-id><pub-id pub-id-type="pmid">22163266</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dosovitskiy</surname> <given-names>A.</given-names></name> <name><surname>Beyer</surname> <given-names>L.</given-names></name> <name><surname>Kolesnikov</surname> <given-names>A.</given-names></name> <name><surname>Weissenborn</surname> <given-names>D.</given-names></name> <name><surname>Zhai</surname> <given-names>X.</given-names></name> <name><surname>Unterthiner</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>An image is worth 16x16 words: transformers for image recognition at scale</article-title>, in <source>Proceedings of the International Conference on Learning Representations ICRL&#x00027;21</source>.</citation>
</ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dowlagar</surname> <given-names>S.</given-names></name> <name><surname>Mamidi</surname> <given-names>R.</given-names></name></person-group> (<year>2021</year>). <article-title>Graph convolutional networks with multi-headed attention for code-mixed sentiment analysis</article-title>, in <source>Proceedings of the First Workshop on Speech and Language Technologies for Dravidian Languages</source> (<publisher-loc>Kyiv</publisher-loc>), <fpage>65</fpage>&#x02013;<lpage>72</lpage>.</citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Feldman</surname> <given-names>R.</given-names></name></person-group> (<year>2013</year>). <article-title>Techniques and applications for sentiment analysis</article-title>. <source>Commun. ACM</source> <volume>56</volume>, <fpage>82</fpage>&#x02013;<lpage>89</lpage>. <pub-id pub-id-type="doi">10.1145/2436256.2436274</pub-id></citation>
</ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="editor"><name><surname>Fellbaum</surname> <given-names>C.</given-names></name></person-group> editor (<year>1998</year>). <article-title>Language, Speech, and Communication. A Bradford Book</article-title>, in <source>WordNet: An Electronic Lexical Database</source> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>A Bradford Book</publisher-name>).</citation>
</ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gal</surname> <given-names>Y.</given-names></name> <name><surname>Ghahramani</surname> <given-names>Z.</given-names></name></person-group> (<year>2016</year>). <article-title>Dropout as a bayesian approximation: representing model uncertainty in deep learning</article-title>, in <person-group person-group-type="editor"><name><surname>Balcan</surname> <given-names>M. F.</given-names></name> <name><surname>Weinberger</surname> <given-names>K. Q.</given-names></name></person-group> editors. <source>Proceedings of The 33rd International Conference on Machine Learning, Proceedings of Machine Learning Research</source> Vol. <volume>48</volume>. (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>PMLR</publisher-name>), <fpage>1050</fpage>&#x02013;<lpage>1059</lpage>.</citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gallagher</surname> <given-names>R. J.</given-names></name> <name><surname>Frank</surname> <given-names>M. R.</given-names></name> <name><surname>Mitchell</surname> <given-names>L.</given-names></name> <name><surname>Schwartz</surname> <given-names>A. J.</given-names></name> <name><surname>Reagan</surname> <given-names>A. J.</given-names></name> <name><surname>Danforth</surname> <given-names>C. M.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Generalized word shift graphs: a method for visualizing and explaining pairwise comparisons between texts</article-title>. <source>EPJ Data Sci.</source> <volume>10</volume>:<fpage>4</fpage>. <pub-id pub-id-type="doi">10.1140/epjds/s13688-021-00260-3</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gohil</surname> <given-names>S.</given-names></name> <name><surname>Vuik</surname> <given-names>S.</given-names></name> <name><surname>Darzi</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>Sentiment analysis of health care tweets: review of the methods used</article-title>. <source>JMIR Public Health Surveillance</source> <volume>4</volume>:<fpage>e43</fpage>. <pub-id pub-id-type="doi">10.2196/publichealth.5789</pub-id><pub-id pub-id-type="pmid">29685871</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hamilton</surname> <given-names>W. L.</given-names></name> <name><surname>Clark</surname> <given-names>K.</given-names></name> <name><surname>Leskovec</surname> <given-names>J.</given-names></name> <name><surname>Jurafsky</surname> <given-names>D.</given-names></name></person-group> (<year>2016</year>). <article-title>Inducing domain-specific sentiment lexicons from unlabeled corpora</article-title>, in <source>Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing</source> (<publisher-loc>Austin, TX</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>595</fpage>&#x02013;<lpage>605</lpage>.<pub-id pub-id-type="pmid">28660257</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hansen</surname> <given-names>L. K.</given-names></name> <name><surname>Salamon</surname> <given-names>P.</given-names></name></person-group> (<year>1990</year>). <article-title>Neural network ensembles</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>12</volume>, <fpage>993</fpage>&#x02013;<lpage>1001</lpage>. <pub-id pub-id-type="doi">10.1109/34.58871</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Haque</surname> <given-names>T. U.</given-names></name> <name><surname>Saber</surname> <given-names>N. N.</given-names></name> <name><surname>Shah</surname> <given-names>F. M.</given-names></name></person-group> (<year>2018</year>). <article-title>Sentiment analysis on large scale Amazon product reviews</article-title>, in <source>2018 IEEE international conference on innovative research and development (ICIRD)</source> (<publisher-loc>Bangkok</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>6</lpage>.</citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Harris</surname> <given-names>Z. S.</given-names></name></person-group> (<year>1954</year>). <article-title>Distributional structure</article-title>. <source>Word</source> <volume>10</volume>, <fpage>146</fpage>&#x02013;<lpage>162</lpage>. <pub-id pub-id-type="doi">10.1080/00437956.1954.11659520</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G.</given-names></name> <name><surname>Vinyals</surname> <given-names>O.</given-names></name> <name><surname>Dean</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Distilling the knowledge in a neural network</article-title>, in <source>NIPS Deep Learning and Representation Learning Workshop</source>. <publisher-loc>Montreal, QC</publisher-loc>.</citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hochreiter</surname> <given-names>S.</given-names></name> <name><surname>Schmidhuber</surname> <given-names>J.</given-names></name></person-group> (<year>1997</year>). <article-title>Long short-term memory</article-title>. <source>Neural Comput.</source> <volume>9</volume>, <fpage>1735</fpage>&#x02013;<lpage>1780</lpage>.</citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hoerl</surname> <given-names>A. E.</given-names></name> <name><surname>Kennard</surname> <given-names>R. W.</given-names></name></person-group> (<year>1970</year>). <article-title>Ridge regression: biased estimation for nonorthogonal problems</article-title>. <source>Technometrics</source> <volume>12</volume>, <fpage>55</fpage>&#x02013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.2307/1271436</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hollis</surname> <given-names>G.</given-names></name> <name><surname>Westbury</surname> <given-names>C.</given-names></name></person-group> (<year>2016</year>). <article-title>The principals of meaning: extracting semantic dimensions from co-occurrence models of semantics</article-title>. <source>Psychon. Bull. Rev.</source> <volume>23</volume>, <fpage>1744</fpage>&#x02013;<lpage>1756</lpage>. <pub-id pub-id-type="doi">10.3758/S13423-016-1053-2</pub-id><pub-id pub-id-type="pmid">27138012</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hollis</surname> <given-names>G.</given-names></name> <name><surname>Westbury</surname> <given-names>C.</given-names></name> <name><surname>Lefsrud</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). <article-title>Extrapolating human judgments from skip-gram vector representations of word meaning</article-title>. <source>Quart. J. Exp. Psychol.</source> <volume>70</volume>, <fpage>1603</fpage>&#x02013;<lpage>1619</lpage>. <pub-id pub-id-type="doi">10.1080/17470218.2016.1195417</pub-id><pub-id pub-id-type="pmid">27251936</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hovy</surname> <given-names>D.</given-names></name> <name><surname>Spruit</surname> <given-names>S. L.</given-names></name></person-group> (<year>2016</year>). <article-title>The social impact of natural language processing</article-title>, in <source>Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)</source> <volume>Berlin</volume>, <fpage>591</fpage>&#x02013;<lpage>598</lpage>.</citation>
</ref>
<ref id="B50">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Joulin</surname> <given-names>A.</given-names></name> <name><surname>Grave</surname> <given-names>E.</given-names></name> <name><surname>Bojanowski</surname> <given-names>P.</given-names></name> <name><surname>Mikolov</surname> <given-names>T.</given-names></name></person-group> (<year>2017</year>). <article-title>Bag of tricks for efficient text classification</article-title>, in <source>Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers</source> (<publisher-loc>Valencia</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>427</fpage>&#x02013;<lpage>431</lpage>.</citation>
</ref>
<ref id="B51">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D. P.</given-names></name> <name><surname>Ba</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Adam: a method for stochastic optimization</article-title>, in <person-group person-group-type="editor"><name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>LeCun</surname> <given-names>Y.</given-names></name></person-group> editors, <source>3rd International Conference on Learning Representations, ICLR Conference Track Proceedings</source> (<publisher-loc>San Diego, CA</publisher-loc>: <publisher-name>ICLR</publisher-name>).</citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kiritchenko</surname> <given-names>S.</given-names></name> <name><surname>Zhu</surname> <given-names>X.</given-names></name> <name><surname>Mohammad</surname> <given-names>S. M.</given-names></name></person-group> (<year>2014</year>). <article-title>Sentiment analysis of short informal texts</article-title>. <source>J. Artif. Intell. Res.</source> <volume>50</volume>, <fpage>723</fpage>&#x02013;<lpage>762</lpage>. <pub-id pub-id-type="doi">10.1613/jair.4272</pub-id></citation>
</ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kissler</surname> <given-names>J.</given-names></name> <name><surname>Herbert</surname> <given-names>C.</given-names></name></person-group> (<year>2013</year>). <article-title>Emotion, etmnooi, or emitoon?&#x02014;faster lexical access to emotional than to neutral words during reading</article-title>. <source>Biol. Psychol.</source> <volume>92</volume>, <fpage>464</fpage>&#x02013;<lpage>479</lpage>. <pub-id pub-id-type="doi">10.1016/j.biopsycho.2012.09.004</pub-id><pub-id pub-id-type="pmid">23059636</pub-id></citation></ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koppel</surname> <given-names>M.</given-names></name> <name><surname>Schler</surname> <given-names>J.</given-names></name></person-group> (<year>2006</year>). <article-title>The importance of neutral examples for learning sentiment</article-title>. <source>Comput. Intell.</source> <volume>22</volume>, <fpage>100</fpage>&#x02013;<lpage>109</lpage>. <pub-id pub-id-type="doi">10.1111/j.1467-8640.2006.00276.x</pub-id><pub-id pub-id-type="pmid">23255960</pub-id></citation></ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Korkontzelos</surname> <given-names>I.</given-names></name> <name><surname>Nikfarjam</surname> <given-names>A.</given-names></name> <name><surname>Shardlow</surname> <given-names>M.</given-names></name> <name><surname>Sarker</surname> <given-names>A.</given-names></name> <name><surname>Ananiadou</surname> <given-names>S.</given-names></name> <name><surname>Gonzalez</surname> <given-names>G. H.</given-names></name></person-group> (<year>2016</year>). <article-title>Analysis of the effect of sentiment analysis on extracting adverse drug reactions from tweets and forum posts</article-title>. <source>J. Biomed. Informat.</source> <volume>62</volume>, <fpage>148</fpage>&#x02013;<lpage>158</lpage>. <pub-id pub-id-type="doi">10.1016/j.jbi.2016.06.007</pub-id><pub-id pub-id-type="pmid">27363901</pub-id></citation></ref>
<ref id="B56">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Krogh</surname> <given-names>A.</given-names></name> <name><surname>Vedelsby</surname> <given-names>J.</given-names></name></person-group> (<year>1994</year>). <article-title>Neural network ensembles, cross validation and active learning</article-title>, in <source>Proceedings of the 7th International Conference on Neural Information Processing Systems NIPS&#x00027;94</source> (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>), <fpage>231</fpage>&#x02013;<lpage>238</lpage>.<pub-id pub-id-type="pmid">24246731</pub-id></citation></ref>
<ref id="B57">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kudo</surname> <given-names>T.</given-names></name> <name><surname>Richardson</surname> <given-names>J.</given-names></name></person-group> (<year>2018</year>). <article-title>SentencePiece: a simple and language independent subword tokenizer and detokenizer for neural text processing</article-title>, in <source>Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations</source> (<publisher-loc>Brussels</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>66</fpage>&#x02013;<lpage>71</lpage>.</citation>
</ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kumar</surname> <given-names>A.</given-names></name> <name><surname>Lee</surname> <given-names>C. M.</given-names></name></person-group> (<year>2006</year>). <article-title>Retail investor sentiment and return comovements</article-title>. <source>J. Finance</source> <volume>61</volume>, <fpage>2451</fpage>&#x02013;<lpage>2486</lpage>. <pub-id pub-id-type="doi">10.1111/j.1540-6261.2006.01063.x</pub-id></citation>
</ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lan</surname> <given-names>Z.</given-names></name> <name><surname>Chen</surname> <given-names>M.</given-names></name> <name><surname>Goodman</surname> <given-names>S.</given-names></name> <name><surname>Gimpel</surname> <given-names>K.</given-names></name> <name><surname>Sharma</surname> <given-names>P.</given-names></name> <name><surname>Soricut</surname> <given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>ALBERT: a lite BERT for self-supervised learning of language representations</article-title>, in <source>Proceedings of the International Conference on Learning Representations</source>.</citation>
</ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Laver</surname> <given-names>M.</given-names></name> <name><surname>Benoit</surname> <given-names>K.</given-names></name> <name><surname>Garry</surname> <given-names>J.</given-names></name></person-group> (<year>2003</year>). <article-title>Extracting policy positions from political texts using words as data</article-title>. <source>Amer. Polit. Sci. Rev.</source> <volume>97</volume>, <fpage>311</fpage>&#x02013;<lpage>331</lpage>. <pub-id pub-id-type="doi">10.1017/S0003055403000698</pub-id><pub-id pub-id-type="pmid">30886898</pub-id></citation></ref>
<ref id="B61">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>He</surname> <given-names>L.</given-names></name> <name><surname>Lewis</surname> <given-names>M.</given-names></name> <name><surname>Zettlemoyer</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). <article-title>End-to-end neural coreference resolution</article-title>, in <source>Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing</source> (<publisher-loc>Copenhagen</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>188</fpage>&#x02013;<lpage>197</lpage>.</citation>
</ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>M.</given-names></name> <name><surname>Lu</surname> <given-names>Q.</given-names></name> <name><surname>Long</surname> <given-names>Y.</given-names></name> <name><surname>Gui</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). <article-title>Inferring affective meanings of words from word embedding</article-title>. <source>IEEE Trans. Affect. Comput.</source> <volume>8</volume>, <fpage>443</fpage>&#x02013;<lpage>456</lpage>. <pub-id pub-id-type="doi">10.1109/TAFFC.2017.2723012</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B63">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liao</surname> <given-names>W.</given-names></name> <name><surname>Zeng</surname> <given-names>B.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Wei</surname> <given-names>P.</given-names></name> <name><surname>Cheng</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name></person-group> (<year>2021</year>). <article-title>Multi-level graph neural network for text sentiment analysis</article-title>. <source>Comput. Elect. Eng.</source> <volume>92</volume>:<fpage>107096</fpage>. <pub-id pub-id-type="doi">10.1016/j.compeleceng.2021.107096</pub-id></citation>
</ref>
<ref id="B64">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ljube&#x00161;i&#x00107;</surname> <given-names>N.</given-names></name> <name><surname>Fi&#x00161;er</surname> <given-names>D.</given-names></name> <name><surname>Peti-Stanti&#x00107;</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>Predicting concreteness and imageability of words within and across languages via word embeddings</article-title>, in <source>Proceedings of The Third Workshop on Representation Learning for NLP</source> (<publisher-loc>Melbourne, VIC</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>217</fpage>&#x02013;<lpage>222</lpage>.</citation>
</ref>
<ref id="B65">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Maas</surname> <given-names>A. L.</given-names></name> <name><surname>Daly</surname> <given-names>R. E.</given-names></name> <name><surname>Pham</surname> <given-names>P. T.</given-names></name> <name><surname>Huang</surname> <given-names>D.</given-names></name> <name><surname>Ng</surname> <given-names>A. Y.</given-names></name> <name><surname>Potts</surname> <given-names>C.</given-names></name></person-group> (<year>2011</year>). <article-title>Learning word vectors for sentiment analysis</article-title>, in <source>Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies</source> (<publisher-loc>Portland, OR</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>142</fpage>&#x02013;<lpage>150</lpage>.</citation>
</ref>
<ref id="B66">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mayzner</surname> <given-names>M. S.</given-names></name> <name><surname>Tresselt</surname> <given-names>M. E.</given-names></name></person-group> (<year>1965</year>). <article-title>Tables of single-letter and digram frequency counts for various word-length and letter-position combinations</article-title>. <source>Psychonomic Monograph Supplements</source> <volume>1</volume>, <fpage>13</fpage>&#x02013;<lpage>32</lpage>.</citation>
</ref>
<ref id="B67">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Medhat</surname> <given-names>W.</given-names></name> <name><surname>Hassan</surname> <given-names>A.</given-names></name> <name><surname>Korashy</surname> <given-names>H.</given-names></name></person-group> (<year>2014</year>). <article-title>Sentiment analysis algorithms and applications: a survey</article-title>. <source>Ain Shams Eng. J.</source> <volume>5</volume>, <fpage>1093</fpage>&#x02013;<lpage>1113</lpage>. <pub-id pub-id-type="doi">10.1016/j.asej.2014.04.011</pub-id></citation>
</ref>
<ref id="B68">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mikolov</surname> <given-names>T.</given-names></name> <name><surname>Chen</surname> <given-names>K.</given-names></name> <name><surname>Corrado</surname> <given-names>G.</given-names></name> <name><surname>Dean</surname> <given-names>J.</given-names></name></person-group> (<year>2013a</year>).<article-title> Efficient estimation of word representations in vector space</article-title>. In <person-group person-group-type="editor"><name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>LeCun</surname> <given-names>Y.</given-names></name></person-group> editors, <source>1st International Conference on Learning Representations, ICLR 2013, Scottsdale, Arizona, USA, May 2-4, 2013, Workshop Track Proceedings</source>.</citation>
</ref>
<ref id="B69">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mikolov</surname> <given-names>T.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Chen</surname> <given-names>K.</given-names></name> <name><surname>Corrado</surname> <given-names>G.</given-names></name> <name><surname>Dean</surname> <given-names>J.</given-names></name></person-group> (<year>2013b</year>). <article-title>Distributed representations of words and phrases and their compositionality</article-title>, in <source>Proceedings of the 26th International Conference on Neural Information Processing Systems - Volume 2, NIPS&#x00027;13</source> (<publisher-loc>Red Hook, NY</publisher-loc>: <publisher-name>Curran Associates Inc</publisher-name>.), <fpage>3111</fpage>&#x02013;<lpage>3119</lpage>.<pub-id pub-id-type="pmid">31840584</pub-id></citation></ref>
<ref id="B70">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Miller</surname> <given-names>G. A.</given-names></name> <name><surname>Newman</surname> <given-names>E. B.</given-names></name> <name><surname>Friedman</surname> <given-names>E. A.</given-names></name></person-group> (<year>1958</year>). <article-title>Length-frequency statistics for written English</article-title>. <source>Inf. Control</source> <volume>1</volume>, <fpage>370</fpage>&#x02013;<lpage>389</lpage>. <pub-id pub-id-type="doi">10.1016/S0019-9958(58)90229-8</pub-id></citation>
</ref>
<ref id="B71">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mohammad</surname> <given-names>S.</given-names></name></person-group> (<year>2018</year>). <article-title>Obtaining reliable human ratings of valence, arousal, and dominance for 20,000 English words</article-title>, in <source>Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source> (<publisher-loc>Melbourne, VIC</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>174</fpage>&#x02013;<lpage>184</lpage>.</citation>
</ref>
<ref id="B72">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nasukawa</surname> <given-names>T.</given-names></name> <name><surname>Yi</surname> <given-names>J.</given-names></name></person-group> (<year>2003</year>). <article-title>Sentiment analysis: capturing favorability using natural language processing</article-title>, in <source>Proceedings of the 2nd International Conference on Knowledge Capture</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>70</fpage>&#x02013;<lpage>77</lpage>.</citation>
</ref>
<ref id="B73">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Osgood</surname> <given-names>C. E.</given-names></name></person-group> (<year>1962</year>). <article-title>Studies on the generality of affective meaning systems</article-title>. <source>Amer. Psychol.</source> <volume>17</volume>:<fpage>10</fpage>. <pub-id pub-id-type="doi">10.1037/h0045146</pub-id></citation>
</ref>
<ref id="B74">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pak</surname> <given-names>A.</given-names></name> <name><surname>Paroubek</surname> <given-names>P.</given-names></name></person-group> (<year>2010</year>). <article-title>Twitter as a corpus for sentiment analysis and opinion mining</article-title>, in <source>Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC&#x00027;10)</source> <publisher-loc>Valletta</publisher-loc>: <publisher-name>European Language Resources Association (ELRA)</publisher-name>.</citation>
</ref>
<ref id="B75">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pang</surname> <given-names>B.</given-names></name> <name><surname>Lee</surname> <given-names>L.</given-names></name></person-group> (<year>2008</year>). <article-title>Opinion mining and sentiment analysis</article-title>. <source>Found. Trends Inf. Retrieval</source> <volume>2</volume>, <fpage>1</fpage>&#x02013;<lpage>135</lpage>. <pub-id pub-id-type="doi">10.1561/1500000011</pub-id></citation>
</ref>
<ref id="B76">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pang</surname> <given-names>B.</given-names></name> <name><surname>Lee</surname> <given-names>L.</given-names></name> <name><surname>Vaithyanathan</surname> <given-names>S.</given-names></name></person-group> (<year>2002</year>). <article-title>Thumbs up? sentiment classification using machine learning techniques</article-title>, in <source>Proceedings of the ACL-02 Conference on Empirical Methods in Natural Language Processing - Volume 10 EMNLP &#x00027;02</source> (<publisher-loc>Philadelphia, PA</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>79</fpage>&#x02013;<lpage>86</lpage></citation>
</ref>
<ref id="B77">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pennington</surname> <given-names>J.</given-names></name> <name><surname>Socher</surname> <given-names>R.</given-names></name> <name><surname>Manning</surname> <given-names>C.</given-names></name></person-group> (<year>2014</year>). <article-title>GloVe: Global vectors for word representation</article-title>, in <source>Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source> (<publisher-loc>Doha</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>1532</fpage>&#x02013;<lpage>1543</lpage>.</citation>
</ref>
<ref id="B78">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Peters</surname> <given-names>M.</given-names></name> <name><surname>Ammar</surname> <given-names>W.</given-names></name> <name><surname>Bhagavatula</surname> <given-names>C.</given-names></name> <name><surname>Power</surname> <given-names>R.</given-names></name></person-group> (<year>2017</year>). <article-title>Semi-supervised sequence tagging with bidirectional language models</article-title>, in <source>Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source> (<publisher-loc>Vancouver, BC</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>1756</fpage>&#x02013;<lpage>176</lpage>.</citation>
</ref>
<ref id="B79">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Peters</surname> <given-names>M.</given-names></name> <name><surname>Neumann</surname> <given-names>M.</given-names></name> <name><surname>Iyyer</surname> <given-names>M.</given-names></name> <name><surname>Gardner</surname> <given-names>M.</given-names></name> <name><surname>Clark</surname> <given-names>C.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>Zettlemoyer</surname> <given-names>L.</given-names></name></person-group> (<year>2018</year>). <article-title>Deep contextualized word representations</article-title>, in <source>Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)</source> (<publisher-loc>New Orleans</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>2227</fpage>&#x02013;<lpage>2237</lpage>.<pub-id pub-id-type="pmid">34343877</pub-id></citation></ref>
<ref id="B80">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Qiu</surname> <given-names>G.</given-names></name> <name><surname>Liu</surname> <given-names>B.</given-names></name> <name><surname>Bu</surname> <given-names>J.</given-names></name> <name><surname>Chen</surname> <given-names>C.</given-names></name></person-group> (<year>2009</year>). <article-title>Expanding domain sentiment lexicon through double propagation</article-title>, in <source>Proceedings of the 21st International Joint Conference on Artificial Intelligence IJCAI&#x00027;09</source> (<publisher-loc>Pasadena, CA</publisher-loc>), <fpage>1199</fpage>&#x02013;<lpage>1204</lpage>.</citation>
</ref>
<ref id="B81">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Radford</surname> <given-names>A.</given-names></name> <name><surname>Narasimhan</surname> <given-names>K.</given-names></name> <name><surname>Salimans</surname> <given-names>T.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name></person-group> (<year>2018</year>). <source>Improving language understanding by generative pre-training</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://www.cs.ubc.ca/&#x0007E;amuham01/LING530/papers/radford2018improving.pdf">https://www.cs.ubc.ca/&#x0007E;amuham01/LING530/papers/radford2018improving.pdf</ext-link></citation>
</ref>
<ref id="B82">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Reagan</surname> <given-names>A. J.</given-names></name> <name><surname>Danforth</surname> <given-names>C. M.</given-names></name> <name><surname>Tivnan</surname> <given-names>B.</given-names></name> <name><surname>Williams</surname> <given-names>J. R.</given-names></name> <name><surname>Dodds</surname> <given-names>P. S.</given-names></name></person-group> (<year>2017</year>). <article-title>Sentiment analysis methods for understanding large-scale texts: a case for using continuum-scored words and word shift graphs</article-title>. <source>EPJ Data Sci.</source> <volume>6</volume>, <fpage>1</fpage>&#x02013;<lpage>21</lpage>. <pub-id pub-id-type="doi">10.1140/epjds/s13688-017-0121-9</pub-id></citation>
</ref>
<ref id="B83">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ribeiro</surname> <given-names>F. N.</given-names></name> <name><surname>Ara&#x000FA;jo</surname> <given-names>M.</given-names></name> <name><surname>Gon&#x000E7;alves</surname> <given-names>P.</given-names></name> <name><surname>Gon&#x000E7;alves</surname> <given-names>M. A.</given-names></name> <name><surname>Benevenuto</surname> <given-names>F.</given-names></name></person-group> (<year>2016</year>). <article-title>SentiBench - a benchmark comparison of state-of-the-practice sentiment analysis methods</article-title>. <source>EPJ Data Sci.</source> <volume>5</volume>, <fpage>1</fpage>&#x02013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.1140/epjds/s13688-016-0085-1</pub-id></citation>
</ref>
<ref id="B84">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Riloff</surname> <given-names>E.</given-names></name></person-group> (<year>1996</year>). <article-title>An empirical study of automated dictionary construction for information extraction in three domains</article-title>. <source>Artif. Intell.</source> <volume>85</volume>, <fpage>101</fpage>&#x02013;<lpage>134</lpage>. <pub-id pub-id-type="doi">10.1016/0004-3702(95)00123-9</pub-id></citation>
</ref>
<ref id="B85">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rumelhart</surname> <given-names>D. E.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Williams</surname> <given-names>R. J.</given-names></name></person-group> (<year>1986</year>). <source>Learning Internal Representations by Error Propagation</source>. (<publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>MIT Press</publisher-name>). <fpage>318</fpage>&#x02013;<lpage>362</lpage>.</citation>
</ref>
<ref id="B86">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>San Vicente</surname> <given-names>I.</given-names></name> <name><surname>Agerri</surname> <given-names>R.</given-names></name> <name><surname>Rigau</surname> <given-names>G.</given-names></name></person-group> (<year>2014</year>). <article-title>Simple, robust and (almost) unsupervised generation of polarity lexicons for multiple languages</article-title>, in <source>Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics</source> (<publisher-loc>Gothenburg</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>88</fpage>&#x02013;<lpage>97</lpage>.</citation>
</ref>
<ref id="B87">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sanh</surname> <given-names>V.</given-names></name> <name><surname>Debut</surname> <given-names>L.</given-names></name> <name><surname>Chaumond</surname> <given-names>J.</given-names></name> <name><surname>Wolf</surname> <given-names>T.</given-names></name></person-group> (<year>2019</year>). <article-title>DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter</article-title>, in <source>Proceedings of the 7th International Conference on Neural Information Processing Systems, 5th Workshop on Energy Efficient Machine Learning and Cognitive Computing</source> <publisher-loc>Vancouver, BC</publisher-loc>: <publisher-name>MIT Press</publisher-name>.</citation>
</ref>
<ref id="B88">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Sennrich</surname> <given-names>R.</given-names></name> <name><surname>Haddow</surname> <given-names>B.</given-names></name> <name><surname>Birch</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>Neural machine translation of rare words with subword units</article-title>, in <source>Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>1715</fpage>&#x02013;<lpage>1725</lpage>.</citation>
</ref>
<ref id="B89">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Shmueli</surname> <given-names>B.</given-names></name> <name><surname>Fell</surname> <given-names>J.</given-names></name> <name><surname>Ray</surname> <given-names>S.</given-names></name> <name><surname>Ku</surname> <given-names>L.-W.</given-names></name></person-group> (<year>2021</year>). <article-title>Beyond fair pay: Ethical implications of NLP crowdsourcing</article-title>, in <source>Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source> (<publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>3758</fpage>&#x02013;<lpage>3769</lpage>.</citation>
</ref>
<ref id="B90">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Snyder</surname> <given-names>B.</given-names></name> <name><surname>Barzilay</surname> <given-names>R.</given-names></name></person-group> (<year>2007</year>). <article-title>Multiple aspect ranking using the good grief algorithm</article-title>, in <source>Human Language Technologies 2007: The Conference of the North American Chapter of the Association for Computational Linguistics; Proceedings of the Main Conference</source> (<publisher-loc>Rochester, NY</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>300</fpage>&#x02013;<lpage>307</lpage></citation>
</ref>
<ref id="B91">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Socher</surname> <given-names>R.</given-names></name> <name><surname>Chen</surname> <given-names>D.</given-names></name> <name><surname>Manning</surname> <given-names>C. D.</given-names></name> <name><surname>Ng</surname> <given-names>A.</given-names></name></person-group> (<year>2013a</year>). <article-title>Reasoning with neural tensor networks for knowledge base completion</article-title>, in <source>Advances in Neural Information Processing Systems</source> (<publisher-loc>Lake Tahoe, NV</publisher-loc>: <publisher-name>Curran Associates Inc</publisher-name>.), <fpage>926</fpage>&#x02013;<lpage>934</lpage>.</citation>
</ref>
<ref id="B92">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Socher</surname> <given-names>R.</given-names></name> <name><surname>Perelygin</surname> <given-names>A.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name> <name><surname>Chuang</surname> <given-names>J.</given-names></name> <name><surname>Manning</surname> <given-names>C. D.</given-names></name> <name><surname>Ng</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2013b</year>). <article-title>Recursive deep models for semantic compositionality over a sentiment treebank</article-title>, in <source>Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing</source> (<publisher-loc>Seattle, WA</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>1631</fpage>&#x02013;<lpage>1642</lpage>.</citation>
</ref>
<ref id="B93">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Srivastava</surname> <given-names>N.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name> <name><surname>Krizhevsky</surname> <given-names>A.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <article-title>Dropout: a simple way to prevent neural networks from overfitting</article-title>. <source>J. Mach. Learn. Res.</source> <volume>15</volume>, <fpage>1929</fpage>&#x02013;<lpage>1958</lpage>. <pub-id pub-id-type="doi">10.5555/2627435.2670313</pub-id><pub-id pub-id-type="pmid">26215079</pub-id></citation></ref>
<ref id="B94">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Stupinski</surname> <given-names>A. M.</given-names></name> <name><surname>Alshaabi</surname> <given-names>T.</given-names></name> <name><surname>Arnold</surname> <given-names>M. V.</given-names></name> <name><surname>Adams</surname> <given-names>J. L.</given-names></name> <name><surname>Minot</surname> <given-names>J. R.</given-names></name> <name><surname>Price</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2021</year>). <source>Quantifying Language Changes Surrounding Mental Health on Twitter</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/2106.01481">https://arxiv.org/abs/2106.01481</ext-link></citation>
</ref>
<ref id="B95">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Taboada</surname> <given-names>M.</given-names></name> <name><surname>Brooke</surname> <given-names>J.</given-names></name> <name><surname>Tofiloski</surname> <given-names>M.</given-names></name> <name><surname>Voll</surname> <given-names>K.</given-names></name> <name><surname>Stede</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <article-title>Lexicon-based methods for sentiment analysis</article-title>. <source>Comput. Linguist.</source> <volume>37</volume>, <fpage>267</fpage>&#x02013;<lpage>307</lpage>. <pub-id pub-id-type="doi">10.1162/COLI_a_00049</pub-id></citation>
</ref>
<ref id="B96">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tang</surname> <given-names>D.</given-names></name> <name><surname>Wei</surname> <given-names>F.</given-names></name> <name><surname>Qin</surname> <given-names>B.</given-names></name> <name><surname>Zhou</surname> <given-names>M.</given-names></name> <name><surname>Liu</surname> <given-names>T.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x00300;Building large-scale Twitter-specific sentiment lexicon: A representation learning approach</article-title>, in <source>Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers</source> (<publisher-loc>Dublin</publisher-loc>: <publisher-name>Dublin City University and Association for Computational Linguistics</publisher-name>), <fpage>172</fpage>&#x02013;<lpage>182</lpage>.</citation>
</ref>
<ref id="B97">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tang</surname> <given-names>H.</given-names></name> <name><surname>Tan</surname> <given-names>S.</given-names></name> <name><surname>Cheng</surname> <given-names>X.</given-names></name></person-group> (<year>2009</year>). <article-title>A survey on sentiment detection of reviews</article-title>. <source>Exp. Syst. Appl.</source> <volume>36</volume>, <fpage>10760</fpage>&#x02013;<lpage>10773</lpage>. <pub-id pub-id-type="doi">10.1016/j.eswa.2009.02.063</pub-id></citation>
</ref>
<ref id="B98">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tatman</surname> <given-names>R.</given-names></name></person-group> (<year>2017</year>). <article-title>Gender and dialect bias in YouTube&#x00027;s automatic captions</article-title>, in <source>Proceedings of the First ACL Workshop on Ethics in Natural Language Processing</source> (<publisher-loc>Valencia</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>53</fpage>&#x02013;<lpage>59</lpage>.</citation>
</ref>
<ref id="B99">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Terveen</surname> <given-names>L.</given-names></name> <name><surname>Hill</surname> <given-names>W.</given-names></name> <name><surname>Amento</surname> <given-names>B.</given-names></name> <name><surname>McDonald</surname> <given-names>D.</given-names></name> <name><surname>Creter</surname> <given-names>J.</given-names></name></person-group> (<year>1997</year>). <article-title>PHOAKS: a system for sharing recommendations</article-title>. <source>Commun. ACM</source> <volume>40</volume>, <fpage>59</fpage>&#x02013;<lpage>62</lpage>. <pub-id pub-id-type="doi">10.1145/245108.245122</pub-id></citation>
</ref>
<ref id="B100">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Thavareesan</surname> <given-names>S.</given-names></name> <name><surname>Mahesan</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Sentiment lexicon expansion using word2vec and fastText for sentiment prediction in Tamil texts</article-title>, in <source>2020 Moratuwa Engineering Research Conference (MERCon)</source> (<publisher-loc>Moratuwa</publisher-loc>), <fpage>272</fpage>&#x02013;<lpage>276</lpage>.</citation>
</ref>
<ref id="B101">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thelwall</surname> <given-names>M.</given-names></name> <name><surname>Buckley</surname> <given-names>K.</given-names></name> <name><surname>Paltoglou</surname> <given-names>G.</given-names></name> <name><surname>Cai</surname> <given-names>D.</given-names></name> <name><surname>Kappas</surname> <given-names>A.</given-names></name></person-group> (<year>2010</year>). <article-title>Sentiment strength detection in short informal text</article-title>. <source>J. Amer. Soc. Inf. Sci. Technol.</source> <volume>61</volume>, <fpage>2544</fpage>&#x02013;<lpage>2558</lpage>. <pub-id pub-id-type="doi">10.1002/asi.21416</pub-id><pub-id pub-id-type="pmid">25855820</pub-id></citation></ref>
<ref id="B102">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Thomas</surname> <given-names>M.</given-names></name> <name><surname>Pang</surname> <given-names>B.</given-names></name> <name><surname>Lee</surname> <given-names>L.</given-names></name></person-group> (<year>2006</year>). <article-title>Get out the vote: determining support or opposition from congressional floor-debate transcripts</article-title>, in <source>Proceedings of the 2006 Conference on Empirical Methods in Natural Language Processing</source> (<publisher-loc>Sydney, NSW</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>327</fpage>&#x02013;<lpage>335</lpage>.</citation>
</ref>
<ref id="B103">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Thompson</surname> <given-names>N. C.</given-names></name> <name><surname>Greenewald</surname> <given-names>K.</given-names></name> <name><surname>Lee</surname> <given-names>K.</given-names></name> <name><surname>Manso</surname> <given-names>G. F.</given-names></name></person-group> (<year>2020</year>). <source>The Computational Limits of Deep Learning</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/2007.05558">https://arxiv.org/abs/2007.05558</ext-link></citation>
</ref>
<ref id="B104">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tumasjan</surname> <given-names>A.</given-names></name> <name><surname>Sprenger</surname> <given-names>T.</given-names></name> <name><surname>Sandner</surname> <given-names>P.</given-names></name> <name><surname>Welpe</surname> <given-names>I.</given-names></name></person-group> (<year>2010</year>). <article-title>Predicting elections with Twitter: what 140 characters reveal about political sentiment</article-title>, in <source>Proceedings of the International AAAI Conference on Web and Social Media</source>, vol. <volume>4</volume>, <publisher-loc>Washington, DC</publisher-loc>.</citation>
</ref>
<ref id="B105">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Turney</surname> <given-names>P. D.</given-names></name></person-group> (<year>2002</year>). <article-title>Thumbs up or thumbs down? Semantic orientation applied to unsupervised classification of reviews</article-title>, in <source>Proceedings of the 40th Annual Meeting on Association for Computational Linguistics ACL &#x00027;02</source> (<publisher-loc>Philadelphia, PA</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>417</fpage>&#x02013;<lpage>424</lpage>.</citation>
</ref>
<ref id="B106">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Turney</surname> <given-names>P. D.</given-names></name> <name><surname>Littman</surname> <given-names>M. L.</given-names></name></person-group> (<year>2003</year>). <article-title>Measuring praise and criticism: Inference of semantic orientation from association</article-title>. <source>ACM Trans. Inf. Syst.</source> <volume>21</volume>, <fpage>315</fpage>&#x02013;<lpage>346</lpage>. <pub-id pub-id-type="doi">10.1145/944012.944013</pub-id></citation>
</ref>
<ref id="B107">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Vaswani</surname> <given-names>A.</given-names></name> <name><surname>Shazeer</surname> <given-names>N.</given-names></name> <name><surname>Parmar</surname> <given-names>N.</given-names></name> <name><surname>Uszkoreit</surname> <given-names>J.</given-names></name> <name><surname>Jones</surname> <given-names>L.</given-names></name> <name><surname>Gomez</surname> <given-names>A. N.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Attention is all you need</article-title>, in <source>Advances in Neural Information Processing Systems</source>, eds <person-group person-group-type="editor"><name><surname>Guyon</surname> <given-names>I.</given-names></name><name><surname>Luxburg</surname> <given-names>U. V.</given-names></name> <name><surname>Bengio</surname> <given-names>S.</given-names></name> <name><surname>Wallach</surname> <given-names>H.</given-names></name> <name><surname>Fergus</surname> <given-names>R.</given-names></name> <name><surname>Vishwanathan</surname> <given-names>S.</given-names></name> <name><surname>Garnett</surname> <given-names>R.</given-names></name></person-group> Vol. <volume>30</volume>. (<publisher-loc>Long Beach, CA</publisher-loc>: <publisher-name>Curran Associates, Inc</publisher-name>.).</citation>
</ref>
<ref id="B108">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Yu</surname> <given-names>L.-C.</given-names></name> <name><surname>Lai</surname> <given-names>K. R.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name></person-group> (<year>2016</year>). <article-title>Community-based weighted graph model for valence-arousal prediction of affective words</article-title>. <source>IEEE/ACM Trans. Audio Speech Lang. Process.</source> <volume>24</volume>, <fpage>1957</fpage>&#x02013;<lpage>1968</lpage>. <pub-id pub-id-type="doi">10.1109/TASLP.2016.2594287</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B109">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wilson</surname> <given-names>T.</given-names></name> <name><surname>Wiebe</surname> <given-names>J.</given-names></name> <name><surname>Hoffmann</surname> <given-names>P.</given-names></name></person-group> (<year>2005</year>). <article-title>Recognizing contextual polarity in phrase-level sentiment analysis</article-title>, in <source>Proceedings of the Conference on Human Language Technology and Empirical Methods in Natural Language Processing HLT &#x00027;05</source> (<publisher-loc>Vancouver, BC</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>), <fpage>347</fpage>&#x02013;<lpage>354</lpage>.</citation>
</ref>
<ref id="B110">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wolf</surname> <given-names>T.</given-names></name> <name><surname>Debut</surname> <given-names>L.</given-names></name> <name><surname>Sanh</surname> <given-names>V.</given-names></name> <name><surname>Chaumond</surname> <given-names>J.</given-names></name> <name><surname>Delangue</surname> <given-names>C.</given-names></name> <name><surname>Moi</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Transformers: state-of-the-art natural language processing</article-title>, in <source>Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations</source> (<publisher-loc>Association for Computational Linguistics</publisher-loc>), <fpage>38</fpage>&#x02013;<lpage>45</lpage>.</citation>
</ref>
<ref id="B111">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>Y.</given-names></name> <name><surname>Schuster</surname> <given-names>M.</given-names></name> <name><surname>Chen</surname> <given-names>Z.</given-names></name> <name><surname>Le</surname> <given-names>Q. V.</given-names></name> <name><surname>Norouzi</surname> <given-names>M.</given-names></name> <name><surname>Macherey</surname> <given-names>W.</given-names></name> <etal/></person-group>. (<year>2016</year>). <source>Google&#x00027;s Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1609.08144">https://arxiv.org/abs/1609.08144</ext-link></citation>
</ref>
<ref id="B112">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yadollahi</surname> <given-names>A.</given-names></name> <name><surname>Shahraki</surname> <given-names>A. G.</given-names></name> <name><surname>Zaiane</surname> <given-names>O. R.</given-names></name></person-group> (<year>2017</year>). <article-title>Current state of text sentiment analysis from opinion to emotion mining</article-title>. <source>ACM Comput. Surveys</source> <volume>50</volume>, <fpage>1</fpage>&#x02013;<lpage>33</lpage>. <pub-id pub-id-type="doi">10.1145/3057270</pub-id></citation>
</ref>
<ref id="B113">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>S.</given-names></name> <name><surname>Xing</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Chang</surname> <given-names>Z.</given-names></name></person-group> (<year>2021</year>). <article-title>Implicit sentiment analysis based on graph attention neural network</article-title>. <source>Eng. Rep.</source> <volume>2012</volume>:<fpage>e12452</fpage>. <pub-id pub-id-type="doi">10.1002/eng2.12452</pub-id><pub-id pub-id-type="pmid">25855820</pub-id></citation></ref>
<ref id="B114">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Z.</given-names></name> <name><surname>Dai</surname> <given-names>Z.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Carbonell</surname> <given-names>J.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R. R.</given-names></name> <name><surname>Le</surname> <given-names>Q. V.</given-names></name></person-group> (<year>2019</year>). <article-title>Xlnet: generalized autoregressive pretraining for language understanding</article-title>, in <source>Advances in Neural Information Processing Systems</source>, Vol. <volume>32</volume>. (<publisher-loc>Vancouver, BC</publisher-loc>: <publisher-name>Curran Associates, Inc</publisher-name>.).<pub-id pub-id-type="pmid">32723719</pub-id></citation></ref>
<ref id="B115">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>Y.</given-names></name> <name><surname>Duan</surname> <given-names>W.</given-names></name> <name><surname>Cao</surname> <given-names>Q.</given-names></name></person-group> (<year>2013</year>). <article-title>The impact of social and conventional media on firm equity value: a sentiment analysis approach</article-title>. <source>Decis. Support Syst.</source> <volume>55</volume>, <fpage>919</fpage>&#x02013;<lpage>926</lpage>. <pub-id pub-id-type="doi">10.1016/j.dss.2012.12.028</pub-id></citation>
</ref>
<ref id="B116">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yuan</surname> <given-names>M.</given-names></name> <name><surname>Lin</surname> <given-names>Y.</given-names></name></person-group> (<year>2006</year>). <article-title>Model selection and estimation in regression with grouped variables</article-title>. <source>J. Roy. Stat. Soc. B (Stati. Methodol.)</source> <volume>68</volume>, <fpage>49</fpage>&#x02013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.1111/J.1467-9868.2005.00532.X</pub-id><pub-id pub-id-type="pmid">28192605</pub-id></citation></ref>
<ref id="B117">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zou</surname> <given-names>H.</given-names></name> <name><surname>Hastie</surname> <given-names>T.</given-names></name></person-group> (<year>2005</year>). <article-title>Regularization and variable selection via the elastic net</article-title>. <source>J. Roy. Stat. Soc. B (Stati. Methodol.)</source> <volume>67</volume>, <fpage>301</fpage>&#x02013;<lpage>320</lpage>. <pub-id pub-id-type="doi">10.1111/j.1467-9868.2005.00503.x</pub-id></citation>
</ref>
</ref-list> 
</back>
</article>