<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comput. Neurosci.</journal-id>
<journal-title>Frontiers in Computational Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comput. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-5188</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fncom.2024.1350916</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Prediction of emotion distribution of images based on weighted <italic>K</italic>-nearest neighbor-attention mechanism</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Cheng</surname> <given-names>Kai</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2597103/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
</contrib>
</contrib-group>
<aff><institution>School of Artificial Intelligence, Xidian University</institution>, <addr-line>Xi&#x00027;an</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Deepika Koundal, University of Petroleum and Energy Studies, India</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Arvind Dhaka, Manipal University Jaipur, India</p>
<p>Mohit Mittal, Institut National de Recherche en Informatique et en Automatique (INRIA), France</p>
<p>Tariq Ahmad, Guilin University of Electronic Technology, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Kai Cheng <email>chengkai7300&#x00040;163.com</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>17</day>
<month>04</month>
<year>2024</year>
</pub-date>
<pub-date pub-type="collection">
<year>2024</year>
</pub-date>
<volume>18</volume>
<elocation-id>1350916</elocation-id>
<history>
<date date-type="received">
<day>06</day>
<month>12</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>28</day>
<month>03</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2024 Cheng.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Cheng</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Existing methods for classifying image emotions often overlook the subjective impact emotions evoke in observers, focusing primarily on emotion categories. However, this approach falls short in meeting practical needs as it neglects the nuanced emotional responses captured within an image. This study proposes a novel approach employing the weighted closest neighbor algorithm to predict the discrete distribution of emotion in abstract paintings. Initially, emotional features are extracted from the images and assigned varying <italic>K</italic>-values. Subsequently, an encoder-decoder architecture is utilized to derive sentiment features from abstract paintings, augmented by a pre-trained model to enhance classification model generalization and convergence speed. By incorporating a blank attention mechanism into the decoder and integrating it with the encoder&#x00027;s output sequence, the semantics of abstract painting images are learned, facilitating precise and sensible emotional understanding. Experimental results demonstrate that the classification algorithm, utilizing the attention mechanism, achieves a higher accuracy of 80.7% compared to current methods. This innovative approach successfully addresses the intricate challenge of discerning emotions in abstract paintings, underscoring the significance of considering subjective emotional responses in image classification. The integration of advanced techniques such as weighted closest neighbor algorithm and attention mechanisms holds promise for enhancing the comprehension and classification of emotional content in visual art.</p></abstract>
<kwd-group>
<kwd>image emotions</kwd>
<kwd>classification</kwd>
<kwd>weighted closest neighbor algorithm</kwd>
<kwd>emotional features</kwd>
<kwd>abstract paintings</kwd>
</kwd-group>
<counts>
<fig-count count="9"/>
<table-count count="4"/>
<equation-count count="15"/>
<ref-count count="33"/>
<page-count count="14"/>
<word-count count="8056"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1 Introduction</title>
<p>Image data are essentially used for transferring information. The amount of picture data is even increasing at an exponential speed owing to the advent of the Internet (Cetinic and She, <xref ref-type="bibr" rid="B6">2022</xref>; Zou et al., <xref ref-type="bibr" rid="B33">2023</xref>). Because of the fast-paced nature of modern society, people&#x00027;s ability to extract information from photos is also accelerating, necessitating more accuracy and efficiency in identifying image data on the network. Based on this necessity, an effective image processing technique that makes use of computer vision is required for humans to manage and use picture data more effectively.</p>
<p>Sentiment analysis, often called opinion mining, is the process of using natural language processing, text analysis, computational linguistics, and biometrics to systematically unpack subjective information and emotional states. The notion was initially introduced by Yang et al. (<xref ref-type="bibr" rid="B27">2023</xref>). Sentiment analysis has gained significant economic and societal significance in the last several years and has been applied extensively in the domains of opinion monitoring (Chen et al., <xref ref-type="bibr" rid="B9">2023</xref>), topic inference (Ngai et al., <xref ref-type="bibr" rid="B17">2022</xref>), and comment analysis and decision-making (Bharadiya, <xref ref-type="bibr" rid="B5">2023</xref>). For monitoring public opinion, the government can make timely policy interventions and accurately determine the direction of public opinion. When it comes to product recommendations, merchants can better understand user needs and suggestions by gauging user satisfaction with product evaluations and enhancing product quality. In the finance domain, trending financial topics can even be used to predict stock direction. Furthermore, sentiment analysis is frequently used for various tasks involving natural language processing. To increase the accuracy of the system, more exact terms for sentiment expression are chosen for machine translation (Chan et al., <xref ref-type="bibr" rid="B7">2023</xref>) by evaluating the sentiment tendency of the input text. The pixel density extraction of the image information is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Example of emotional distribution in images.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-18-1350916-g0001.tif"/>
</fig>
<p>Various classification techniques will be broken down into different levels for the sentiment analysis task: output results will categorize the methods into sentiment intensity classification and sentiment polarity classification; granularity of the processed text will divide them into three research levels: word level, sentence level, and chapter level; research methodology will separate them into unsupervised learning, semi-supervised learning, and supervised learning, and so on. The majority of the conventional sentiment classification algorithms employ manually created feature selection techniques for feature extraction, such as the maximum entropy model (Chandrasekaran et al., <xref ref-type="bibr" rid="B8">2022</xref>), plain Bayes (Wang et al., <xref ref-type="bibr" rid="B26">2022</xref>), support vector machines (Zhao et al., <xref ref-type="bibr" rid="B30">2021a</xref>), and so on. However, these techniques have limitations, such as being labor-intensive, time-consuming, and hard to train. As a result, they are not well-suited for use in the current large-scale application scenarios.</p>
<p>With advancements in machine learning, research efforts (Milani and Fraternali, <xref ref-type="bibr" rid="B16">2021</xref>) led to the development of deep learning methods that give neural networks a hierarchical structure. This development subsequently resulted in an explosion of deep learning research. Feature learning, at the heart of deep learning, uses hierarchical networks to convert unprocessed input into more abstract and higher-level feature information. With its superior learning capacity to optimize automated feature extraction, deep learning has produced remarkable research achievements in recent years in the domains of speech recognition, picture processing, and natural language processing. The application of deep learning techniques to text sentiment analysis has gained popularity as a natural language processing study area. Among these techniques, Song et al. (<xref ref-type="bibr" rid="B22">2021</xref>) used a convolutional neural network to classify text emotion for the first time, and the results were superior to those of conventional machine learning techniques.</p>
<p>The study of human eyesight is where attention mechanism first emerged. According to cognitive science, humans have a tendency to ignore other observable information in favor of focusing on a certain portion of the information based on the demand imposed by the information processing bottleneck. The primary objective of attention mechanism is to efficiently separate valuable information from a vast quantity of data. To understand the word dependencies inside the phrase and grasp the internal structure of the sentence, the self-attention mechanism&#x02014;a unique form of attention mechanism&#x02014;is incorporated into the sentiment classification job. To establish an accurate and efficient technique for sentiment analysis based on deep learning technology and self-attention mechanism, this study examines the present technical issues in the field of sentiment analysis from the standpoint of the real demands of sentiment analysis.</p></sec>
<sec id="s2">
<title>2 Related studies</title>
<p>Natural language processing has attracted extensive research attention (McCormack and Lomas, <xref ref-type="bibr" rid="B15">2021</xref>) because it introduced the idea of sentiment analysis. There are three prominent methods for conducting sentiment analysis at present: the sentiment dictionary approach, the classical machine learning approach, and the deep learning approach.</p>
<p>Experts must annotate the sentiment polarity of the text&#x00027;s terms in order for researchers to perform sentiment analysis based on sentiment dictionary. Based on semantic rules and sentiment dictionary, researchers compute the text&#x00027;s sentiment score and determine the sentiment tendency. Among these researchers, Toisoul et al. (<xref ref-type="bibr" rid="B25">2021</xref>) demonstrated positive findings on a multi-domain dataset by expanding the domain-specific vocabulary by extracting subject terms from the corpus using latent Dirichlet allocation (LDA) modeling based on the pre-existing sentiment lexicon. Peng et al. (<xref ref-type="bibr" rid="B18">2022</xref>) used the point mutual information (PMI) technique to assess the similarity of adjectives in WordNet. The polar semantics (ISA) approach was then used to generate numerous fixed sentence constructions in order to examine the target text sentiment tendency. To create a Chinese microblogging sentiment dictionary, Liu et al. (<xref ref-type="bibr" rid="B13">2021a</xref>) first identified microblogging sentences using information entropy and then filtered network sentiment terms using the sentiment-oriented pointwise mutual information (SO-PMI) method.</p>
<p>Ding et al. (<xref ref-type="bibr" rid="B10">2021</xref>) introduced the idea of the primary word and used weight priority calculations to determine the text&#x00027;s semantic inclination degree. These developments paved the way for accomplishing more difficult sentiment analysis tasks. The approach based on sentiment dictionary has the benefit of being more accurate in classifying text at the word or phrase level. However, the system migration is not good, and the sentiment dictionaries are often geared to certain domains. These days, one of the most popular techniques for sentiment analysis is classical machine learning-based techniques. Using simple bag-of-words features from a movie review dataset, Yang et al. (<xref ref-type="bibr" rid="B28">2021</xref>) was the first to use machine learning techniques to the sentiment binary classification issue and produced superior experimental outcomes. Utilizing Twitter comments as test data, Roy et al. (<xref ref-type="bibr" rid="B19">2023</xref>) classified emotions into six categories&#x02014;happiness, sadness, disgust, fear, surprise, and anger&#x02014;and employed plain Bayes for text sentiment analysis. The data were processed with consideration for lexical and expression features, leading to a high classification accuracy. To address the sentiment classification problem, Sahoo et al. (<xref ref-type="bibr" rid="B20">2021</xref>) merged a genetic algorithm with simple Bayes, and the results of the experiments indicated that the combined model outperformed the individual models. To extract rich sentiment data and include them in the basic feature model, Liu et al. (<xref ref-type="bibr" rid="B14">2021b</xref>) used machine learning techniques with numerous rules, which increased the classification result in microblog sentiment classification trials. In order to complete the study of sentiment analysis, Sampath et al. (<xref ref-type="bibr" rid="B21">2021</xref>) included semantic rules into the support vector machine model. The experiment confirmed that the support vector machine model with the inclusion of semantic rules performed better in the sentiment classification task. Deeper text semantic information is hard to learn, even while machine learning-based techniques enhance the sentiment classification performance and lower the reliance on sentiment lexicon.</p>
<p>Text sentiment analysis based on deep learning has garnered a much interest from academics at both national and international levels due to its superior performance in the fields of picture processing and natural language processing. Zhang et al. (<xref ref-type="bibr" rid="B29">2023</xref>) used deep neural network training to create the Collobert and Weston (C&#x00026;W) model, which was then used to perform well on natural language processing tasks including sentiment classification and lexical annotation. To demonstrate the efficacy of single-layer convolutional neural networks (CNNs) in sentiment classification tasks, Zhao et al. (<xref ref-type="bibr" rid="B31">2021b</xref>) combined different sizes of convolutional kernels with maximum pooling and performed comparison tests on seven datasets. The study employed convolutional neural networks for sentiment analysis tasks. A number of recurrent neural networks, including recurrent neural network (RNN), multiplicative RNN (MRNN), recursive neural tensor network (RNTN), and others, were progressively suggested by Szubielska et al. (<xref ref-type="bibr" rid="B23">2021</xref>). The RNTN model, for example, uses a syntactic analysis tree to determine word sentiment and then outputs the sentence&#x00027;s sentiment classification result in the form of word sentiment summation. To tackle the sentiment analysis problem utilizing a long short-term memory (LSTM) network with an expanded gate structure, which increases the model&#x00027;s flexibility, Li et al. (<xref ref-type="bibr" rid="B11">2022</xref>) employed Twitter comments as the experimental data. RNNs were utilized by Zhou et al. (<xref ref-type="bibr" rid="B32">2023</xref>) to model texts by taking into account their temporal information. Li et al. (<xref ref-type="bibr" rid="B12">2023</xref>) achieved outstanding results in a sentiment classification test by modeling utterances using a tree LSTM model to approximate the sentence structure. By segmenting a text according to sentences, obtaining vectors through convolutional pooling operation, and then inputting them into LSTM according to temporal relations to construct a CNN-LSTM model and apply it to the task of sentiment analysis, Alirezazadeh et al. (<xref ref-type="bibr" rid="B4">2023</xref>) primarily addressed the issue of temporal and long-range dependencies in a chapter-level text. Teodoro et al. (<xref ref-type="bibr" rid="B24">2023</xref>) constructed an experimental minimal convolutional neural network (EMCNN) model using microblog comments as the experimental data, combining lexical and emoji characteristics. The model produced experimental findings that outperformed the benchmark model&#x00027;s performance.</p></sec>
<sec id="s3">
<title>3 Attention given</title>
<p>We propose an emotion classification method based on the attention mechanism that sets blank attention in the decoder and fuses the output sequence of the encoder to learn the image semantics to guide the model to learn the image emotion more accurately and reasonably via the learning mechanism of the decoder. This method is intended to address the characteristics of small numbers of abstract painting samples and rich image semantics. <xref ref-type="fig" rid="F2">Figure 2</xref> depicts the general flowchart of the procedure used in this article, along with the encoder&#x02013;decoder architecture, the emotion classification module, and the backbone network for extracting picture feature sequences.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Model of this article.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-18-1350916-g0002.tif"/>
</fig>
<sec>
<title>3.1 Image sequence generation</title>
<p>Since the encoder anticipates a sequence as input, the abstract painting dataset in this study has been uniformly normalized, meaning that its length and width are 224 and its number of channels is 3. To extract the image&#x00027;s features, the image is supplied into the backbone network. The residual network has a strong feature learning ability and adapts to the characteristics of the backbone convolutional network architecture. In this study, ResNet-50 is adopted as the backbone network to solve the network degradation problem brought by fewer samples of abstract paintings to simplify the model training parameters of this article to a certain extent, improve the training efficiency, and carry out comparative experiments with the residual network variant in the ablation experiments, and to assess the influence of the backbone network on the model accuracy rate (Ahmad et al., <xref ref-type="bibr" rid="B2">2023</xref>). The abstract painting dataset is generated by the backbone network to generate canonical image features with a length and width of 7 and a channel count of 256 and is spread into a one-dimensional sequence, resulting in an image sequence of length 49 and a channel count of 256 to be fed to the encoder.</p>
</sec>
<sec>
<title>3.2 Encoders</title>
<p>By adjusting the number of encoder layers, the model demonstrates the significance of global image-level self-attention, guarantees that there is no appreciable loss of accuracy when <italic>N</italic> = 6, and prevents an increase in training difficulty brought on by the addition of too many parameters. This article adopts the position coding method of detection transformer (DETR), which uses the sine and cosine functions to encode the positions of rows and columns of the parity channel of the abstract painting feature map, adapting to the sequence input of the encoder&#x02013;decoder architecture (Ahmad and Wu, <xref ref-type="bibr" rid="B1">2023</xref>). The encoder&#x02013;decoder architecture is not sensitive to the order of the image sequence and does not have the ability to learn the sequence position information. The calculation for the position coding as shown in <xref ref-type="disp-formula" rid="E1">Equation (1)</xref>:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mrow><mml:mi>f</mml:mi><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mi>i</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mi>sin</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mi>x</mml:mi><mml:mrow><mml:msup><mml:mrow><mml:mn>10000</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>/</mml:mo><mml:mn>128</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mi>k</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mo stretchy='false'>[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>127</mml:mn><mml:mo stretchy='false'>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>cos</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mfrac><mml:mi>x</mml:mi><mml:mrow><mml:msup><mml:mrow><mml:mn>10000</mml:mn></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>/</mml:mo><mml:mn>28</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mo stretchy='false'>[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>127</mml:mn><mml:mo stretchy='false'>]</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <italic>x</italic> is the row and column spread value of point (<italic>p, q</italic>) and <italic>i</italic> is the channel of the feature map. For a feature map with a length and width of 7 and a channel count of 256, respectively, the row and column position encoding on the point with a channel of 10 and a coordinate value of (1,2) is sin [((1 &#x000D7; 7) &#x0002B;2)/(1,000,010/128)] and sin [((2 &#x000D7; 7) &#x0002B;1)/(1,000,010/128)], respectively, and the position encoding of the remaining image sequences of the channels is computed by this rule. Encoding finally generates a one-dimensional feature sequence with a length of 49 and a channel count of 256 with position information.</p>
<p>The <italic>Q, K, V</italic> in the encoder is a one-dimensional sequence of a fixed length of 49 and a channel count of 256, which is used as sentiment weights in translating the image sequence and ordering its position in each encoding session. As the model learns the feature dependencies between image sequences, the multi-head self-attention module supports the model by reinforcing the original features with sequence global information. This support enables the model to learn discriminative features for sentiment classification. The original image sequence serves as the input for the first coding layer, and the input for each succeeding layer is the image sequence encoded in the preceding layer. The picture feature sequences are given to the decoder after being encoded and learned by many coding layers of the encoder, avoiding the issue of delayed network convergence and poorer accuracy brought on by the increased depth of the model.</p>
</sec>
<sec>
<title>3.3 Decoders</title>
<p>The blank attention in this study has the same format as the feature sequence of the model input, that is, a sequence with a fixed length of 49 and a channel count of 256. Similar to the encoding phase, the blank attention is weighted as a query statement with <italic>Q, K, V</italic> of the first self-attention layer in the decoder, but at this point, the blank attention does not need to focus on the location information. At each decoding stage, the multi-head attention module transforms the blank attention sequences and generates the output of the attention sequences with weights by avoiding the problem of slower model convergence through the residuals and normalization module.</p>
<p>The attention sequence with weights from the upper layer and the output sequence from the encoder are fed into the second self-attention layer. This study uses the same sine and cosine functions in the decoder as in the encoder to encode the position of the weighted attention sequences from the upper layers since the output sequence of the encoder contains positional information and needs to accommodate its positional connection. The positional encoding of the picture sequence for each channel is calculated for the weighted attention sequence of length 49 and a channel count of 256. This positional encoding is applied to the rows and columns of the parity channels. Ultimately, a weighted attention sequence of length 49 and a channel count of 256 with position information are obtained. It is combined with the output sequence of the encoder as a query statement and weighted with <italic>Q, K, V</italic> from the second layer in the decoder. In each decoding stage, the output sequence of the encoder is translated, and the sequence positions are sorted.</p>
</sec>
<sec>
<title>3.4 Classification of emotions</title>
<p><xref ref-type="fig" rid="F3">Figure 3</xref> depicts the emotion classification module. The sentiment classification module combines the output sequences of the encoder and decoder to produce weighted sentiment sequences, which suppress redundant sentiment information in the model, direct the model to concentrate on deep and shallow sentiment information, and improve the model&#x00027;s ability to classify sentiment. The fully connected layer is used to map the weighted sentiment sequences, and the cross-entropy loss is minimized to produce stable sentiment classification results (Ahmad et al., <xref ref-type="bibr" rid="B3">2021</xref>). The normalized exponential function is used to calculate the probability value of each type of sentiment; the abstract painting sentiment predicted by the model has the highest probability value.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Emotional classification.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-18-1350916-g0003.tif"/>
</fig>
<p>The normalized exponential function is as shown in <xref ref-type="disp-formula" rid="E2">Equation (2)</xref>:</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:munderover></mml:mstyle><mml:msup><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>S</italic><sub><italic>i</italic></sub> is the normalized value of a particular sentiment and computation max(<italic>S</italic><sub><italic>i</italic></sub>) is the abstract painting sentiment label predicted by the model.</p>
<p>One popular loss function for handling classification difficulties is the cross-entropy function, which is primarily used to quantify the difference between two probability distributions. The cross-entropy loss function is as shown in <xref ref-type="disp-formula" rid="E3">Equation (3)</xref>:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M3"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>c</mml:mi><mml:mi>h</mml:mi><mml:mtext>_</mml:mtext><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>z</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mo class="qopname">log</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Each term in the cross-entropy function is <italic>p</italic> and <italic>q</italic> and <italic>p</italic> indicates the true probability distribution and <italic>q</italic> represents the predicted probability distribution. The cross-entropy function describes the difference between the two probability distributions. For the special case, the cross-entropy function of the binary classification problem, there are a total of two terms, i.e., probability distributions of classes 0 and 1, and there is <italic>p</italic>(0) &#x0003D; 1&#x02212;<italic>p</italic>(1), so we can get the expression for the binary classification cross entropy loss function, where <italic>y</italic><sub><italic>ic</italic></sub> is the true value and <italic>p</italic><sub><italic>ic</italic></sub> is the probability of the predicted value.</p></sec>
</sec>
<sec id="s4">
<title>4 Image preprocessing</title>
<sec>
<title>4.1 Datasets</title>
<p>The abstract dataset, which includes 280 abstract paintings, was created by Machajdik. These paintings are better suited for challenges requiring the prediction of emotion distribution because they simply feature colors and textures and not any clearly discernible objects. The 230 participants in the dataset expressed their emotions by identifying these 280 photographs, with an average of 14 people doing so. The final sentiment category is determined by which of these sentiment markers received the most votes. Due to the ambiguity of emotions, several categories may have extremely similar or identical numbers of votes, making the classification process unclear. Therefore, the ratio of votes for each emotion category is used as a probability distribution to form a probability distribution of emotions corresponding to the image, as shown in <xref ref-type="fig" rid="F3">Figure 3</xref>.</p>
</sec>
<sec>
<title>4.2 Feature extraction</title>
<p>Since abstract paintings contain only colors and textures and do not generate emotions through specific objects, the features extracted are emotional features based on the theory of artistry.</p>
<sec>
<title>4.2.1 Color histogram</title>
<p>Artists use colors to express or trigger different emotions in observers, and extracting color histograms from color features is a common and effective method. The color histogram space <italic>H</italic> is defined as <xref ref-type="disp-formula" rid="E4">Equation (4)</xref>:</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>H</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>h</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>h</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mi>h</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>h</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>h</italic>(<italic>L</italic><sub><italic>k</italic></sub>) denotes the frequency of the <italic>k</italic>th color. The similarity of the color histograms of the two images are measured using the Euclidean distance as shown in <xref ref-type="disp-formula" rid="E5">Equation (5)</xref>:</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M5"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>D</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
</sec>
<sec>
<title>4.2.2 Itten comparison</title>
<p>Itten successfully used the strategy of color combination by defining seven contrast attributes. Machajdik used seven contrast attributes such as light and dark contrast, saturation contrast, extension contrast, complementary contrast, hue contrast, warm and cool contrast, and simultaneous contrast of images as the emotional characteristics of artistry theory.</p>
<p>As in the case of light and dark contrast, the image is segmented into <italic>R</italic><sub>1</sub>, <italic>R</italic><sub>2</sub>...<italic>R</italic><sub><italic>N</italic></sub>, small chunks using the watershed segmentation algorithm, and the average <italic>h</italic><sub><italic>n</italic></sub> (Chroma) <italic>b</italic><sub><italic>n</italic></sub> (Brightness) <italic>s</italic><sub><italic>n</italic></sub> (Saturation) is calculated for each chunk. Calculation <italic>b</italic><sub><italic>n</italic></sub> belongs to five fuzzy luminance: <inline-formula><mml:math id="M6"><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mi>Very</mml:mi><mml:mi>Dark</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>V</mml:mi><mml:mi>D</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mtext>,</mml:mtext><mml:mi>Dark</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>D</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mtext>,&#x000A0;middle&#x000A0;</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:mi>M</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>Light</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>L</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mtext>,</mml:mtext><mml:mi>Very</mml:mi><mml:mi>Light</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mi>V</mml:mi><mml:mi>L</mml:mi><mml:mo stretchy='false'>)</mml:mo></mml:mtd></mml:mtr></mml:mtable><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:math></inline-formula> affiliation function as shown in <xref ref-type="disp-formula" rid="E6">Equations (6</xref>&#x02013;<xref ref-type="disp-formula" rid="E10">10)</xref>.</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M7"><mml:mrow><mml:mtext>VD</mml:mtext><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mn>1</mml:mn><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x02A7D;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>21</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mfrac><mml:mrow><mml:mn>39</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mn>18</mml:mn></mml:mrow></mml:mfrac><mml:mtext>&#x02003;</mml:mtext><mml:mn>21</mml:mn><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x0003C;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x02A7D;</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>39</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:math></disp-formula>
<disp-formula id="E7"><label>(7)</label><mml:math id="M8"><mml:mtext>D</mml:mtext><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mfrac><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mn>21</mml:mn></mml:mrow><mml:mrow><mml:mn>18</mml:mn></mml:mrow></mml:mfrac><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:mn>21</mml:mn><mml:mo>&#x0003C;</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x02264;</mml:mo><mml:mn>39</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mfrac><mml:mrow><mml:mn>55</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mn>16</mml:mn></mml:mrow></mml:mfrac><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:mn>39</mml:mn><mml:mo>&#x0003C;</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x02264;</mml:mo><mml:mn>55</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<disp-formula id="E8"><label>(8)</label><mml:math id="M9"><mml:mtext>M</mml:mtext><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mfrac><mml:mrow><mml:mn>55</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mn>16</mml:mn></mml:mrow></mml:mfrac><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:mn>39</mml:mn><mml:mo>&#x0003C;</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x02264;</mml:mo><mml:mn>55</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mfrac><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mn>55</mml:mn></mml:mrow><mml:mrow><mml:mn>13</mml:mn></mml:mrow></mml:mfrac><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:mn>55</mml:mn><mml:mo>&#x0003C;</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x02264;</mml:mo><mml:mn>68</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<disp-formula id="E9"><label>(9)</label><mml:math id="M10"><mml:mrow><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x02212;</mml:mo><mml:mn>55</mml:mn></mml:mrow><mml:mrow><mml:mn>13</mml:mn></mml:mrow></mml:mfrac><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:mn>55</mml:mn><mml:mo>&#x0003C;</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x02264;</mml:mo><mml:mn>68</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mfrac><mml:mrow><mml:mn>84</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mn>16</mml:mn></mml:mrow></mml:mfrac><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:mn>68</mml:mn><mml:mo>&#x0003C;</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x02264;</mml:mo><mml:mn>84</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<disp-formula id="E10"><label>(10)</label><mml:math id="M11"><mml:mrow><mml:mtext>VL</mml:mtext><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable><mml:mtr><mml:mtd><mml:mrow><mml:mfrac><mml:mrow><mml:mn>84</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mn>16</mml:mn></mml:mrow></mml:mfrac><mml:mtext>&#x02003;</mml:mtext><mml:mn>68</mml:mn><mml:mo>&#x0003C;</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x02264;</mml:mo><mml:mn>84</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn>1</mml:mn><mml:mtext>&#x02003;</mml:mtext><mml:msub><mml:mi>b</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x0003E;</mml:mo><mml:mn>84</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>Thus, a 1<sup>&#x0002A;</sup>5 dimensional vector for each small block of image <italic>R</italic><sub>1</sub>, <italic>R</italic><sub>2</sub>...<italic>R</italic><sub><italic>N</italic></sub> is obtained, and for the whole image, the light/dark contrast is defined as <xref ref-type="disp-formula" rid="E11">Equation (11)</xref>:</p>
<disp-formula id="E11"><label>(11)</label><mml:math id="M12"><mml:mrow><mml:mtext>B</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:mtext>i</mml:mtext><mml:mo stretchy='false'>)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:mstyle></mml:mrow></mml:mfrac><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:msub><mml:mi>R</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:mstyle><mml:msup><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>B</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02212;</mml:mo><mml:mover accent='true'><mml:mi>B</mml:mi><mml:mo stretchy='true'>&#x000AF;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></disp-formula>
<p>where <italic>i</italic> &#x0003D; 1, &#x02026;, 5, <italic>R</italic><sub><italic>n</italic></sub> is the number of pixels in the split block.</p>
<p>In this way, the vector expression of the contrasting attributes of the images is obtained as features, and the similarity of the different images is calculated by the Euclidean distance.</p>
<p>The Itten model is also used to determine whether or not an image is harmonic, and it can also be used to identify an image&#x00027;s emotional expression. Select three to four of the image&#x00027;s prominent colors, connect them to the colors on the Itten hue wheel, and if they form a positive polygon, the image is harmonic. To determine the dominant chromaticity of an image, make a histogram of its N colors. Ignore the colors with a proportion of &#x0003C; 5%. The harmony of a polygon can be assessed by comparing its internal angles to those of a square polygon built from the same number of vertices.</p></sec>
<sec>
<title>4.2.3 Texture</title>
<p>The main idea behind the statistical approach to texture analysis is to symbolize textures by the randomness of the distribution of gray levels in a graph. We define <italic>z</italic> as a random variable representing the gray levels, <italic>L</italic> as the maximum gray level of the image, <italic>Z</italic><sub><italic>i</italic></sub> as the number of pixels with gray level <italic>i</italic>, 01 denotes the gray level histogram, and with respect to <italic>z</italic>, the nth order moments are calculated as shown in <xref ref-type="disp-formula" rid="E12">Equation (12)</xref>:</p>
<disp-formula id="E12"><label>(12)</label><mml:math id="M13"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>Z</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi>m</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msup><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p><inline-formula><mml:math id="M14"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></inline-formula> is the mean value of z.</p>
<p>The second-order moments are more important in texture description; it is a measure of grayscale contrast, where <inline-formula><mml:math id="M15"><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></inline-formula> indicates the smoothness of the image, and a smaller value of <italic>u</italic><sub><italic>n</italic></sub>(<italic>z</italic>) corresponds to a smaller <italic>R</italic> value, indicating that the smaller the value of R, the smoother the image.</p>
</sec>
</sec>
<sec>
<title>4.3 Weighted <italic>K</italic>-nearest neighbor sentiment distribution prediction algorithm</title>
<p>Assuming that there are <italic>M</italic> sentiment categories <italic>C</italic><sub>1</sub>, &#x022EF;&#x02009;, <italic>C</italic><sub><italic>M</italic></sub> and <italic>N</italic> training images, <italic>x</italic><sub>1</sub>&#x022EF;&#x02009;, <italic>x</italic><sub><italic>N</italic></sub> (which also denote the corresponding features of the images) use <inline-formula><mml:math id="M16"><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">p</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>&#x022EF;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> to denote the sentiment distribution of <italic>X</italic><sub><italic>n</italic></sub>, where <italic>P</italic><sub><italic>nm</italic></sub> denotes the probability that <italic>x</italic><sub><italic>n</italic></sub> expresses a sentiment of <italic>c</italic><sub><italic>m</italic></sub>, and for each image, there is <inline-formula><mml:math id="M17"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. Assuming that <italic>y</italic> is a test image, the goal of this study is to find the sentiment distribution <inline-formula><mml:math id="M18"><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">p</mml:mtext></mml:mstyle><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>M</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> of y, i.e., as shown in <xref ref-type="disp-formula" rid="E13">Equation (13)</xref>.</p>
<disp-formula id="E13"><label>(13)</label><mml:math id="M19"><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mtext>f</mml:mtext><mml:mo stretchy='false'>(</mml:mo><mml:msubsup><mml:mrow><mml:mo>&#x0007B;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>&#x0007D;</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x02192;</mml:mo><mml:mtext>p</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<p>Training sets that are very far away have little effect on <italic>y</italic>. Considering that including all training sets can slow down the run and irrelevant training samples can also mislead the algorithm&#x00027;s classification, the effect of isolated noise samples can be eliminated by taking a weighted average of the <italic>K</italic>-nearest neighbors.</p>
<p>Weighted <italic>K</italic>-nearest neighbor option denotes only the drizzle functions corresponding to the <italic>K</italic> training images that assign the larger weights to the closer nearest neighbors. denotes the sentiment distribution of the <italic>K</italic> training images nearest to the test image, which is considered as a basis function, and the sentiment distribution <italic>P</italic> of the test image <italic>y</italic> is computed by performing a distance-weighted summation of the basis function, i.e.,</p>
<disp-formula id="E14"><label>(14)</label><mml:math id="M20"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mstyle displaystyle="false"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mstyle displaystyle="false"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>s</italic> is the similarity between the test sample and the training sample, as shown in <xref ref-type="disp-formula" rid="E15">Equation (15)</xref>.</p>
<disp-formula id="E15"><label>(15)</label><mml:math id="M21"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mi>e</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mfrac><mml:mrow><mml:mi>d</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>d</italic> is the Euclidean distance and &#x003B2; is the average distance of <italic>y</italic> from the training images.</p>
<p>Algorithm: Weighted <italic>K</italic>-nearest neighbor sentiment distribution prediction algorithm.</p>
<p>Input: Training set (<italic>x</italic><sub><italic>n</italic></sub>, <italic>p</italic><sub><italic>n</italic></sub>), test set <italic>y</italic>.</p>
<p>Output: Sentiment distribution <italic>p</italic> for the test set.</p>
<list list-type="order">
<list-item><p>Calculate the distance <italic>d</italic> between the test set image <italic>y</italic> and each image in the training set.</p></list-item>
<list-item><p>Select the first <italic>k</italic> images <italic>x</italic><sub>1</sub>...<italic>x</italic><sub><italic>k</italic></sub> that are closest to <italic>y</italic> in the increasing order of distance.</p></list-item>
<list-item><p><inline-formula><mml:math id="M22"><mml:mi>&#x003B2;</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:mfrac><mml:msqrt><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:mo>&#x0002B;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:msqrt></mml:math></inline-formula> is brought into <xref ref-type="disp-formula" rid="E14">Equation (14)</xref> in order to compute the similarity <italic>s</italic>.</p></list-item>
<list-item><p>Calculate the sentiment distribution of the test image <italic>y</italic> <inline-formula><mml:math id="M23"><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></inline-formula>.</p></list-item>
</list></sec></sec>
<sec id="s5">
<title>5 Experimentation and analysis</title>
<sec>
<title>5.1 Landscape image</title>
<p>Experiments on a large number of landscape images (resolution of &#x0007E;1 million pixels, downloaded from &#x0201C;Baidu images&#x0201D;) to achieve the simulation of ink and wash painting and to achieve a more satisfactory simulation effect. <xref ref-type="fig" rid="F4">Figure 4</xref> represents the algorithm from shallow to deep ink &#x0201C;drawing&#x0201D; simulation process: <xref ref-type="fig" rid="F4">Figure 4A</xref> shows the layer effect, <xref ref-type="fig" rid="F4">Figures 4B</xref>, <xref ref-type="fig" rid="F4">C</xref> represent the first two layers and the first three layers of the superposition effect, and <xref ref-type="fig" rid="F4">Figure 4D</xref> shows the seven-layer superposition effect, that is, the final eight-ink effect [<italic>p</italic><sub><italic>u</italic></sub>(<italic>u</italic><sub><italic>j</italic></sub>) = (0.2, 0.15, 0.15, 0.15, 0.15, 0.15, 0.15, 0.15, 0.15, 0.15, 0.15, 0.15, 0.15, 0.15, 0.15, 0.15, 0.15, 0.15, 0.15, 0.15, 0.15 and <xref ref-type="fig" rid="F4">Figure 4A</xref> for the 00 layer effect. 15, 0.15, 0.15, 0.11, 0.06, 0.03)].</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Layer stacking process effect and ink simulation results. <bold>(A)</bold> One grayscale layer overlay effect. <bold>(B)</bold> Two grayscale layer overlay effect. <bold>(C)</bold> Three grayscale layer overlay effect. <bold>(D)</bold> Seven grayscale layer overlay effect.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-18-1350916-g0004.tif"/>
</fig>
<p>In the algorithm, the histogram specification serves to preset the weight of each ink color and enhance the recognition of the inked area. <xref ref-type="fig" rid="F5">Figures 5</xref>, <xref ref-type="fig" rid="F6">6</xref> show a comparison of the ink simulation experiments for two other sets of landscape images and reveal the role of histogram specification in the simulation effect. <xref ref-type="fig" rid="F5">Figures 5B</xref>, <xref ref-type="fig" rid="F6">6B</xref> show the effect of the algorithm based on direct eighth order grayscale color reduction, and <xref ref-type="fig" rid="F5">Figures 5C</xref>, <xref ref-type="fig" rid="F6">6C</xref> show the effect of the algorithm based on histogram specification. The values of <italic>p</italic><sub><italic>u</italic></sub>(<italic>u</italic><sub><italic>j</italic></sub>) were [0.2, 0.15, 0.15, 0.15, 0.15, 0.15, 0.11, 0.06, 0.03] and [0.3, 0.125, 0.15, 0.125, 0.125, 0.1, 0.05, 0.025]. The direct eight-order color reduction approach is governed by the color values of the original diagram, which is easily the source of the imbalance of the weight of each ink color and the lack of distinctiveness, as can be seen from the comparison of the two sets of diagrams. The histogram specification method can better control the amount of ink colors and especially strengthen the weight of <italic>Gray</italic>(0) (i.e., white area). Ink simulation has a better sense of hierarchy and differentiation. In summary, the algorithm in this article simulates the ink effect of the landscape map through the method of layer simulation ink overlay, the simulation map has a strong sense of hierarchy, and the layers of ink can be integrated with each other and also has a natural paper-ink penetration effect.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Comparison of direct eighth order color reduction and histogram prescribed ink simulation 1. <bold>(A)</bold> Landscape original 2. <bold>(B)</bold> Direct eighth order color reduction link effect. <bold>(C)</bold> Histogram normalization link effect.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-18-1350916-g0005.tif"/>
</fig>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Comparison of direct eighth order color reduction and histogram prescribed ink simulation 2. <bold>(A)</bold> Landscape original 3. <bold>(B)</bold> Direct eighth order color reduction link effect. <bold>(C)</bold> Histogram normalization link effect.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-18-1350916-g0006.tif"/>
</fig>
</sec>
<sec>
<title>5.2 Abstract paintings</title>
<p>The existing sentiment classification networks ResNet and Swin Transformer and their variants are compared under the sentiment classification accuracy metrics in order to assess the effectiveness of the model in this article. The encoder&#x02013;decoder structure with various numbers of layers is set up for this article&#x00027;s method; the one-layer encoder&#x02013;decoder structure is defined as Tiny and the six-layer encoder-decoder structure is defined as Base. By training five batches of experimental findings and averaging them as the final results of the experimental data, five rounds of cross-validation were used to test the models. To accelerate the convergence of abstract painting sentiment classification, each model is fine-tuned based on the ImageNet pre-trained model, using the Adam W optimizer with a weight decay of 0.1/30 epoch and an initial learning rate of 0.0001, and trained based on the NVIDIA RTX 2080Ti.</p>
<p>The actual Naxi Dongba abstract paintings were gathered from the literature on Na xi abstract paintings, and the abstract paintings were divided into four categories based on the subject matter of the painting&#x00027;s creation. For instance, in the abstract painting data set shown in <xref ref-type="table" rid="T1">Table 1</xref>, the figures, ghosts and monsters, animals, and plants are shown from left to right, and the abstract paintings were divided into 12 different emotion categories based on the emotions they conveyed.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Example of part of the abstract painting dataset.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Emotion</bold></th>
<th valign="top" align="center" colspan="4"><bold>Painting theme</bold></th>
</tr>
</thead>
<tbody>
<tr style="background-color:#919498;color:#ffffff">
<td/>
<td valign="top" align="left"><bold>Character</bold></td>
<td valign="top" align="left"><bold>Ghosts</bold></td>
<td valign="top" align="left"><bold>Animal</bold></td>
<td valign="top" align="left"><bold>Plant</bold></td>
</tr> <tr>
<td valign="top" align="left" rowspan="2">Negative</td>
<td valign="top" align="left"><inline-graphic xlink:href="fncom-18-1350916-i0001.tif"/></td>
<td valign="top" align="left"><inline-graphic xlink:href="fncom-18-1350916-i0002.tif"/></td>
<td valign="top" align="left"><inline-graphic xlink:href="fncom-18-1350916-i0003.tif"/></td>
<td valign="top" align="left"><inline-graphic xlink:href="fncom-18-1350916-i0004.tif"/></td>
</tr>
 <tr>
<td valign="top" align="left">Be cautious</td>
<td valign="top" align="left">Aversion and resistance</td>
<td valign="top" align="left">Anxiety and tension</td>
<td valign="top" align="left">Faint and wilting</td>
</tr> <tr>
<td valign="top" align="left" rowspan="2">Neutral</td>
<td valign="top" align="left"><inline-graphic xlink:href="fncom-18-1350916-i0005.tif"/></td>
<td valign="top" align="left"><inline-graphic xlink:href="fncom-18-1350916-i0006.tif"/></td>
<td valign="top" align="left"><inline-graphic xlink:href="fncom-18-1350916-i0007.tif"/></td>
<td valign="top" align="left"><inline-graphic xlink:href="fncom-18-1350916-i0008.tif"/></td>
</tr>
 <tr>
<td valign="top" align="left">Harmony and friendliness</td>
<td valign="top" align="left">Neutral and Pure</td>
<td valign="top" align="left">Be pragmatic and responsive</td>
<td valign="top" align="left">Impartial and impartial</td>
</tr> <tr>
<td valign="top" align="left" rowspan="2">Positive</td>
<td valign="top" align="left"><inline-graphic xlink:href="fncom-18-1350916-i0009.tif"/></td>
<td valign="top" align="left"><inline-graphic xlink:href="fncom-18-1350916-i0010.tif"/></td>
<td valign="top" align="left"><inline-graphic xlink:href="fncom-18-1350916-i0011.tif"/></td>
<td valign="top" align="left"><inline-graphic xlink:href="fncom-18-1350916-i0012.tif"/></td>
</tr>
 <tr>
<td valign="top" align="left">Diligent and simple</td>
<td valign="top" align="left">Enthusiastic and proactive</td>
<td valign="top" align="left">Elegant and gentle</td>
<td valign="top" align="left">Beautiful and graceful</td>
</tr></tbody>
</table>
</table-wrap>
<p>ResNet50 was used as the backbone network in order to extract image features and tested on the test set for sentiment classification of abstract paintings.</p>
<p>The experimental findings in <xref ref-type="table" rid="T2">Table 2</xref> demonstrate that the algorithm presented in this article is superior to ResNet, Swin, and their variation network topologies for the job of sentiment recognition for abstract paintings. Established sentiment classification techniques like ResNet-101 and Swin-B achieved classification accuracies of 71.4 and 73.2%, respectively, whereas this article&#x00027;s method-Tiny and method-Base produced the best classification outcomes with classification accuracies of 74.3 and 80.8%, respectively.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Abstract painting emotion classification experiment.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Model</bold></th>
<th valign="top" align="center"><bold>Classification accuracy (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">ResNet-18</td>
<td valign="top" align="center">64.6</td>
</tr> <tr>
<td valign="top" align="left">ResNet-34</td>
<td valign="top" align="center">68.4</td>
</tr> <tr>
<td valign="top" align="left">ResNet-50</td>
<td valign="top" align="center">70.3</td>
</tr> <tr>
<td valign="top" align="left">ResNet-101</td>
<td valign="top" align="center">71.4</td>
</tr> <tr>
<td valign="top" align="left">Swin-T</td>
<td valign="top" align="center">70.1</td>
</tr> <tr>
<td valign="top" align="left">Swin-S</td>
<td valign="top" align="center">72.7</td>
</tr> <tr>
<td valign="top" align="left">Swin-B</td>
<td valign="top" align="center">73.2</td>
</tr> <tr>
<td valign="top" align="left">Vit-T</td>
<td valign="top" align="center">72.7</td>
</tr> <tr>
<td valign="top" align="left">Vit-B</td>
<td valign="top" align="center">76.8</td>
</tr> <tr>
<td valign="top" align="left">Method of this article&#x02014;Tiny</td>
<td valign="top" align="center">74.3</td>
</tr> <tr>
<td valign="top" align="left">Method of this article&#x02014;Base</td>
<td valign="top" align="center">80.8</td>
</tr></tbody>
</table>
</table-wrap>
<p>The existing sentiment classification techniques do not account for the deeper sentiment elements that are buried in abstract paintings; instead, they focus on predicting the sentiment labels of abstract paintings while neglecting their linguistically complex and emotionally varied properties. The method in this article, in contrast, uses blank attention in the decoder and fuses the encoder&#x00027;s output sequence while learning the semantics of the abstract painting image as the emotion attention through the decoder&#x00027;s decoding learning mechanism. As a result, the method employed in this study is able to achieve a higher classification accuracy rate.</p>
<p>This study first conducts ablation experiments on the backbone network, compares a variety of Res Net variants to replace the backbone network, and keeps the structure of this article&#x00027;s model unchanged for the experiments in order to assess the impact of the number of parameters of the backbone network on the accuracy of sentiment classification. The results of the experiments are shown in <xref ref-type="table" rid="T3">Table 3</xref>.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Backbone network experiment.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Backbone</bold></th>
<th valign="top" align="center"><bold>Parameter quantity (M)</bold></th>
<th valign="top" align="center"><bold>Accuracy (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">1</td>
<td valign="top" align="center">ResNet-18</td>
<td valign="top" align="center">29</td>
<td valign="top" align="center">72.5</td>
</tr> <tr>
<td valign="top" align="left">2</td>
<td valign="top" align="center">ResNet-34</td>
<td valign="top" align="center">39</td>
<td valign="top" align="center">76.7</td>
</tr> <tr>
<td valign="top" align="left">3</td>
<td valign="top" align="center">ResNet-50</td>
<td valign="top" align="center">42</td>
<td valign="top" align="center">80.8</td>
</tr> <tr>
<td valign="top" align="left">4</td>
<td valign="top" align="center">ResNet-101</td>
<td valign="top" align="center">61</td>
<td valign="top" align="center">81.9</td>
</tr></tbody>
</table>
</table-wrap>
<p>The model parameter amount was 42 M and the classification accuracy was 80.8% when ResNet-50 was used as the backbone network. The number of model parameters was cut to 29 M with the use of ResNet-18, however the model&#x00027;s classification accuracy dropped by 8.3%. ResNet-34, on the other hand, reduced the number of model parameters by 3 M while increasing the classification accuracy of the model by 4.1% when utilized as the backbone network. The number of model parameters rises by 19 M when ResNet-101 is used as the backbone network, yet the classification accuracy increases by 1.1%. In this article, choosing ResNet-50 as the backbone network ensures that there is no significant decrease in the accuracy rate and avoids the increase in training difficulty due to the introduction of too many parameters.</p>
<p><xref ref-type="fig" rid="F7">Figure 7</xref> displays the line graph of the experimental analysis of the number of coding&#x02013;decoding layers; as the number of coding&#x02013;decoding layers increases, the model&#x00027;s accuracy gradually increases, suggesting that adding more coding&#x02013;decoding layers can, to a certain extent, increase the accuracy of the classification of the emotions in abstract paintings. The model uses six coding&#x02013;decoding layers to achieve 80.8% classification accuracy, avoiding the overfitting issue that results from the stacking of coding&#x02013;decoding layers. However, as the number of coding&#x02013;decoding layers increases, the improvement in accuracy eventually slows down and becomes flat.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Analysis of the number of encoding and decoding layers.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-18-1350916-g0007.tif"/>
</fig>
<p>To prove the effectiveness of this article&#x00027;s attention mechanism for classifying the emotions of abstract paintings, two types of ablation models are set up to eliminate the decoder and encoder outputs, based on keeping the backbone network of the model as ResNet-50: The attention mechanism setup is not used in the ablation model, which eliminates the output of the decoder. Instead, the model uses the coded sequence output from the encoder as the basis for emotion classification. The classifier then normalizes the coded sequence to determine the likelihood of outputting emotion labels through the full connectivity layer spreading. The fully linked layer disperses the coded sequence, and the normalization of the classifier determines the likelihood of producing emotion labels; The attention mechanism established in this research is kept in the ablation model that removes the encoder output, and the model continues to mine the picture semantics of the abstract paintings without fusing the encoder output with the attention mechanism. In <xref ref-type="table" rid="T4">Table 4</xref>, the experimental findings are displayed.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Ablation experiment.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Decoder output</bold></th>
<th valign="top" align="center"><bold>Encoder output</bold></th>
<th valign="top" align="center"><bold>Accuracy (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">1</td>
<td valign="top" align="center">X</td>
<td valign="top" align="center">&#x0221A;</td>
<td valign="top" align="center">74.3</td>
</tr> <tr>
<td valign="top" align="left">2</td>
<td valign="top" align="center">&#x0221A;</td>
<td valign="top" align="center">X</td>
<td valign="top" align="center">77.9</td>
</tr> <tr>
<td valign="top" align="left">3</td>
<td valign="top" align="center">&#x0221A;</td>
<td valign="top" align="center">&#x0221A;</td>
<td valign="top" align="center">80.8</td>
</tr></tbody>
</table>
</table-wrap>
<p>When the encoder is used to help classify the abstraction of drawing sentiments, the accuracy of sentiment classification decreased after eliminating the decoder output by 6.5%, showing higher classification accuracy than that of the ResNet-50 classification model; however, after eliminating the encoder output, the accuracy of sentiment classification of the ablation model decreased by 2.9%, which is higher than that of the ResNet-101 classification model and close to that of the ResNet-50 classification model. This finding shows that the attention mechanism in this study can help the model recognize abstract paintings&#x00027; emotions more accurately by acting as a facilitator.</p>
<p>In this study, we used a full convolutional network to calculate the emotional weights of the model, visualize the weight heat map of the model, and simultaneously highlight and locate the regions in the heat map that significantly influence the expression of emotion.</p>
<p><xref ref-type="fig" rid="F8">Figure 8A</xref> provides an illustration of an abstract painting&#x00027;s original image, which is tagged with the predicted emotions derived from the image by the model test and contains 12 emotions as determined by the experimental data, respectively. The ablation model produced by the elimination decoder is depicted in <xref ref-type="fig" rid="F8">Figure 8B</xref>, with loose regions of attention and unfocused regions of interest in the model&#x00027;s heat map; the regions of interest for abstract paintings of various subjects also differ significantly from one another. <xref ref-type="fig" rid="F8">Figure 8C</xref> demonstrates that, despite being more compact, the model heat map&#x00027;s zone of interest suffers from ambiguous regions of interest and incorrect localization. It is also unresponsive to a smaller percentage of the neutral emotion image. The focus in the figure paintings is on the behavior and movements of the Dong ba figures, and the areas highlighted by the model labeled colors in the different image emotions correspond to the areas of the abstract paintings where the figures are holding arms, dancing, and making gestures, respectively. <xref ref-type="fig" rid="F8">Figure 8D</xref> shows the model heat map of this article, which has a more concentrated region of interest and more stable localization. For emotionally complex animal paintings, the model expands the emotional expression to the animal&#x00027;s body area; for the plant paintings, the color highlighting points out the plant petal area, which corresponds to the plant&#x00027;s budding or blossoming gesture. In the ghost paintings, the model heat map focuses on the ghost behavior and action area.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Presentation of visualization results: <bold>(A)</bold> initial image; <bold>(B)</bold> elimination decoder; <bold>(C)</bold> elimination encoder; and <bold>(D)</bold> model of this article.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-18-1350916-g0008.tif"/>
</fig>
<p>The visualization experiments demonstrate the comparison experiments of the ablation model and the region of interest of the model described in this article. They also show how the relationship between the abstract painting emotion attention and the image emotion learned by this article model is more intimate and how this has a more immediate effect on the results of the emotion classification. It demonstrates how well the model in this study extracts the emotion from images of abstract paintings, making it more appropriate for classifying the emotions of abstract paintings.</p>
</sec>
<sec>
<title>5.3 Predictive distribution</title>
<p>The effect of different values of <italic>K</italic> (<italic>K</italic> = 5, 10, 20, 40, 50, 100, 252 using 10-fold cross-validation, where <italic>K</italic> = 252 is the global weighting of the training set) on the prediction of the sentiment distribution in a weighted <italic>KKN</italic> is shown in <xref ref-type="fig" rid="F9">Figure 9</xref>.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p>Effect of different <italic>K</italic>-values on the prediction of sentiment distribution.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fncom-18-1350916-g0009.tif"/>
</fig>
<p>In this example, the optimal <italic>K</italic>-value is affected by the sentiment category, and considering the average performance, it is considered that the best prediction is achieved at <italic>K</italic> = 40, 50, which outperforms the global weighting, and when <italic>K</italic> = 252, all the training images are used for distribution prediction.</p></sec>
</sec>
<sec sec-type="conclusions" id="s6">
<title>6 Conclusion</title>
<p>The majority of early algorithms employed for sentiment classification were based on shallow machine learning and extract features using manually constructed feature selection techniques that have weak generalization ability, require extensive training times, and entail high labor costs. Because of its superior learning capacity to optimize feature extraction and prevent the flaws of manual feature selection, deep learning has produced positive research outcomes in the field of text sentiment categorization. The attention mechanism&#x00027;s primary objective is to swiftly separate valuable information from a vast amount of data. When applied to the sentiment classification task, it is capable of identifying word dependencies within sentences and identifying the internal organization of the sentence. Using a weighted closest neighbor technique, we provide a novel approach in this study to predict the discrete sentiment distribution of each picture in an abstract painting. Testing shows that the attention mechanism-based classification algorithm achieves a better classification accuracy of 80.7% when compared to state-of-the-art techniques, thereby resolving the issues of rich material and difficulties in identifying the emotions shown in abstract paintings. Nevertheless, there are several drawbacks to the attention mechanism in this article, such as its incapacity to create the positional link between objects and scenes in abstract paintings. Furthermore, it is restricted by the dataset on abstract paintings and is unable to sufficiently address the issues of imprecise sentiment categorization and imprecise attention learnt from datasets that are made publicly available. Future research methods might thus expand the sentiment dataset to a broader picture data domain and further expand the abstract painting sentiment classification system to a multimodal level in order to overcome these problems.</p></sec>
<sec sec-type="data-availability" id="s7">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p></sec>
<sec sec-type="author-contributions" id="s8">
<title>Author contributions</title>
<p>KC: Writing &#x02013; original draft.</p></sec>
</body>
<back>
<sec sec-type="funding-information" id="s9">
<title>Funding</title>
<p>The author(s) declare that no financial support was received for the research, authorship, and/or publication of this article.</p>
</sec>
<ack><p>The author would like to show sincere appreciation to the development of those techniques that have contributed to this research.</p>
</ack>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ahmad</surname> <given-names>T.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name></person-group> (<year>2023</year>). <article-title>SDIGRU: spatial and deep features integration using multilayer gated recurrent unit for human activity recognition</article-title>. <source>IEEE Trans. Comput. Soc. Syst</source>. <volume>2023</volume>:<fpage>3249152</fpage>. <pub-id pub-id-type="doi">10.1109/TCSS.2023.3249152</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ahmad</surname> <given-names>T.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name> <name><surname>Alwageed</surname> <given-names>H. S.</given-names></name> <name><surname>Khan</surname> <given-names>F.</given-names></name> <name><surname>Khan</surname> <given-names>J.</given-names></name> <name><surname>Lee</surname> <given-names>Y.</given-names></name></person-group> (<year>2023</year>). <article-title>Human activity recognition based on deep-temporal learning using convolution neural networks features and bidirectional gated recurrent unit with features selection</article-title>. <source>IEEE Access</source> <volume>11</volume>, <fpage>33148</fpage>&#x02013;<lpage>33159</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2023.3263155</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ahmad</surname> <given-names>T.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name> <name><surname>Khan</surname> <given-names>I.</given-names></name> <name><surname>Rahim</surname> <given-names>A.</given-names></name> <name><surname>Khan</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>Human action recognition in video sequence using logistic regression by features fusion approach based on CNN features</article-title>. <source>Int. J. Adv. Comput. Sci. Appl.</source> <volume>11</volume>:<fpage>121103</fpage>. <pub-id pub-id-type="doi">10.14569/IJACSA.2021.0121103</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Alirezazadeh</surname> <given-names>P.</given-names></name> <name><surname>Schirrmann</surname> <given-names>M.</given-names></name> <name><surname>Stolzenburg</surname> <given-names>F.</given-names></name></person-group> (<year>2023</year>). <article-title>Improving deep learning-based plant disease classification with attention mechanism</article-title>. <source>Gesunde Pflanzen</source> <volume>75</volume>, <fpage>49</fpage>&#x02013;<lpage>59</lpage>. <pub-id pub-id-type="doi">10.1007/s10343-022-00796-y</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bharadiya</surname> <given-names>J.</given-names></name></person-group> (<year>2023</year>). <article-title>Convolutional neural networks for image classification</article-title>. <source>Int. J. Innov. Sci. Res. Technol</source>. <volume>8</volume>, <fpage>673</fpage>&#x02013;<lpage>677</lpage>. <pub-id pub-id-type="doi">10.5281/zenodo.8020781</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cetinic</surname> <given-names>E.</given-names></name> <name><surname>She</surname> <given-names>J.</given-names></name></person-group> (<year>2022</year>). <article-title>Understanding and creating art with AI: review and outlook</article-title>. <source>ACM Trans. Multimed. Comput. Commun. Appl</source>. <volume>18</volume>, <fpage>1</fpage>&#x02013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.1145/3475799</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chan</surname> <given-names>J. Y. L.</given-names></name> <name><surname>Bea</surname> <given-names>K. T.</given-names></name> <name><surname>Leow</surname> <given-names>S. M. H.</given-names></name> <name><surname>Phoong</surname> <given-names>S. W.</given-names></name> <name><surname>Cheng</surname> <given-names>W. K.</given-names></name></person-group> (<year>2023</year>). <article-title>State of the art: a review of sentiment analysis based on sequential transfer learning</article-title>. <source>Artif. Intell. Rev</source>. <volume>56</volume>, <fpage>749</fpage>&#x02013;<lpage>780</lpage>. <pub-id pub-id-type="doi">10.1007/s10462-022-10183-8</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chandrasekaran</surname> <given-names>G.</given-names></name> <name><surname>Antoanela</surname> <given-names>N.</given-names></name> <name><surname>Andrei</surname> <given-names>G.</given-names></name> <name><surname>Monica</surname> <given-names>C.</given-names></name> <name><surname>Hemanth</surname> <given-names>J.</given-names></name></person-group> (<year>2022</year>). <article-title>Visual sentiment analysis using deep learning models with social media data</article-title>. <source>Appl. Sci</source>. <volume>12</volume>:<fpage>1030</fpage>. <pub-id pub-id-type="doi">10.3390/app12031030</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Hua</surname> <given-names>Z.</given-names></name></person-group> (<year>2023</year>). <article-title>Retinex low-light image enhancement network based on attention mechanism</article-title>. <source>Multimed. Tools Appl</source>. <volume>82</volume>, <fpage>4235</fpage>&#x02013;<lpage>4255</lpage>. <pub-id pub-id-type="doi">10.1007/s11042-022-13411-z</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ding</surname> <given-names>F.</given-names></name> <name><surname>Yu</surname> <given-names>K.</given-names></name> <name><surname>Gu</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Shi</surname> <given-names>Y.</given-names></name></person-group> (<year>2021</year>). <article-title>Perceptual enhancement for autonomous vehicles: restoring visually degraded images for context prediction via adversarial training</article-title>. <source>IEEE Trans. Intell. Transport. Syst</source>. <volume>23</volume>, <fpage>9430</fpage>&#x02013;<lpage>9441</lpage>. <pub-id pub-id-type="doi">10.1109/TITS.2021.3120075</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Shao</surname> <given-names>W.</given-names></name> <name><surname>Ji</surname> <given-names>S.</given-names></name> <name><surname>Cambria</surname> <given-names>E.</given-names></name></person-group> (<year>2022</year>). <article-title>BiERU: bidirectional emotional recurrent unit for conversational sentiment analysis</article-title>. <source>Neurocomputing</source> <volume>467</volume>, <fpage>73</fpage>&#x02013;<lpage>82</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2021.09.057</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Li</surname> <given-names>M.</given-names></name> <name><surname>Yan</surname> <given-names>P.</given-names></name> <name><surname>Li</surname> <given-names>G.</given-names></name> <name><surname>Jiang</surname> <given-names>Y.</given-names></name> <name><surname>Luo</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Deep learning attention mechanism in medical image analysis: basics and beyonds</article-title>. <source>Int. J. Netw. Dyn. Intell.</source> <volume>2023</volume>, <fpage>93</fpage>&#x02013;<lpage>116</lpage>. <pub-id pub-id-type="doi">10.53941/ijndi0201006</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Nie</surname> <given-names>H.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>Y. F.</given-names></name></person-group> (<year>2021a</year>). <article-title>Anisotropic angle distribution learning for head pose estimation and attention understanding in human-computer interaction</article-title>. <source>Neurocomputing</source> <volume>433</volume>, <fpage>310</fpage>&#x02013;<lpage>322</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2020.09.068</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>T.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Yang</surname> <given-names>B.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name></person-group> (<year>2021b</year>). <article-title>NGDNet: nonuniform Gaussian-label distribution learning for infrared head pose estimation and on-task behavior understanding in the classroom</article-title>. <source>Neurocomputing</source> <volume>436</volume>, <fpage>210</fpage>&#x02013;<lpage>220</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2020.12.090</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McCormack</surname> <given-names>J.</given-names></name> <name><surname>Lomas</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>Deep learning of individual aesthetics</article-title>. <source>Neural Comput. Appl</source>. <volume>33</volume>, <fpage>3</fpage>&#x02013;<lpage>17</lpage>. <pub-id pub-id-type="doi">10.1007/s00521-020-05376-7</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Milani</surname> <given-names>F.</given-names></name> <name><surname>Fraternali</surname> <given-names>P.</given-names></name></person-group> (<year>2021</year>). <article-title>A dataset and a convolutional model for iconography classification in paintings</article-title>. <source>J. Comput. Cult. Herit</source>. <volume>14</volume>, <fpage>1</fpage>&#x02013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1145/3458885</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ngai</surname> <given-names>W. K.</given-names></name> <name><surname>Xie</surname> <given-names>H.</given-names></name> <name><surname>Zou</surname> <given-names>D.</given-names></name> <name><surname>Chou</surname> <given-names>K. L.</given-names></name></person-group> (<year>2022</year>). <article-title>Emotion recognition based on convolutional neural networks and heterogeneous bio-signal data sources</article-title>. <source>Inform. Fusion</source> <volume>77</volume>, <fpage>107</fpage>&#x02013;<lpage>117</lpage>. <pub-id pub-id-type="doi">10.1016/j.inffus.2021.07.007</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Peng</surname> <given-names>S.</given-names></name> <name><surname>Cao</surname> <given-names>L.</given-names></name> <name><surname>Zhou</surname> <given-names>Y.</given-names></name> <name><surname>Ouyang</surname> <given-names>Z.</given-names></name> <name><surname>Yang</surname> <given-names>A.</given-names></name> <name><surname>Li</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>A survey on deep learning for textual emotion analysis in social networks</article-title>. <source>Digit. Commun. Netw</source>. <volume>8</volume>, <fpage>745</fpage>&#x02013;<lpage>762</lpage>. <pub-id pub-id-type="doi">10.1016/j.dcan.2021.10.003</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roy</surname> <given-names>S. K.</given-names></name> <name><surname>Deria</surname> <given-names>A.</given-names></name> <name><surname>Shah</surname> <given-names>C.</given-names></name> <name><surname>Haut</surname> <given-names>J. M.</given-names></name> <name><surname>Du</surname> <given-names>Q.</given-names></name> <name><surname>Plaza</surname> <given-names>A.</given-names></name></person-group> (<year>2023</year>). <article-title>Spectral-spatial morphological attention transformer for hyperspectral image classification</article-title>. <source>IEEE Trans. Geosci. Rem. Sens</source>. <volume>61</volume>, <fpage>1</fpage>&#x02013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2023.3242346</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sahoo</surname> <given-names>K. K.</given-names></name> <name><surname>Dutta</surname> <given-names>I.</given-names></name> <name><surname>Ijaz</surname> <given-names>M. F.</given-names></name> <name><surname>Wozniak</surname> <given-names>M.</given-names></name> <name><surname>Singh</surname> <given-names>P. K.</given-names></name></person-group> (<year>2021</year>). <article-title>TLEFuzzyNet: fuzzy rank-based ensemble of transfer learning models for emotion recognition from human speeches</article-title>. <source>IEEE Access</source> <volume>9</volume>, <fpage>166518</fpage>&#x02013;<lpage>166530</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2021.3135658</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sampath</surname> <given-names>V.</given-names></name> <name><surname>Maurtua</surname> <given-names>I.</given-names></name> <name><surname>Aguilar Martin</surname> <given-names>J. J.</given-names></name> <name><surname>Gutierrez</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>A survey on generative adversarial networks for imbalance problems in computer vision tasks</article-title>. <source>J. Big Data</source> <volume>8</volume>, <fpage>1</fpage>&#x02013;<lpage>59</lpage>. <pub-id pub-id-type="doi">10.1186/s40537-021-00414-0</pub-id><pub-id pub-id-type="pmid">33552840</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Song</surname> <given-names>T.</given-names></name> <name><surname>Zheng</surname> <given-names>W.</given-names></name> <name><surname>Liu</surname> <given-names>S.</given-names></name> <name><surname>Zong</surname> <given-names>Y.</given-names></name> <name><surname>Cui</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name></person-group> (<year>2021</year>). <article-title>Graph-embedded convolutional neural network for image-based EEG emotion recognition</article-title>. <source>IEEE Trans. Emerg. Top. Comput</source>. <volume>10</volume>, <fpage>1399</fpage>&#x02013;<lpage>1413</lpage>. <pub-id pub-id-type="doi">10.1109/TETC.2021.3087174</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Szubielska</surname> <given-names>M.</given-names></name> <name><surname>Imbir</surname> <given-names>K.</given-names></name> <name><surname>Szyma&#x00144;ska</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>The influence of the physical context and knowledge of artworks on the aesthetic experience of interactive installations</article-title>. <source>Curr. Psychol</source>. <volume>40</volume>, <fpage>3702</fpage>&#x02013;<lpage>3715</lpage>. <pub-id pub-id-type="doi">10.1007/s12144-019-00322-w</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Teodoro</surname> <given-names>A. A.</given-names></name> <name><surname>Silva</surname> <given-names>D. H.</given-names></name> <name><surname>Rosa</surname> <given-names>R. L.</given-names></name> <name><surname>Saadi</surname> <given-names>M.</given-names></name> <name><surname>Wuttisittikulkij</surname> <given-names>L.</given-names></name> <name><surname>Mumtaz</surname> <given-names>R. A.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>A skin cancer classification approach using gan and roi-based attention mechanism</article-title>. <source>J. Sign. Process. Syst.</source> <volume>95</volume>, <fpage>211</fpage>&#x02013;<lpage>224</lpage>. <pub-id pub-id-type="doi">10.1007/s11265-022-01757-4</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Toisoul</surname> <given-names>A.</given-names></name> <name><surname>Kossaifi</surname> <given-names>J.</given-names></name> <name><surname>Bulat</surname> <given-names>A.</given-names></name> <name><surname>Tzimiropoulos</surname> <given-names>G.</given-names></name> <name><surname>Pantic</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>Estimation of continuous valence and arousal levels from faces in naturalistic conditions</article-title>. <source>Nat. Machine Intell</source>. <volume>3</volume>, <fpage>42</fpage>&#x02013;<lpage>50</lpage>. <pub-id pub-id-type="doi">10.1038/s42256-020-00280-0</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Song</surname> <given-names>W.</given-names></name> <name><surname>Tao</surname> <given-names>W.</given-names></name> <name><surname>Liotta</surname> <given-names>A.</given-names></name> <name><surname>Yang</surname> <given-names>D.</given-names></name> <name><surname>Li</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>A systematic review on affective computing: emotion models, databases, and recent advances</article-title>. <source>Inform. Fusion</source> <volume>83</volume>, <fpage>19</fpage>&#x02013;<lpage>52</lpage>. <pub-id pub-id-type="doi">10.1016/j.inffus.2022.03.009</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>H.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name></person-group> (<year>2023</year>). <article-title>CovidViT: a novel neural network with self-attention mechanism to detect COVID-19 through X-ray images</article-title>. <source>Int. J. Machine Learn. Cybernet</source>. <volume>14</volume>, <fpage>973</fpage>&#x02013;<lpage>987</lpage>. <pub-id pub-id-type="doi">10.1007/s13042-022-01676-7</pub-id><pub-id pub-id-type="pmid">36274812</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Z.</given-names></name> <name><surname>Baraldi</surname> <given-names>P.</given-names></name> <name><surname>Zio</surname> <given-names>E.</given-names></name></person-group> (<year>2021</year>). <article-title>A multi-branch deep neural network model for failure prognostics based on multimodal data</article-title>. <source>J. Manufact. Syst</source>. <volume>59</volume>, <fpage>42</fpage>&#x02013;<lpage>50</lpage>. <pub-id pub-id-type="doi">10.1016/j.jmsy.2021.01.007</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>M.</given-names></name> <name><surname>Wu</surname> <given-names>D.</given-names></name></person-group> (<year>2023</year>). <article-title>Federated multidomain learning with graph ensemble autoencoder GMM for emotion recognition</article-title>. <source>IEEE Trans. Intell. Transp. Syst.</source> <volume>24</volume>, <fpage>7631</fpage>&#x02013;<lpage>7641</lpage>. <pub-id pub-id-type="doi">10.1109/TITS.2022.3203800</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>S.</given-names></name> <name><surname>Jia</surname> <given-names>G.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Ding</surname> <given-names>G.</given-names></name> <name><surname>Keutzer</surname> <given-names>K.</given-names></name></person-group> (<year>2021a</year>). <article-title>Emotion recognition from multiple modalities: fundamentals and methodologies</article-title>. <source>IEEE Sign. Process. Mag</source>. <volume>38</volume>, <fpage>59</fpage>&#x02013;<lpage>73</lpage>. <pub-id pub-id-type="doi">10.1109/MSP.2021.3106895</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>W.</given-names></name> <name><surname>Zhou</surname> <given-names>D.</given-names></name> <name><surname>Qiu</surname> <given-names>X.</given-names></name> <name><surname>Jiang</surname> <given-names>W.</given-names></name></person-group> (<year>2021b</year>). <article-title>Compare the performance of the models in art classification</article-title>. <source>PLoS ONE</source> <volume>16</volume>:<fpage>e0248414</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0248414</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>J.</given-names></name> <name><surname>Pang</surname> <given-names>L.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name></person-group> (<year>2023</year>). <article-title>Underwater image enhancement method by multi-interval histogram equalization</article-title>. <source>IEEE J. Ocean. Eng</source>. <volume>48</volume>, <fpage>474</fpage>&#x02013;<lpage>488</lpage>. <pub-id pub-id-type="doi">10.1109/JOE.2022.3223733</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zou</surname> <given-names>Q.</given-names></name> <name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Yang</surname> <given-names>S.</given-names></name> <name><surname>Chen</surname> <given-names>B.</given-names></name></person-group> (<year>2023</year>). <article-title>A compact periocular recognition system based on deep learning framework AttenMidNet with the attention mechanism</article-title>. <source>Multimed. Tools Appl</source>. <volume>82</volume>, <fpage>15837</fpage>&#x02013;<lpage>15857</lpage>. <pub-id pub-id-type="doi">10.1007/s11042-022-14017-1</pub-id></citation>
</ref>
</ref-list>
</back>
</article>