<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurosci.</journal-id>
<journal-title>Frontiers in Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-453X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnins.2021.739138</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>IE-IQA: Intelligibility Enriched Generalizable No-Reference Image Quality Assessment</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Song</surname> <given-names>Tianshu</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1402306/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Li</surname> <given-names>Leida</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1473172/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhu</surname> <given-names>Hancheng</given-names></name>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Qian</surname> <given-names>Jiansheng</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c002"><sup>&#x0002A;</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>School of Information and Control Engineering, China University of Mining and Technology</institution>, <addr-line>Xuzhou</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>School of Artificial Intelligence, Xidian University</institution>, <addr-line>Xi&#x00027;an</addr-line>, <country>China</country></aff>
<aff id="aff3"><sup>3</sup><institution>Pazhou Lab</institution>, <addr-line>Guangzhou</addr-line>, <country>China</country></aff>
<aff id="aff4"><sup>4</sup><institution>School of Computer Science and Technology, China University of Mining and Technology</institution>, <addr-line>Xuzhou</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Guangtao Zhai, Shanghai Jiao Tong University, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Yucheng Zhu, Shanghai Jiao Tong University, China; Weisheng Li, Chongqing University of Posts and Telecommunications, China; Wei Sun, Shanghai Jiao Tong University, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Leida Li <email>ldli&#x00040;xidian.edu.cn</email></corresp>
<corresp id="c002">Jiansheng Qian <email>qianjsh&#x00040;cumt.edu.cn</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Perception Science, a section of the journal Frontiers in Neuroscience</p></fn></author-notes>
<pub-date pub-type="epub">
<day>21</day>
<month>10</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>15</volume>
<elocation-id>739138</elocation-id>
<history>
<date date-type="received">
<day>10</day>
<month>06</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>08</day>
<month>09</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2021 Song, Li, Zhu and Qian.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Song, Li, Zhu and Qian</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> 
</permissions>
<abstract><p>Image quality assessment (IQA) for authentic distortions in the wild is challenging. Though current IQA metrics have achieved decent performance for synthetic distortions, they still cannot be satisfactorily applied to realistic distortions because of the generalization problem. Improving generalization ability is an urgent task to make IQA algorithms serviceable in real-world applications, while relevant research is still rare. Fundamentally, image quality is determined by both distortion degree and intelligibility. However, current IQA metrics mostly focus on the distortion aspect and do not fully investigate the intelligibility, which is crucial for achieving robust quality estimation. Motivated by this, this paper presents a new framework for building highly generalizable image quality model by integrating the intelligibility. We first analyze the relation between intelligibility and image quality. Then we propose a bilateral network to integrate the above two aspects of image quality. During the fusion process, feature selection strategy is further devised to avoid negative transfer. The framework not only catches the conventional distortion features but also integrates intelligibility features properly, based on which a highly generalizable no-reference image quality model is achieved. Extensive experiments are conducted based on five intelligibility tasks, and the results demonstrate that the proposed approach outperforms the state-of-the-art metrics, and the intelligibility task consistently improves metric performance and generalization ability.</p></abstract>
<kwd-group>
<kwd>image quality assessment</kwd>
<kwd>NR-IQA</kwd>
<kwd>intelligibility</kwd>
<kwd>distortion</kwd>
<kwd>generalization</kwd>
<kwd>semantic</kwd>
</kwd-group>
<counts>
<fig-count count="7"/>
<table-count count="5"/>
<equation-count count="7"/>
<ref-count count="50"/>
<page-count count="12"/>
<word-count count="7776"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Image quality assessment (IQA) plays a vital role in image acquisition, compression, enhancement, retrieval, etc. The existing IQA metrics are mainly designed for synthetic distortions and cannot be applied to wild images satisfactorily due to the limited generalization ability. Fundamentally, image quality embodies two aspects: distortion and intelligibility (Abdou and Dusaussoy, <xref ref-type="bibr" rid="B1">1986</xref>). Most IQA algorithms only focus on the distortion measurement and the intelligibility aspect is rarely investigated. In this paper, we mainly investigate the role of intelligibility in building a highly generalizable IQA model.</p>
<p>Intelligibility refers to the ability of an image to provide information to a person or a machine (Abdou and Dusaussoy, <xref ref-type="bibr" rid="B1">1986</xref>), that is, the degree to which the image could be understood. Distortions affect image intelligibility, and accordingly, intelligibility is indicative of image quality when humans make judgments. Traditional handcrafted feature-based IQA metrics mainly focus on distortions and cannot commendably describe image intelligibility. Deep learning-based methods learn the IQA task in a data-driven manner, and consequently do not directly pay attention to image intelligibility, either.</p>
<p>Since the most essential function of image is to convey information, when distortions seriously undermine the expression of information, the intelligibility will also become low, which in turn indicates poor image quality. Real-world images are typically contaminated by complicated distortions, which lead to different degrees of intelligibility. <xref ref-type="fig" rid="F1">Figure 1</xref> explains how intelligibility indicates image quality. <xref ref-type="fig" rid="F1">Figures 1A,B</xref> both suffer from severe motion blur, and both contain human as the main content. The human face in <xref ref-type="fig" rid="F1">Figure 1A</xref> is too blurred to be recognized, whereas a woman&#x00027;s face in <xref ref-type="fig" rid="F1">Figure 1B</xref> can still be easily identified. Thus, <xref ref-type="fig" rid="F1">Figure 1B</xref> has higher intelligibility and accordingly higher quality score. The distortion in <xref ref-type="fig" rid="F1">Figure 1C</xref> is not heavier than <xref ref-type="fig" rid="F1">Figure 1D</xref>, but <xref ref-type="fig" rid="F1">Figure 1D</xref> is easier to be recognized; hence, <xref ref-type="fig" rid="F1">Figure 1D</xref> has higher intelligibility and accordingly higher quality score. Finally, <xref ref-type="fig" rid="F1">Figures 1E,F</xref> was mainly underexposed with locally overexposed. The main content in <xref ref-type="fig" rid="F1">Figure 1E</xref> is illegible, whereas <xref ref-type="fig" rid="F1">Figure 1F</xref> can still be distinguished as a singing stage with performers. Therefore, the quality of <xref ref-type="fig" rid="F1">Figure 1F</xref> is better than that of <xref ref-type="fig" rid="F1">Figure 1E</xref>. It can be concluded from <xref ref-type="fig" rid="F1">Figure 1</xref> that images with similar distortions may have significantly different quality due to different degrees of intelligibility. Therefore, a robust quality assessment metric should also take intelligibility into account, especially for severe distortions.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Relation between intelligibility and image quality. <bold>(A&#x02013;F)</bold> Compared to images in the first row, images in the second row have higher intelligibility and accordingly higher mean opinion score (MOS). Images are from the KonIQ-10k (Hosu et al., <xref ref-type="bibr" rid="B16">2020</xref>) dataset. The range of MOS is [1, 5], and higher MOS means better quality.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-739138-g0001.tif"/>
</fig>
<p>Motivated by the above facts, this paper presents a new framework to achieve highly generalizable image quality assessment by integrating intelligibility and distortion measure. The intelligibility of an image can be represented from different perspectives, such as &#x0201C;whether the content of the image is recognizable,&#x0201D; &#x0201C;which category does the main object in the image belong to,&#x0201D; and &#x0201C;what scene does the image show.&#x0201D; The results of these questions are all important information conveyed by the image, and through the mining of these questions, we can obtain descriptions of image intelligibility. These questions can be described by popular computer vision tasks, such as image classification, scene recognition, object detection, and instance segmentation. Therefore, we calculate intelligibility features based on these semantic tasks. Then, we propose a bilateral network to combine the distortion features and intelligibility features. Further, we design different feature selection strategies for different semantic understanding tasks. This produces highly generalizable intelligibility features. The distortion network is applied to extract distortion features that are complementary to those intelligibility features. With the bilateral network, highly generalizable intelligibility features with rich semantic information can be fused with distortion features, producing the final IQA model.</p>
<p>The contributions of this work are summarized as follows:</p>
<list list-type="bullet">
<list-item><p>We propose a new framework for designing highly generalizable image quality models by integrating intelligibility and distortion, two fundamental aspects of image quality. In the proposed framework, intelligibility features can be extracted based on popular semantic tasks, such as image recognition, scene classification, and object detection.</p></list-item>
<list-item><p>We propose a bilateral network with an intelligibility enhanced module to fuse intelligibility features with distortion features for building a robust IQA model. A feature selection strategy is proposed to extract intelligibility features instead of doing direct training. This strategy can avoid the risk of damaging generalizable features.</p></list-item>
<list-item><p>We have verified the effectiveness of the proposed method through extensive experiments and compared with the state-of-the-arts. The experimental results demonstrate that the proposed model can achieve significantly better generalization performance.</p></list-item>
</list>
</sec>
<sec id="s2">
<title>2. Related Work</title>
<p>Early no-reference IQA (NR-IQA) metrics typically train a regressor to obtain quality scores based on handcrafted features. For example, BLIINDS-II (Saad et al., <xref ref-type="bibr" rid="B31">2012</xref>), BRISQUE (Mittal et al., <xref ref-type="bibr" rid="B26">2012</xref>), and BIQI (Moorthy and Bovik, <xref ref-type="bibr" rid="B27">2010</xref>) designed features meticulously through natural scene statistics (NSS). NFERM (Gu et al., <xref ref-type="bibr" rid="B14">2014</xref>) incorporated features inspired by the free energy theory, human visual system, and NSS. CORNIA (Ye et al., <xref ref-type="bibr" rid="B42">2012</xref>) and HOSA (Xu et al., <xref ref-type="bibr" rid="B40">2016</xref>) trained large-scale visual codebooks from natural image to make predictions. The above handcrafted feature-based IQA models are usually limited in handling the diversified scenes and distortion types in real-world images.</p>
<p>With the boom of deep learning, convolutional neural networks have been widely applied in IQA. Early attempts utilized relatively shallow networks (Kang et al., <xref ref-type="bibr" rid="B18">2014</xref>; Kang et al., <xref ref-type="bibr" rid="B19">2015</xref>; Kottayil et al., <xref ref-type="bibr" rid="B21">2016</xref>) to extract features for assessing synthetic distortions. Then, deeper networks were utilized to handle more complex distortions (Bosse et al., <xref ref-type="bibr" rid="B3">2017</xref>; Kim and Lee, <xref ref-type="bibr" rid="B20">2017</xref>; Ma et al., <xref ref-type="bibr" rid="B25">2017</xref>; Yan et al., <xref ref-type="bibr" rid="B41">2019</xref>; Zhai et al., <xref ref-type="bibr" rid="B43">2020</xref>; Zhang J. et al., <xref ref-type="bibr" rid="B45">2020</xref>). It is widely acknowledged that large datasets are needed for training deep neural networks. However, so far the largest IQA dataset only has 11,125 images, which are still limited. Thus, recent deep IQA metrics (Bianco et al., <xref ref-type="bibr" rid="B2">2018</xref>; Varga et al., <xref ref-type="bibr" rid="B37">2018</xref>; Zhang W. et al., <xref ref-type="bibr" rid="B47">2020</xref>) utilize networks pre-trained on large-scale computer vision tasks and then fine-tune on them. For example, Bianco et al. (<xref ref-type="bibr" rid="B2">2018</xref>) made fine-tuning on the model pre-trained on subset of ImageNet (Imagenet large scale visual recognition challenge, 1.3M images) (Russakovsky et al., <xref ref-type="bibr" rid="B30">2015</xref>) and Places-205 (Wang et al., <xref ref-type="bibr" rid="B39">2015</xref>) (2.5M images). Varga et al. (<xref ref-type="bibr" rid="B37">2018</xref>) made fine-tuning on deep pre-trained network (ResNet101 He et al., <xref ref-type="bibr" rid="B15">2016</xref>) to learn the distribution of mean opinion score (MOS). Zhang W. et al. (<xref ref-type="bibr" rid="B47">2020</xref>) utilized two different networks to evaluate synthetic and authentic distortions, respectively, and the authentic network was fine-tuned on pre-trained network (VGG16, Simonyan and Zisserman, <xref ref-type="bibr" rid="B33">2015</xref>). Make fine-tuning on pre-trained model of recognition task is a suboptimal method because IQA task is different from recognition tasks. Recognition tasks should be robust to distortions while IQA should distinguish distortions. Though fine-tuning with IQA images can improve IQA performance, the generalizable features trained with large-scale dataset were damaged during further training. And due to the small sample property of IQA, generalization ability of new features is still unsatisfying and cannot be adopted to real-world applications.</p>
<p>Until recently, the generalization problem of IQA models began to receive attention. Zhu et al. (<xref ref-type="bibr" rid="B50">2020</xref>) adopted meta-learning to learn the prior knowledge of distortions in synthetic distortions and then fine-tune on authentic distortions to achieve better generalization ability. Hosu et al. (<xref ref-type="bibr" rid="B16">2020</xref>) built a large dataset (KonIQ-10k: 10,073 images) for model training and obtained better generalization performance. Su et al. (<xref ref-type="bibr" rid="B34">2020</xref>) incorporated semantic features and multi-scale content features to handle challenges of distortion diversity and content variation. The above methods have achieved better generalization performance than earlier metrics, but their generalization ability is still far from ideal and further explorations are needed. In this paper, we work toward this direction by proposing a new framework to address the generalization problem, where the intelligibility property of images is investigated.</p>
</sec>
<sec id="s3">
<title>3. Proposed Method</title>
<sec>
<title>3.1. Relation Between Intelligibility and Quality</title>
<p>As aforementioned, image intelligibility can be described by semantic understanding tasks. The most popular one is the classification task on Imagenet Large Scale Visual Recognition Challenge, which has 1.3 M images belonging to 1,000 classes (Russakovsky et al., <xref ref-type="bibr" rid="B30">2015</xref>). Therefore, we take the deep convolutional neural network (DCNN) trained on this task as an example. The output of the classification network is a probability distribution <italic>o</italic><sub><italic>i</italic></sub>, <italic>i</italic> &#x0003D; 1, 2, ..., 1, 000, and 1, 000 is the total number of classes. The prediction confidence <italic>c</italic> can be obtained by</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>c</mml:mi><mml:mo>=</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>o</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:mn>1000</mml:mn><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The confidence <italic>c</italic> in Equation (1) also represents the top1-probability. If the intelligibility of an image is high, the model may easily recognize the category and the top1-probability may be notably high. When the intelligibility is low, the model will be unconfident of its predictions and the top1-probability also tends to be low.</p>
<p>To have an intuitive understanding of the above characteristic, we compare the average classification confidence score obtained from images of different quality. First, we divide images from an IQA dataset into several groups according to their MOS values in ascending order. (Specifically, MOS are divided into 6 equal intervals of [<italic>m</italic><sub><italic>i</italic></sub>, <italic>m</italic><sub><italic>i</italic>&#x0002B;1</sub>] where <italic>i</italic> &#x0003D; 1 &#x02212; 6, <italic>m</italic><sub>1</sub> &#x0003D; <italic>min</italic>(<italic>MOS</italic>), <italic>m</italic><sub>7</sub> &#x0003D; <italic>max</italic>(<italic>MOS</italic>).) Then, we utilize an image classification model trained on ImageNet to obtain the confidence score of images in each group. Finally, we calculate the average confidence score of each quality group, and illustrate them in <xref ref-type="fig" rid="F2">Figure 2A</xref>. We can observe that images with poor quality tend to have lower prediction confidence than those with high quality. That is, image quality does have a significant impact on intelligibility.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Relation between recognition confidence and image quality. <bold>(A)</bold> Image quality affects recognition confidence; <bold>(B)</bold> recognition confidence indicates image quality; <bold>(C&#x02013;H)</bold> representative images with different prediction confidence and mean opinion score (MOS). Panels <bold>(C&#x02013;H)</bold> correspond to six ascending bins of B whose MOS increases with confidence. All results are obtained from the KonIQ-10k dataset based on EfficientNet-B0 network (Tan and Le, <xref ref-type="bibr" rid="B35">2019</xref>).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-739138-g0002.tif"/>
</fig>
<p>In this paper, we are more interested in how intelligibility indicates image quality. Therefore, we do another experiment by dividing images according to the prediction confidence and compare the average MOS value of different confidence intervals. The results are presented in <xref ref-type="fig" rid="F2">Figure 2B</xref>. Furthermore, to show the relation more intuitively, we also show sample images in <xref ref-type="fig" rid="F2">Figures 2C&#x02013;H</xref> that corresponds to the six ascending bins of <xref ref-type="fig" rid="F2">Figure 2B</xref>. We can observe from <xref ref-type="fig" rid="F2">Figures 2B&#x02013;H</xref> that intelligibility described by image recognition task can distinctly indicate image quality.</p>
</sec>
<sec>
<title>3.2. Our Framework</title>
<p>In this paper, we propose an intelligibility enriched IQA (IE-IQA) framework, as illustrated in <xref ref-type="fig" rid="F3">Figure 3</xref>. In our framework, we propose a bilateral network to integrate intelligibility features and conventional distortion features. Since intelligibility can be represented using different image understanding tasks, it is reasonable to utilize features from these tasks as intelligibility features. However, IQA is different from image understanding tasks, and directly utilizing features of image understanding tasks may lead to negative transfer, which has been proved by many transfer learning researches (Pan and Yang, <xref ref-type="bibr" rid="B28">2010</xref>; Cao et al., <xref ref-type="bibr" rid="B4">2018</xref>; Zhang J. et al., <xref ref-type="bibr" rid="B44">2018</xref>). Since intelligibility is vital to our framework, utilizing features that are most relevant to intelligibility is a better way. Thus, we first propose a feature selection module to pick out more relevant features, and then fuse them with distortion features through an intelligibility enhanced module.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Proposed framework of IE-IQA. Our framework contains intelligibility and distortion backbone, and colorful blocks in the distortion backbone are trainable while gray blocks in the intelligibility backbone are not trainable. An intelligibility enhanced module is adopted to fuse distortion features with intelligibility features obtained from the proposed feature selection module.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-739138-g0003.tif"/>
</fig>
<p>The distortion backbone with parameter &#x003B8;<sup>&#x0002A;</sup> in <xref ref-type="fig" rid="F3">Figure 3</xref> is denoted as <inline-formula><mml:math id="M2"><mml:msub><mml:mrow><mml:mi>G</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:math></inline-formula>, which is adopted for extracting distortion features <italic>g</italic><sub><italic>j</italic></sub> from image <italic>I</italic>. The intelligibility backbone <italic>F</italic><sub>&#x003B8;</sub> with parameter &#x003B8; is adopted for extracting intelligibility features <italic>f</italic><sub><italic>j</italic></sub>. We select the most important features <inline-formula><mml:math id="M3"><mml:msubsup><mml:mi>f</mml:mi><mml:mi>j</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msubsup></mml:math></inline-formula> from <italic>f</italic><sub><italic>j</italic></sub> and then fuse them with distortion features <italic>g</italic><sub><italic>j</italic></sub> (denoted as <inline-formula><mml:math id="M4"><mml:msubsup><mml:mi>f</mml:mi><mml:mi>j</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msubsup><mml:mo>&#x021D4;</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:math></inline-formula>) to obtain quality score <italic>q</italic> through a regressor <inline-formula><mml:math id="M5"><mml:msub><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:math></inline-formula>. The whole process is explained as follows:</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M6"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:msub><mml:mi>f</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mi>&#x003B8;</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>I</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:msubsup><mml:mi>f</mml:mi><mml:mi>j</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msubsup><mml:mo>&#x02190;</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x000A0;&#x000A0;&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:msub><mml:mi>g</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>G</mml:mi><mml:mrow><mml:msup><mml:mi>&#x003B8;</mml:mi><mml:mo>*</mml:mo></mml:msup></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>I</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>q</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:msup><mml:mi>&#x003B8;</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msup></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mi>j</mml:mi><mml:mo>&#x02032;</mml:mo></mml:msubsup><mml:mo>&#x021D4;</mml:mo><mml:msub><mml:mi>g</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<p>In this paper, four extensively studied semantic understanding tasks are utilized to obtain intelligibility features, including image recognition on subset of ImageNet (Russakovsky et al., <xref ref-type="bibr" rid="B30">2015</xref>), scene classification on Places-365 (Zhou et al., <xref ref-type="bibr" rid="B49">2017</xref>), object detection and instance segmentation on MS-COCO (Lin et al., <xref ref-type="bibr" rid="B24">2014</xref>). In addition, we also utilize a relevant unrecognizability prediction task, which predicts the unrecognizable degree of an image. This task is trained on the VizWiz-QualityIssues dataset (Chiu et al., <xref ref-type="bibr" rid="B6">2020</xref>), containing images with labels of the unrecognizable degree. Even if intelligibility features of heavily distorted images cannot obtain desired results in original tasks, they can still be distinguished from features of high-quality images, which is beneficial to the IQA task.</p>
<p>In our framework, the distortion backbone works in a data-driven manner to search for the best distortion features, and the intelligibility backbone is guaranteed to obtain features with high generalization ability and rich semantic information. To achieve these goals, we propose to freeze parameters &#x003B8; of the intelligibility backbone during the training process while keeping parameters &#x003B8;<sup>&#x0002A;</sup> in the distortion backbone trainable. On the one hand, the distortion network loads the pre-trained model trained on ImageNet. Though the pre-trained model has decent generalization ability, we still need to train the feature extractor with image quality data so that the network can adapt to IQA task and obtain better performance. Therefore, we make parameters of the distortion backbone trainable. On the other hand, training the intelligibility backbone may be problematic. High level features of image understanding tasks are rich in semantic information which is generalizable. If we train the intelligibility network using the IQA data, the generalization ability of intelligibility features (which are already generalizable) may be destroyed. Therefore, we freeze the intelligibility backbone to handle this problem.</p>
<p>In the proposed intelligibility enhanced module, we tried several feature fusion strategies: (1) utilize one/two/three fully connected (FC) layers to regress the quality score and fuse intelligibility features to different FC layer with add/multiply/concatenate operation; (2) utilize other layers to align intelligibility features with distortion features and then use other FC layers to regress the quality score; (3) utilize auxiliary layers and loss fuction to train intelligibility features along with strategy-(1) or strategy-(2); (4) replace low-dimensional features with sparse selected features (features that are not selected are set to zero) and then utilize strategy-(1) or strategy-(2). In implementation, we have found that these strategies achieve similar results. Due to the feature selection module, it is easy to combine lower dimensional intelligibility features and simple strategy can obtain satisfying results. The loss function we utilized is the mean square error (MSE).</p>
</sec>
<sec>
<title>3.3. Feature Selection</title>
<p>During the feature fusion process, we propose strategies to select intelligibility features. For a specific semantic understanding task, only a part of neural units and corresponding features in a DCNN are significantly activated during the inference process, while others are not vital to the final prediction and intelligibility (Hu et al., <xref ref-type="bibr" rid="B17">2016</xref>; Zhang Q. et al., <xref ref-type="bibr" rid="B46">2018</xref>; Zhou et al., <xref ref-type="bibr" rid="B48">2019</xref>). Since introducing too many features are not conducive (even harmful in many transfer learning experiments) to IQA performance and generalization, we design feature selection strategies for different tasks based on contribution and sensitivity. Contribution-based strategy chooses features with greater contributions to predictions while sensitivity-based strategy chooses features that predictions are more sensitive to.</p>
<sec>
<title>3.3.1. Contribution-Based Strategy</title>
<p>We propose to select features that have prominent contributions to final predictions. Theoretically, this strategy is not limited to any specified network as long as the network can be separated into a backbone and one FC layer. In fact, this kind of network architecture is very common in the image classification and scene recognition. Specifically, the output of backbone can be denoted by <italic>f</italic><sub><italic>j</italic></sub>, <italic>j</italic> &#x0003D; 1, 2..., <italic>N</italic><sub><italic>d</italic></sub>, where <italic>N</italic><sub><italic>d</italic></sub> is the dimension of features and the output-logits of the FC layer can be described by</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>w</italic><sub><italic>ij</italic></sub>, <italic>b</italic><sub><italic>i</italic></sub>, <italic>z</italic><sub><italic>i</italic></sub> are weights, bias and logits of the FC layer, <italic>i</italic> &#x0003D; 1, 2, ..., <italic>C</italic>, and <italic>C</italic> is the number of total classes. The feature selection strategy is shown in <xref ref-type="table" rid="T6">Algorithm 1</xref>. In <xref ref-type="table" rid="T6">Algorithm 1</xref>, we locate the top1-probability first. Then, we calculate the contribution of each dimension of feature <italic>f</italic><sub><italic>j</italic></sub> by</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M8"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<table-wrap position="float" id="T6">
<label>Algorithm 1</label>
<caption><p>Feature selection strategy based on contributions.</p></caption>
<graphic xlink:href="fnins-15-739138-i0001.tif"/>
</table-wrap>
<p>In Equation (4), the contributions of features are determined by both weights and activation values. Finally, features that contribute significantly to the top1-probability are selected.</p>
</sec>
<sec>
<title>3.3.2. Sensitivity-Based Strategy</title>
<p>Some networks have several non-linear FC layers and it is not easy to measure their contributions. Consider the unrecognizability prediction task for example. First, we train a model with the backbone of EfficientNet-B0 (Tan and Le, <xref ref-type="bibr" rid="B35">2019</xref>) and three FC layers with RELU to regress the unrecognizability score. Then, we adopt a sensitivity-based method to select features and the sensitivity can be obtained by gradients. Specifically, the input feature is <italic>f</italic><sub><italic>j</italic></sub>, <italic>j</italic> &#x0003D; 1, 2, ...<italic>N</italic><sub><italic>d</italic></sub> and the FC layers with active function are represented by function <italic>F</italic>. The unrecognizability score <italic>s</italic> can be obtained by</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M9"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mi>F</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The sensitivity of features can be described by</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M10"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>g</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>s</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x02202;</mml:mi><mml:mi>s</mml:mi><mml:mo>/</mml:mo><mml:mi>&#x02202;</mml:mi><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Equation (6) represents the importance of features through partial derivatives, which is widely used in sensitivity analysis and model interpreting (Garson, <xref ref-type="bibr" rid="B11">1991</xref>; Dimopoulos et al., <xref ref-type="bibr" rid="B8">1995</xref>). After obtaining the importance of features, the selected number is calculated. Then, the index of sorted features can be obtained through</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M11"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>g</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x0201C;<italic>argsort</italic>&#x0201D; means that sort the sequence and return corresponding index (it is the same with &#x0201C;<italic>argsort</italic>&#x0201D; in <xref ref-type="table" rid="T6">Algorithm 1</xref>). Finally, a selection operation is executed.</p>
<p>In contrast to directly merging all intelligibility features with distortion features, fusing features with lower dimension after feature selection exhibits better performance and generalization ability during the test process. Different from attention mechanism, the proposed feature selection strategy can reduce the dimension of the intelligibility feature and does not need any additional module or further training.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<title>4. Experiments</title>
<sec>
<title>4.1. Datasets</title>
<p>In our experiments, five image quality datasets with authentic distortions are adopted, including KonIQ-10k (Hosu et al., <xref ref-type="bibr" rid="B16">2020</xref>), Smartphone Photography Attribute and Quality (SPAQ) (Fang et al., <xref ref-type="bibr" rid="B10">2020</xref>), LIVE in the Wild Image Quality Challenge (LIVEW) (Ghadiyaram and Bovik, <xref ref-type="bibr" rid="B12">2016</xref>), CID2013 (Virtanen et al., <xref ref-type="bibr" rid="B38">2015</xref>), and BID (Ciancio et al., <xref ref-type="bibr" rid="B7">2011</xref>). Specifically, the KonIQ-10k dataset has 10,073 labeled images selected from a massive public database YFCC100M (Thomee et al., <xref ref-type="bibr" rid="B36">2016</xref>), and the labels are obtained from 1.2 million ratings. The SPAQ dataset contains 11,125 labeled images obtained from 66 smartphones with exchangeable image file format data tags and rich opinion annotations. The annotations include MOS, attribute scores (such as brightness, noisiness, and sharpness) as well as scene category labels. LIVEW contains 1,162 labeled images and CID2013 contains 480 images from eight scenes. Different from the other four datasets, the BID dataset focuses on blur images and contains 586 images.</p>
</sec>
<sec>
<title>4.2. Implementation and Evaluation Protocol</title>
<p>In our experiments, the distortion network adopts the backbone of EfficientNet-B0 and the intelligibility network for the image recognition task is EfficientNet-B0 as well. EfficientNet-B0 consists of one convolutional layer followed by seven mobile inverted bottleneck modules, and then another convolutional layer followed by global average pooling. EfficientNet-B0 has an input size of 224 &#x000D7; 224 and 5.3 M parameters, and the dimension of its output feature is 1280. Network for scene classification task is ResNet-18 (He et al., <xref ref-type="bibr" rid="B15">2016</xref>), and object detection is Faster-RCNN (Ren et al., <xref ref-type="bibr" rid="B29">2017</xref>) with ResNet50-FPN (Lin et al., <xref ref-type="bibr" rid="B23">2017</xref>) backbone. The instance segmentation task is DeeplabV3&#x0002B; (Chen et al., <xref ref-type="bibr" rid="B5">2018</xref>) with the backbone of ResNet101. During the training process, SGD optimizer with initial learning-rate 0.03 is utilized (we train FC layers first and then utilize warm-up strategy when training the distortion backbone). For all of our experiments, we first resize images into 244 &#x000D7; 244, then we randomly crop them to 224 &#x000D7; 224 with a randomly horizontal flip to augment training images. During the test process, we directly resize test images into 224 &#x000D7; 224 and then predict once, which is more efficient in real applications. We tried different selection ratios of 1, 5, 20, and 50%. The final selection ratio of the recognition task, class task, detection task, segmentation task, and unrecognization task are 5, 5, 20, 50, and 50%, respectively.</p>
<p>Our evaluation criteria are two widely used correlation coefficients: Pearson&#x00027;s linear correlation coefficient (PLCC) and Spearman&#x00027;s rank order correlation coefficient (SRCC).</p>
</sec>
<sec>
<title>4.3. Performance Comparison</title>
<p>This paper aims to propose a highly generalizable NR-IQA model, thus we train our model in one dataset and then test on other datasets directly without doing any fine-tuning. For comparison, we also re-train some popular handcrafted feature-based methods, such as BRISQUE, CORNIA, HOSA, and deep learning-based methods, including DBCNN (Zhang W. et al., <xref ref-type="bibr" rid="B47">2020</xref>), MetaIQA (Zhu et al., <xref ref-type="bibr" rid="B50">2020</xref>), and WaDIQaM-NR (Bosse et al., <xref ref-type="bibr" rid="B3">2017</xref>) (codes are publically available) with the same setting. All results trained on KonIQ-10k are shown in <xref ref-type="table" rid="T1">Table 1</xref>. The middle group in <xref ref-type="table" rid="T1">Table 1</xref> shows deep learning-based methods, and the results of methods without public codes are obtained from the original papers. The bottom group shows our results.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Pearson&#x00027;s linear correlation coefficient (PLCC)/Spearman&#x00027;s rank order correlation coefficient (SRCC) results of cross-dataset test.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>PLCC/SRCC</bold></th>
<th valign="top" align="center"><bold>KonIQ-10k</bold></th>
<th valign="top" align="center"><bold>SPAQ</bold></th>
<th valign="top" align="center"><bold>LIVEW</bold></th>
<th valign="top" align="center"><bold>CID</bold></th>
<th valign="top" align="center"><bold>BID</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">BIQI Moorthy and Bovik, <xref ref-type="bibr" rid="B27">2010</xref></td>
<td valign="top" align="center">0.637/0.595</td>
<td valign="top" align="center">0.622/0.661</td>
<td valign="top" align="center">0.492/0.471</td>
<td valign="top" align="center">0.612/0.599</td>
<td valign="top" align="center">0.478/0.493</td>
</tr>
<tr>
<td valign="top" align="left">NFERM Gu et al., <xref ref-type="bibr" rid="B14">2014</xref></td>
<td valign="top" align="center">0.725/0.689</td>
<td valign="top" align="center">0.697/0.711</td>
<td valign="top" align="center">0.551/0.540</td>
<td valign="top" align="center">0.708/0.680</td>
<td valign="top" align="center">0.529/0.530</td>
</tr>
<tr>
<td valign="top" align="left">BRISQUE Mittal et al., <xref ref-type="bibr" rid="B26">2012</xref></td>
<td valign="top" align="center">0.689/0.647</td>
<td valign="top" align="center">0.660/0.682</td>
<td valign="top" align="center">0.576/0.554</td>
<td valign="top" align="center">0.553/0.533</td>
<td valign="top" align="center">0.589/0.597</td>
</tr>
<tr>
<td valign="top" align="left">BLINDSII Saad et al., <xref ref-type="bibr" rid="B31">2012</xref></td>
<td valign="top" align="center">0.440/0.447</td>
<td valign="top" align="center">0.466/0.460</td>
<td valign="top" align="center">0.331/0.319</td>
<td valign="top" align="center">0.278/0.301</td>
<td valign="top" align="center">0.393/0.401</td>
</tr>
<tr>
<td valign="top" align="left">GWH-GLBP Li et al., <xref ref-type="bibr" rid="B22">2016</xref></td>
<td valign="top" align="center">0.549/0.514</td>
<td valign="top" align="center">0.614/0.628</td>
<td valign="top" align="center">0.464/0.435</td>
<td valign="top" align="center">0.071/0.002</td>
<td valign="top" align="center">0.477/0.483</td>
</tr>
<tr>
<td valign="top" align="left">FISBLIM Gu et al., <xref ref-type="bibr" rid="B13">2013</xref></td>
<td valign="top" align="center">0.375/0.347</td>
<td valign="top" align="center">0.566/0.569</td>
<td valign="top" align="center">0.376/0.289</td>
<td valign="top" align="center">-0.219/-0.234</td>
<td valign="top" align="center">0.392/0.344</td>
</tr>
<tr>
<td valign="top" align="left">CORNIA Ye et al., <xref ref-type="bibr" rid="B42">2012</xref></td>
<td valign="top" align="center">0.773/0.738</td>
<td valign="top" align="center">0.727/0.766</td>
<td valign="top" align="center">0.672/0.639</td>
<td valign="top" align="center">0.599/0.538</td>
<td valign="top" align="center">0.692/0.688</td>
</tr>
<tr>
<td valign="top" align="left">HOSA Xu et al., <xref ref-type="bibr" rid="B40">2016</xref></td>
<td valign="top" align="center">0.791/0.761</td>
<td valign="top" align="center">0.743/0.771</td>
<td valign="top" align="center">0.677/0.652</td>
<td valign="top" align="center">0.684/0.664</td>
<td valign="top" align="center">0.694/0.679</td>
</tr>
<tr>
<td valign="top" align="left">NSSADNN Yan et al., <xref ref-type="bibr" rid="B41">2019</xref></td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">0.813/0.745&#x0002A;</td>
<td valign="top" align="center">0.825/0.748&#x0002A;</td>
<td valign="top" align="center">/</td>
</tr>
<tr>
<td valign="top" align="left">MEON Ma et al., <xref ref-type="bibr" rid="B25">2017</xref></td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">0.693/0.688&#x0002A;</td>
<td valign="top" align="center">0.703/0.701&#x0002A;</td>
<td valign="top" align="center">/</td>
</tr>
<tr>
<td valign="top" align="left">BIECON Kim and Lee, <xref ref-type="bibr" rid="B20">2017</xref></td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">0.613/0.595&#x0002A;</td>
<td valign="top" align="center">0.620/0.606&#x0002A;</td>
<td valign="top" align="center">/</td>
</tr>
<tr>
<td valign="top" align="left">DeepRN (ResNet101) Varga et al., <xref ref-type="bibr" rid="B37">2018</xref></td>
<td valign="top" align="center">0.880/0.867</td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">0.750/0.726</td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">/</td>
</tr>
<tr>
<td valign="top" align="left">DeepBIQ (InceptionV2) Bianco et al., <xref ref-type="bibr" rid="B2">2018</xref></td>
<td valign="top" align="center">0.911/0.907</td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">0.821/0.804</td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">/</td>
</tr>
<tr>
<td valign="top" align="left">HyperNet Su et al., <xref ref-type="bibr" rid="B34">2020</xref></td>
<td valign="top" align="center">0.917/<bold>0.906</bold></td>
<td valign="top" align="center">0.843/0.846<sup>&#x0002B;</sup></td>
<td valign="top" align="center">NA/0.785</td>
<td valign="top" align="center">0.808/0.782<sup>&#x0002B;</sup></td>
<td valign="top" align="center">NA/<bold>0.819</bold></td>
</tr>
<tr>
<td valign="top" align="left">MetaIQA Zhu et al., <xref ref-type="bibr" rid="B50">2020</xref></td>
<td valign="top" align="center">0.876/0.846</td>
<td valign="top" align="center">0.804/0.822</td>
<td valign="top" align="center">0.748/0.716</td>
<td valign="top" align="center">0.726/0.682</td>
<td valign="top" align="center">0.740/0.738</td>
</tr>
<tr>
<td valign="top" align="left">WaDIQaM-NR Bosse et al., <xref ref-type="bibr" rid="B3">2017</xref></td>
<td valign="top" align="center">0.657/0.631</td>
<td valign="top" align="center">0.675/0.702</td>
<td valign="top" align="center">0.521/0.523</td>
<td valign="top" align="center">0.584/0.495</td>
<td valign="top" align="center">0.499/0.538</td>
</tr>
<tr>
<td valign="top" align="left">DBCNN Zhang W. et al., <xref ref-type="bibr" rid="B47">2020</xref></td>
<td valign="top" align="center">0.892/0.868</td>
<td valign="top" align="center">0.827/0.836</td>
<td valign="top" align="center">0.802/0.775</td>
<td valign="top" align="center">0.788/0.758</td>
<td valign="top" align="center">0.769/0.769</td>
</tr>
<tr>
<td valign="top" align="left" colspan="6"><bold>Our Results</bold></td>
</tr>
<tr>
<td valign="top" align="left"><bold>IE-IQA (</bold><italic>w/ recognition task)</italic></td>
<td valign="top" align="center"><bold>0.921</bold>/0.900</td>
<td valign="top" align="center"><bold>0.863/0.859</bold></td>
<td valign="top" align="center"><bold>0.839/0.829</bold></td>
<td valign="top" align="center">0.815/0.788</td>
<td valign="top" align="center"><bold>0.822</bold>/0.817</td>
</tr>
<tr>
<td valign="top" align="left"><bold>IE-IQA (</bold><italic>w/ classification task)</italic></td>
<td valign="top" align="center">0.920/0.900</td>
<td valign="top" align="center">0.862/0.858</td>
<td valign="top" align="center">0.835/0.828</td>
<td valign="top" align="center">0.818/0.795</td>
<td valign="top" align="center">0.819/0.813</td>
</tr>
<tr>
<td valign="top" align="left"><bold>IE-IQA (</bold><italic>w/ detection task)</italic></td>
<td valign="top" align="center"><bold>0.921/</bold>0.901</td>
<td valign="top" align="center">0.862/0.857</td>
<td valign="top" align="center">0.835/0.826</td>
<td valign="top" align="center">0.819/0.800</td>
<td valign="top" align="center">0.816/0.810</td>
</tr>
<tr>
<td valign="top" align="left"><bold>IE-IQA (</bold><italic>w/ segmentation task)</italic></td>
<td valign="top" align="center">0.917/0.900</td>
<td valign="top" align="center">0.862/0.857</td>
<td valign="top" align="center">0.825/0.826</td>
<td valign="top" align="center"><bold>0.827/0.801</bold></td>
<td valign="top" align="center">0.812/0.809</td>
</tr>
<tr>
<td valign="top" align="left"><bold>IE-IQA (</bold><italic>w/ unrecognization task)</italic></td>
<td valign="top" align="center">0.920/ 0.902</td>
<td valign="top" align="center"><bold>0.863</bold>/0.858</td>
<td valign="top" align="center">0.835/<bold>0.829</bold></td>
<td valign="top" align="center">0.819/0.794</td>
<td valign="top" align="center">0.816/0.813</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The model is trained on 80% images of KonIQ-10k and directly tested on rest 20% KonIQ-10k images and other datasets. Results with &#x0201C;&#x0002A;&#x0201D; are obtained after fine-tuning on the dataset and reported in original papers. Results with NA of HyperNet (only three datasets) are reported in original papers (Su et al., <xref ref-type="bibr" rid="B34">2020</xref>) and results with &#x0201C;&#x0002B;&#x0201D; are obtained from the released model. Best results are in bold</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>From <xref ref-type="table" rid="T1">Table 1</xref>, we can observe that our framework with five intelligibility tasks can consistently achieve the best cross-dataset performance for most cases. It should be emphasized that our models are only trained with KonIQ-10k (80% images) and directly tested on other datasets without any fine-tuning. Though NSSADNN, MEON, and BIECON made fine-tuning on the target dataset, our generalization performance can still maintain a significant advantage.</p>
<p>Efficient-B0 has 5.3M parameters, which is less than ResNet18 (11.7 M parameters, the backbone of MetaIQA), ResNet50 (26 M parameters, the backbone of HyperNet), and ResNet101 (44.5 M parameters, the backbone of DeepRN). Efficient-B0 is easy to converge, and we show the loss and PLCC results during training and test in <xref ref-type="fig" rid="F4">Figure 4</xref>. We can observe from <xref ref-type="fig" rid="F4">Figure 4</xref> that the test loss decreases with the training loss and the test performance increases with training performance. This means that the network is trained without overfitting.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Loss and Pearson&#x00027;s linear correlation coefficient (PLCC) during training and test. <bold>(A)</bold> Loss of training and test. Two enlarged subfigures shows results of epochs 2&#x02013;50 and epochs 200&#x02013;250. <bold>(B)</bold> PLCC of training and test. The model is trained with recognition task on KonIQ-10k.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-739138-g0004.tif"/>
</fig>
<p>To make a further comparison, we also train our methods on SPAQ and perform cross-dataset tests on the other four datasets. The results are shown in <xref ref-type="table" rid="T2">Table 2</xref>.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Pearson&#x00027;s linear correlation coefficient (PLCC)/Spearman&#x00027;s rank order correlation coefficient (SRCC) results of cross-dataset test.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>PLCC/SRCC</bold></th>
<th valign="top" align="center"><bold>SPAQ</bold></th>
<th valign="top" align="center"><bold>KonIQ-10k</bold></th>
<th valign="top" align="center"><bold>LIVEW</bold></th>
<th valign="top" align="center"><bold>CID</bold></th>
<th valign="top" align="center"><bold>BID</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">NFERM Gu et al., <xref ref-type="bibr" rid="B14">2014</xref></td>
<td valign="top" align="center">0.832/0.823</td>
<td valign="top" align="center">0.455/0.447</td>
<td valign="top" align="center">0.591/0.542</td>
<td valign="top" align="center">0.437/0.342</td>
<td valign="top" align="center">0.578/0.570</td>
</tr>
<tr>
<td valign="top" align="left">BRSIQUE Mittal et al., <xref ref-type="bibr" rid="B26">2012</xref></td>
<td valign="top" align="center">0.833/0.822</td>
<td valign="top" align="center">0.446/0.433</td>
<td valign="top" align="center">0.593/0.553</td>
<td valign="top" align="center">0.499/0.504</td>
<td valign="top" align="center">0.589/0.578</td>
</tr>
<tr>
<td valign="top" align="left">CORNIA Ye et al., <xref ref-type="bibr" rid="B42">2012</xref></td>
<td valign="top" align="center">0.867/0.859</td>
<td valign="top" align="center">0.532/0.516</td>
<td valign="top" align="center">0.663/0.621</td>
<td valign="top" align="center">0.552/0.465</td>
<td valign="top" align="center">0.676/0.673</td>
</tr>
<tr>
<td valign="top" align="left">HOSA Xu et al., <xref ref-type="bibr" rid="B40">2016</xref></td>
<td valign="top" align="center">0.873/0.866</td>
<td valign="top" align="center">0.559/0.534</td>
<td valign="top" align="center">0.682/0.650</td>
<td valign="top" align="center">0.593/0.536</td>
<td valign="top" align="center">0.681/0.670</td>
</tr>
<tr>
<td valign="top" align="left">Baseline Fang et al., <xref ref-type="bibr" rid="B10">2020</xref></td>
<td valign="top" align="center">0.909/0.908</td>
<td valign="top" align="center">0.532/0.523<sup>&#x0002B;</sup></td>
<td valign="top" align="center">0.564/0.517<sup>&#x0002B;</sup></td>
<td valign="top" align="center">0.518/0.569<sup>&#x0002B;</sup></td>
<td valign="top" align="center">0.574/0.566<sup>&#x0002B;</sup></td>
</tr>
<tr>
<td valign="top" align="left">MT-S Fang et al., <xref ref-type="bibr" rid="B10">2020</xref></td>
<td valign="top" align="center"><bold>0.921/0.917</bold></td>
<td valign="top" align="center">0.486/0.485<sup>&#x0002B;</sup></td>
<td valign="top" align="center">0.539/0.493<sup>&#x0002B;</sup></td>
<td valign="top" align="center">0.342/0.389<sup>&#x0002B;</sup></td>
<td valign="top" align="center">0.530/0.529<sup>&#x0002B;</sup></td>
</tr>
<tr>
<td valign="top" align="left">HyperNet Su et al., <xref ref-type="bibr" rid="B34">2020</xref></td>
<td valign="top" align="center">0.917/0.915</td>
<td valign="top" align="center">0.679/0.645</td>
<td valign="top" align="center">0.695/0.680</td>
<td valign="top" align="center">0.624/0.585</td>
<td valign="top" align="center">0.648/0.647</td>
</tr>
<tr>
<td valign="top" align="left">MetaIQA Zhu et al., <xref ref-type="bibr" rid="B50">2020</xref></td>
<td valign="top" align="center">0.871/0.870</td>
<td valign="top" align="center">0.722/0.686</td>
<td valign="top" align="center">0.765/0.731</td>
<td valign="top" align="center">0.737/0.695</td>
<td valign="top" align="center">0.743/0.735</td>
</tr>
<tr>
<td valign="top" align="center" colspan="6"><bold>Our Results</bold></td>
</tr>
<tr>
<td valign="top" align="left"><bold>IE-IQA (</bold><italic>w/ recognition task)</italic></td>
<td valign="top" align="center">0.918/0.913</td>
<td valign="top" align="center">0.768/0.710</td>
<td valign="top" align="center">0.779/0.764</td>
<td valign="top" align="center">0.743/0.713</td>
<td valign="top" align="center">0.744/0.742</td>
</tr>
<tr>
<td valign="top" align="left"><bold>IE-IQA (</bold><italic>w/ classification task)</italic></td>
<td valign="top" align="center">0.917/0.915</td>
<td valign="top" align="center">0.761/0.720</td>
<td valign="top" align="center">0.764/0.758</td>
<td valign="top" align="center">0.737/0.702</td>
<td valign="top" align="center">0.737/0.737</td>
</tr>
<tr>
<td valign="top" align="left"><bold>IE-IQA (</bold><italic>w/ detection task)</italic></td>
<td valign="top" align="center">0.920/0.916</td>
<td valign="top" align="center"><bold>0.777/0.728</bold></td>
<td valign="top" align="center"><bold>0.782/0.772</bold></td>
<td valign="top" align="center">0.742/0.702</td>
<td valign="top" align="center"><bold>0.748/0.749</bold></td>
</tr>
<tr>
<td valign="top" align="left"><bold>IE-IQA (</bold><italic>w/ segmentation task)</italic></td>
<td valign="top" align="center">0.918/0.914</td>
<td valign="top" align="center">0.775/0.724</td>
<td valign="top" align="center">0.781/0.768</td>
<td valign="top" align="center"><bold>0.752/0.737</bold></td>
<td valign="top" align="center">0.744/0.746</td>
</tr>
<tr>
<td valign="top" align="left"><bold>IE-IQA (</bold><italic>w/ unrecognization task)</italic></td>
<td valign="top" align="center">0.920/0.916</td>
<td valign="top" align="center">0.770/0.721</td>
<td valign="top" align="center">0.774/0.764</td>
<td valign="top" align="center"><bold>0.752</bold>/0.725</td>
<td valign="top" align="center">0.747/0.746</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The model is trained on 80% images of Smartphone Photography Attribute and Quality (SPAQ) and directly tested on rest 20% SPAQ images and other datasets. Results with &#x0201C;&#x0002B;&#x0201D; are obtained from the released model. HyperNet are retrained with image size of 244 &#x000D7; 244. Best results are in bold</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>The model &#x0201C;Baseline&#x0201D; in Fang et al. (<xref ref-type="bibr" rid="B10">2020</xref>) means the baseline model (ResNet50) and &#x0201C;MT-S&#x0201D; means the model jointly trained with MOS and scene labels (The SPAQ dataset has scene category labels). We can observe that compared to MT-S, our method can achieve similar performance on the training dataset. However, by combining intelligibility features, the generalization performance of the proposed method is apparently much better.</p>
<p>Comparing <xref ref-type="table" rid="T2">Table 2</xref> with <xref ref-type="table" rid="T1">Table 1</xref>, we can observe that models trained on KonIQ-10k have better cross-dataset performance. One possible reason is the source of images. The SPAQ dataset is obtained from smartphones only, while the image sources of KonIQ-10k are more diversified. Another possible reason is that the image size of the SPAQ dataset is very large (4000 &#x000D7; 3000 is common) and our model has an input size of 224 &#x000D7; 224. Small size input may lose much information and the interpolation algorithm may bring new distortions.</p>
<p>Another phenomenon observed from <xref ref-type="table" rid="T1">Tables 1</xref>, <xref ref-type="table" rid="T2">2</xref> is that the proposed method achieves slightly worse generalization performance on the BID/CID databases than the other three datasets. The BID dataset focuses on blur images and the CID dataset consists of limited scenes of images (eight scenes). This may lead to a more pronounced distribution discrepancy between CID/BID and the training datasets.</p>
<p>Though our metric aims to achieve high generalization ability, we still make further experiments on intra-dataset tests. The results are listed in <xref ref-type="table" rid="T3">Table 3</xref>. We can summarize from <xref ref-type="table" rid="T3">Table 3</xref> that our metric can achieve state-of-the-arts intra-dataset performance. Though HyperNet achieves better performance for some cases, it needs to evaluate crop 25 patches during evaluating, costing much more time than the proposed metric. For example, when evaluating 1,000 images with the resolution of 1024 &#x000D7; 768 (batchsize = 1, using one TITANXp GPU and Intel Xeon E5-2630V4 CPU), HyperNet costs 2,040 s, while the proposed metric only costs 84 s.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Pearson&#x00027;s linear correlation coefficient (PLCC)/Spearman&#x00027;s rank order correlation coefficient (SRCC) results on intra-dataset tests.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Dataset</bold></th>
<th valign="top" align="center"><bold>KonIQ-10k</bold></th>
<th valign="top" align="center"><bold>SPAQ</bold></th>
<th valign="top" align="center"><bold>LIVEW</bold></th>
<th valign="top" align="center"><bold>CID2013</bold></th>
<th valign="top" align="center"><bold>RBID</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">NFERM Gu et al., <xref ref-type="bibr" rid="B14">2014</xref></td>
<td valign="top" align="center">0.725/0.689</td>
<td valign="top" align="center">0.832/0.823</td>
<td valign="top" align="center">0.562 /0.517</td>
<td valign="top" align="center">0.825/0.823</td>
<td valign="top" align="center">0.585/0.559</td>
</tr>
<tr>
<td valign="top" align="left">BRISQUE Mittal et al., <xref ref-type="bibr" rid="B26">2012</xref></td>
<td valign="top" align="center">0.689/0.647</td>
<td valign="top" align="center">0.833/0.822</td>
<td valign="top" align="center">0.574/0.557</td>
<td valign="top" align="center">0.810/0.814</td>
<td valign="top" align="center">0.617/0.594</td>
</tr>
<tr>
<td valign="top" align="left">CORNIA Ye et al., <xref ref-type="bibr" rid="B42">2012</xref></td>
<td valign="top" align="center">0.773/0.738</td>
<td valign="top" align="center">0.867/0.859</td>
<td valign="top" align="center">0.692/0.655</td>
<td valign="top" align="center">0.822/0.803</td>
<td valign="top" align="center">0.712/0.695</td>
</tr>
<tr>
<td valign="top" align="left">HOSA Xu et al., <xref ref-type="bibr" rid="B40">2016</xref></td>
<td valign="top" align="center">0.791/0.761</td>
<td valign="top" align="center">0.873/0.866</td>
<td valign="top" align="center">0.703/0.667</td>
<td valign="top" align="center">0.835/0.833</td>
<td valign="top" align="center">0.716/0.684</td>
</tr>
<tr>
<td valign="top" align="left">NSSADNN Yan et al., <xref ref-type="bibr" rid="B41">2019</xref></td>
<td valign="top" align="center">&#x000A0;/</td>
<td valign="top" align="center">&#x000A0;/</td>
<td valign="top" align="center">0.813<sup>&#x0002A;</sup>/0.745<sup>&#x0002A;</sup></td>
<td valign="top" align="center">0.825<sup>&#x0002A;</sup>/0.748<sup>&#x0002A;</sup></td>
<td valign="top" align="center">&#x000A0;/</td>
</tr>
<tr>
<td valign="top" align="left">MEON Ma et al., <xref ref-type="bibr" rid="B25">2017</xref></td>
<td valign="top" align="center">&#x000A0;/</td>
<td valign="top" align="center">&#x000A0;/</td>
<td valign="top" align="center">0.693<sup>&#x0002A;</sup>/0.688<sup>&#x0002A;</sup></td>
<td valign="top" align="center">0.703<sup>&#x0002A;</sup> / 0.701<sup>&#x0002A;</sup></td>
<td valign="top" align="center">&#x000A0;/</td>
</tr>
<tr>
<td valign="top" align="left">BIECON Kim and Lee, <xref ref-type="bibr" rid="B20">2017</xref></td>
<td valign="top" align="center">&#x000A0;/</td>
<td valign="top" align="center">&#x000A0;/</td>
<td valign="top" align="center">0.613<sup>&#x0002A;</sup>/0.595<sup>&#x0002A;</sup></td>
<td valign="top" align="center">0.620<sup>&#x0002A;</sup>/0.606<sup>&#x0002A;</sup></td>
<td valign="top" align="center">&#x000A0;/</td>
</tr>
<tr>
<td valign="top" align="left">Baseline Fang et al., <xref ref-type="bibr" rid="B10">2020</xref></td>
<td valign="top" align="center">0.908/0.889</td>
<td valign="top" align="center">0.909<sup>&#x0002A;</sup>/0.908<sup>&#x0002A;</sup></td>
<td valign="top" align="center">0.825/0.794</td>
<td valign="top" align="center">0.876/0.881</td>
<td valign="top" align="center">0.802/0.794</td>
</tr>
<tr>
<td valign="top" align="left">WaDIQaM-NR Bosse et al., <xref ref-type="bibr" rid="B3">2017</xref></td>
<td valign="top" align="center">0.805<sup>&#x0002A;</sup>/0.797<sup>&#x0002A;</sup></td>
<td valign="top" align="center">&#x000A0;/</td>
<td valign="top" align="center">0.680<sup>&#x0002A;</sup>/0.671<sup>&#x0002A;</sup></td>
<td valign="top" align="center">0.729<sup>&#x0002A;</sup>/0.708<sup>&#x0002A;</sup></td>
<td valign="top" align="center">0.742<sup>&#x0002A;</sup>/0.725<sup>&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">HyperNet Su et al., <xref ref-type="bibr" rid="B34">2020</xref></td>
<td valign="top" align="center">0.917<sup>&#x0002A;</sup>/<bold>0.906</bold><sup>&#x0002A;</sup></td>
<td valign="top" align="center">0.914/0.909</td>
<td valign="top" align="center"><bold>0.882</bold><sup>&#x0002A;</sup><bold>/0.859</bold><sup>&#x0002A;</sup></td>
<td valign="top" align="center">/</td>
<td valign="top" align="center"><bold>0.878</bold><sup>&#x0002A;</sup>/<bold>0.869</bold><sup>&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">DBCNN Zhang W. et al., <xref ref-type="bibr" rid="B47">2020</xref></td>
<td valign="top" align="center">0.892/0.868</td>
<td valign="top" align="center">0.915<sup>&#x0002A;</sup>/0.911<sup>&#x0002A;</sup></td>
<td valign="top" align="center">0.869<sup>&#x0002A;</sup>/0.851<sup>&#x0002A;</sup></td>
<td valign="top" align="center">/</td>
<td valign="top" align="center">0.859<sup>&#x0002A;</sup>/0.845<sup>&#x0002A;</sup></td>
</tr>
<tr>
<td valign="top" align="left">MetaIQA Zhu et al., <xref ref-type="bibr" rid="B50">2020</xref></td>
<td valign="top" align="center">0.887<sup>&#x0002A;</sup>/0.850<sup>&#x0002A;</sup></td>
<td valign="top" align="center">0.871/0.870</td>
<td valign="top" align="center">0.835<sup>&#x0002A;</sup>/0.802<sup>&#x0002A;</sup></td>
<td valign="top" align="center">0.784<sup>&#x0002A;</sup>/0.766<sup>&#x0002A;</sup></td>
<td valign="top" align="center">0.777/0.746</td>
</tr>
<tr>
<td valign="top" align="left"><bold>IE-IQA (</bold><italic>w/ recognition task)</italic></td>
<td valign="top" align="center"><bold>0.921/</bold>0.900</td>
<td valign="top" align="center"><bold>0.918/0.913</bold></td>
<td valign="top" align="center">0.868/0.838</td>
<td valign="top" align="center"><bold>0.934/0.934</bold></td>
<td valign="top" align="center">0.838/0.837</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Results with &#x0002A; are obtained from published papers. Other results are obtained from retrained model. Best results are marked in bold</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>To explore how feature selection strategy affects prediction results, we make a comparison of the results with/without feature selection strategy, and show them in <xref ref-type="fig" rid="F5">Figure 5</xref>. The results show that removing noisy features and utilizing features having significant influence on final predictions tend to achieve higher performance and better generalization ability with only one exception (recognition task) on the CID dataset. One possible reason is that the CID dataset has only eight specific scenes and many images in CID contain the same objects. In this situation, selected features may not provide rich distinguished information for evaluating quality of images with similar contents.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Performance comparison for different tasks with/without feature selection strategy on KonIQ-10k. <bold>(A)</bold> Pearson&#x00027;s linear correlation coefficient (PLCC) results. <bold>(B)</bold> Spearman&#x00027;s rank order correlation coefficient (SRCC) results. The recognition and classification task utilize contribution-based strategy, and the unrecognizability task utilizes gradient-based strategy.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-739138-g0005.tif"/>
</fig>
<p>To demonstrate the effectiveness of intelligibility features, we make ablation studies and show the results in <xref ref-type="fig" rid="F6">Figure 6</xref>. The baseline means the model with distortion backbone alone. From <xref ref-type="fig" rid="F6">Figure 6</xref>, we can observe that intelligibility features do improve both performance and generalization ability. Therefore, it is necessary to combine both intelligibility aspect and distortion aspect in IQA metrics.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Ablation study of intelligibility enriched IQA (IE-IQA). <bold>(A)</bold> Pearson&#x00027;s linear correlation coefficient (PLCC) results of models trained with 80% KonIQ-10k. <bold>(B)</bold> Spearman&#x00027;s rank order correlation coefficient (SRCC) results of models trained with 80% KonIQ-10k. <bold>(C)</bold> PLCC results of models trained with 80% Smartphone Photography Attribute and Quality (SPAQ). <bold>(D)</bold> SRCC results of models trained with 80% SPAQ.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-739138-g0006.tif"/>
</fig>
<p>During the training process, the distortion network loads the pre-trained model, and some semantic information and intelligibility features may have already existed in the pre-trained model. To further investigate the effects of original intelligibility features on the distortion network, we train the distortion network from scratch. Then we fuse the intelligibility network with the distortion network. The results are shown in <xref ref-type="table" rid="T4">Tables 4</xref>, <xref ref-type="table" rid="T5">5</xref>. From the tables, we can observe that the introduced intelligibility network still benefits the performance of the whole framework even the distortion network is not pre-trained.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Results of training the distortion network from scratch on 80% KonIQ-10k.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>PLCC</bold></th>
<th valign="top" align="center"><bold>KonIQ-10k(20%)</bold></th>
<th valign="top" align="center"><bold>SPAQ</bold></th>
<th valign="top" align="center"><bold>LIVEW</bold></th>
<th valign="top" align="center"><bold>CID</bold></th>
<th valign="top" align="center"><bold>BID</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Only distortion</td>
<td valign="top" align="center">0.784</td>
<td valign="top" align="center">0.756</td>
<td valign="top" align="center">0.638</td>
<td valign="top" align="center">0.676</td>
<td valign="top" align="center">0.645</td>
</tr>
<tr>
<td valign="top" align="left">W/recognition</td>
<td valign="top" align="center">0.814</td>
<td valign="top" align="center"><bold>0.812</bold></td>
<td valign="top" align="center"><bold>0.689</bold></td>
<td valign="top" align="center"><bold>0.714</bold></td>
<td valign="top" align="center"><bold>0.706</bold></td>
</tr>
<tr>
<td valign="top" align="left">W/classification</td>
<td valign="top" align="center">0.814</td>
<td valign="top" align="center">0.763</td>
<td valign="top" align="center">0.653</td>
<td valign="top" align="center">0.695</td>
<td valign="top" align="center">0.659</td>
</tr>
<tr>
<td valign="top" align="left">W/detection</td>
<td valign="top" align="center">0.812</td>
<td valign="top" align="center">0.740</td>
<td valign="top" align="center">0.644</td>
<td valign="top" align="center">0.682</td>
<td valign="top" align="center">0.640</td>
</tr>
<tr>
<td valign="top" align="left">W/segmentation</td>
<td valign="top" align="center"><bold>0.826</bold></td>
<td valign="top" align="center">0.758</td>
<td valign="top" align="center">0.676</td>
<td valign="top" align="center">0.688</td>
<td valign="top" align="center">0.684</td>
</tr>
<tr>
<td valign="top" align="left">W/unrecognization</td>
<td valign="top" align="center">0.811</td>
<td valign="top" align="center">0.749</td>
<td valign="top" align="center">0.651</td>
<td valign="top" align="center">0.691</td>
<td valign="top" align="center">0.666</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Best results are in bold</italic>.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Results of training the distortion network from scratch on 80% Smartphone Photography Attribute and Quality (SPAQ).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>PLCC</bold></th>
<th valign="top" align="center"><bold>SPAQ(20%)</bold></th>
<th valign="top" align="center"><bold>KonIQ-10k</bold></th>
<th valign="top" align="center"><bold>LIVEW</bold></th>
<th valign="top" align="center"><bold>CID</bold></th>
<th valign="top" align="center"><bold>BID</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">only distortion</td>
<td valign="top" align="center">0.878</td>
<td valign="top" align="center">0.568</td>
<td valign="top" align="center">0.605</td>
<td valign="top" align="center">0.665</td>
<td valign="top" align="center">0.598</td>
</tr>
<tr>
<td valign="top" align="left">w/recognition</td>
<td valign="top" align="center">0.883</td>
<td valign="top" align="center">0.591</td>
<td valign="top" align="center"><bold>0.628</bold></td>
<td valign="top" align="center"><bold>0.702</bold></td>
<td valign="top" align="center">0.628</td>
</tr>
<tr>
<td valign="top" align="left">w/classification</td>
<td valign="top" align="center"><bold>0.884</bold></td>
<td valign="top" align="center">0.585</td>
<td valign="top" align="center">0.625</td>
<td valign="top" align="center">0.695</td>
<td valign="top" align="center"><bold>0.631</bold></td>
</tr>
<tr>
<td valign="top" align="left">w/detection</td>
<td valign="top" align="center">0.882</td>
<td valign="top" align="center"><bold>0.598</bold></td>
<td valign="top" align="center">0.626</td>
<td valign="top" align="center">0.698</td>
<td valign="top" align="center">0.623</td>
</tr>
<tr>
<td valign="top" align="left">w/segmentation</td>
<td valign="top" align="center">0.883</td>
<td valign="top" align="center">0.592</td>
<td valign="top" align="center">0.627</td>
<td valign="top" align="center">0.697</td>
<td valign="top" align="center">0.629</td>
</tr>
<tr>
<td valign="top" align="left">w/unrecognization</td>
<td valign="top" align="center">0.881</td>
<td valign="top" align="center">0.585</td>
<td valign="top" align="center">0.621</td>
<td valign="top" align="center">0.686</td>
<td valign="top" align="center">0.615</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Best results are in bold</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>To explore how intelligibility affects quality assessment results intuitively, we utilize the method of Grad-CAM (Selvaraju et al., <xref ref-type="bibr" rid="B32">2017</xref>) to investigate which area of an image affects the prediction most. Examples are shown in <xref ref-type="fig" rid="F7">Figure 7</xref>, where red areas have more conspicuous influence to the prediction than blue areas. As shown in <xref ref-type="fig" rid="F7">Figure 7</xref>, the intelligibility features do play an important role in the quality assessment. The baseline model with distortion network only (<xref ref-type="fig" rid="F6">Figure 6B</xref>) cannot effectively locate salient objects which people may pay attention to. The intelligibility features (<xref ref-type="fig" rid="F6">Figure 6C</xref>) alone mainly focus on relatively local regions and cannot well utilize global information of images. In contrast, the proposed model (<xref ref-type="fig" rid="F6">Figure 6D</xref>) not only meticulously locate salient objects (important for intelligibility), but also pay more attention to wider areas, which catches global information. It is widely acknowledged that both global and local information are vital to IQA metrics (Fang et al., <xref ref-type="bibr" rid="B9">2018</xref>); hence, from this point of view, it is not hard to understand that by combining the intelligibility features, our model can achieve better performance.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Visualization results of Grad-CAM. <bold>(A)</bold> Original images; <bold>(B)</bold> heat-maps of the baseline model; <bold>(C)</bold> heat-maps of the intelligibility network; <bold>(D)</bold> heat-maps of proposed model with image recognition task.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnins-15-739138-g0007.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="conclusions" id="s5">
<title>5. Conclusions</title>
<p>In this paper, we first analyzed the relation between intelligibility and image quality. The results reveal that intelligibility is indicative of image quality. Therefore, we proposed a new framework, i.e., Intelligibility-Enriched-IQA, to combine intelligibility with conventional distortion measure. Feature selection strategy was proposed to select the most important intelligibility features, which alleviates negative transfer and avoids damaging highly generalizable features. Extensive experimental results show the effectiveness of proposed method, and our model achieves state-of-the-art performance in terms of the generalization ability. These results demonstrate that introducing intelligibility is a promising way in building highly generalizable IQA metrics.</p>
</sec>
<sec sec-type="data-availability" id="s6">
<title>Data Availability Statement</title>
<p>The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found in the article/supplementary material.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>TS and LL contributed to conception and design of the study. TS performed the experiment and wrote the first draft of the manuscript. LL, HZ, and JQ wrote sections of the manuscript. All authors contributed to manuscript revision, read, and approved the submitted version.</p>
</sec>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>This work was supported in part by the National Natural Science Foundation of China under Grants 62171340, 61771473, 61991451, and 61379143, the Key Project of Shaanxi Provincial Department of Education (Collaborative Innovation Center) under Grant 20JY024, the Fundamental Research Funds for the Central Universities under Grant JBF211902, the Science and Technology Plan of Xi&#x00027;an under Grant 20191122015KYPT011JC013, the Natural Science Foundation of Jiangsu Province under Grants BK20181354 and BK20200649, and the Six Talent Peaks High-level Talents in Jiangsu Province under Grant XYDXX-063.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec> 
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Abdou</surname> <given-names>I. E.</given-names></name> <name><surname>Dusaussoy</surname> <given-names>N. J.</given-names></name></person-group> (<year>1986</year>). <article-title>&#x0201C;Survey of image quality measurements,&#x0201D;</article-title> in <source>Proceedings of 1986 ACM Fall Joint Computer Conference, ACM &#x00027;86</source> (<publisher-loc>Washington, DC</publisher-loc>: <publisher-name>IEEE Computer Society Press</publisher-name>), <fpage>71</fpage>&#x02013;<lpage>78</lpage>.</citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bianco</surname> <given-names>S.</given-names></name> <name><surname>Celona</surname> <given-names>L.</given-names></name> <name><surname>Napoletano</surname> <given-names>P.</given-names></name> <name><surname>Schettini</surname> <given-names>R.</given-names></name></person-group> (<year>2018</year>). <article-title>On the use of deep learning for blind image quality assessment</article-title>. <source>Signal Image Video Process</source>. <volume>12</volume>, <fpage>355</fpage>&#x02013;<lpage>362</lpage>. <pub-id pub-id-type="doi">10.1007/s11760-017-1166-8</pub-id><pub-id pub-id-type="pmid">31995493</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bosse</surname> <given-names>S.</given-names></name> <name><surname>Maniry</surname> <given-names>D.</given-names></name> <name><surname>M&#x000FC;ller</surname> <given-names>K.-R.</given-names></name> <name><surname>Wiegand</surname> <given-names>T.</given-names></name> <name><surname>Samek</surname> <given-names>W.</given-names></name></person-group> (<year>2017</year>). <article-title>Deep neural networks for no-reference and full-reference image quality assessment</article-title>. <source>IEEE Trans. Image Process</source>. <volume>27</volume>, <fpage>206</fpage>&#x02013;<lpage>219</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2017.2760518</pub-id><pub-id pub-id-type="pmid">29028191</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cao</surname> <given-names>Z.</given-names></name> <name><surname>Long</surname> <given-names>M.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Jordan</surname> <given-names>M. I.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Partial transfer learning with selective adversarial networks,&#x0201D;</article-title> in <source>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>), <fpage>2724</fpage>&#x02013;<lpage>2732</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00288</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>L. C.</given-names></name> <name><surname>Zhu</surname> <given-names>Y.</given-names></name> <name><surname>Papandreou</surname> <given-names>G.</given-names></name> <name><surname>Schroff</surname> <given-names>F.</given-names></name> <name><surname>Adam</surname> <given-names>H.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Encoder-decoder with atrous separable convolution for semantic image segmentation,&#x0201D;</article-title> in <source>2018 Proceedings of European Conference on Computer Vision (ECCV)</source> (<publisher-loc>Munich</publisher-loc>), <fpage>833</fpage>&#x02013;<lpage>851</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-01234-2_49</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chiu</surname> <given-names>T.</given-names></name> <name><surname>Zhao</surname> <given-names>Y.</given-names></name> <name><surname>Gurari</surname> <given-names>D.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Assessing image quality issues for real-world problems,&#x0201D;</article-title> in <source>2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>, <fpage>3643</fpage>&#x02013;<lpage>3653</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00370</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ciancio</surname> <given-names>A.</given-names></name> <name><surname>Targino da Costa</surname> <given-names>A. L. N. T.</given-names></name> <name><surname>da Silva</surname> <given-names>E. A. B.</given-names></name> <name><surname>Said</surname> <given-names>A.</given-names></name> <name><surname>Samadani</surname> <given-names>R.</given-names></name> <name><surname>Obrador</surname> <given-names>P.</given-names></name></person-group> (<year>2011</year>). <article-title>No-reference blur assessment of digital pictures based on multifeature classifiers</article-title>. <source>IEEE Trans. Image Process</source>. <volume>20</volume>, <fpage>64</fpage>&#x02013;<lpage>75</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2010.2053549</pub-id><pub-id pub-id-type="pmid">21172744</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dimopoulos</surname> <given-names>Y.</given-names></name> <name><surname>Bourret</surname> <given-names>P.</given-names></name> <name><surname>Lek</surname> <given-names>S.</given-names></name></person-group> (<year>1995</year>). <article-title>Use of some sensitivity criteria for choosing networks with good generalization ability</article-title>. <source>Neural Process. Lett</source>. <volume>2</volume>, <fpage>1</fpage>&#x02013;<lpage>4</lpage>. <pub-id pub-id-type="doi">10.1007/BF02309007</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fang</surname> <given-names>Y.</given-names></name> <name><surname>Yan</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name> <name><surname>Lin</surname> <given-names>W.</given-names></name></person-group> (<year>2018</year>). <article-title>No reference quality assessment for screen content images with both local and global feature representation</article-title>. <source>IEEE Trans. Image Process</source>. <volume>27</volume>, <fpage>1600</fpage>&#x02013;<lpage>1610</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2017.2781307</pub-id><pub-id pub-id-type="pmid">29324414</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fang</surname> <given-names>Y.</given-names></name> <name><surname>Zhu</surname> <given-names>H.</given-names></name> <name><surname>Zeng</surname> <given-names>Y.</given-names></name> <name><surname>Ma</surname> <given-names>K.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Perceptual quality assessment of smartphone photography,&#x0201D;</article-title> in <source>2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>, <fpage>3674</fpage>&#x02013;<lpage>3683</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00373</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Garson</surname> <given-names>G. D.</given-names></name></person-group> (<year>1991</year>). <article-title>Interpreting neural-network connection weights</article-title>. <source>AI Expert</source> <volume>6</volume>, <fpage>46</fpage>&#x02013;<lpage>51</lpage>.</citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ghadiyaram</surname> <given-names>D.</given-names></name> <name><surname>Bovik</surname> <given-names>A. C.</given-names></name></person-group> (<year>2016</year>). <article-title>Massive online crowdsourced study of subjective and objective picture quality</article-title>. <source>IEEE Trans. Image Process</source>. <volume>25</volume>, <fpage>372</fpage>&#x02013;<lpage>387</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2015.2500021</pub-id><pub-id pub-id-type="pmid">26571530</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gu</surname> <given-names>K.</given-names></name> <name><surname>Zhai</surname> <given-names>G.</given-names></name> <name><surname>Liu</surname> <given-names>M.</given-names></name> <name><surname>Yang</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name> <name><surname>Sun</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>&#x0201C;FISBLIM: a five-step blind metric for quality assessment of multiply distorted images,&#x0201D;</article-title> in <source>SiPS 2013 Proceedings</source> (<publisher-loc>Taipei</publisher-loc>), <fpage>241</fpage>&#x02013;<lpage>246</lpage>. <pub-id pub-id-type="doi">10.1109/SiPS.2013.6674512</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gu</surname> <given-names>K.</given-names></name> <name><surname>Zhai</surname> <given-names>G.</given-names></name> <name><surname>Yang</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name></person-group> (<year>2014</year>). <article-title>Using free energy principle for blind image quality assessment</article-title>. <source>IEEE Trans. Multimedia</source> <volume>17</volume>, <fpage>50</fpage>&#x02013;<lpage>63</lpage>. <pub-id pub-id-type="doi">10.1109/TMM.2014.2373812</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Deep residual learning for image recognition,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Las Vegas, NV</publisher-loc>), <fpage>770</fpage>&#x02013;<lpage>778</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id><pub-id pub-id-type="pmid">32166560</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hosu</surname> <given-names>V.</given-names></name> <name><surname>Lin</surname> <given-names>H.</given-names></name> <name><surname>Sziranyi</surname> <given-names>T.</given-names></name> <name><surname>Saupe</surname> <given-names>D.</given-names></name></person-group> (<year>2020</year>). <article-title>KonIQ-10k: an ecologically valid database for deep learning of blind image quality assessment</article-title>. <source>IEEE Trans. Image Process</source>. <volume>29</volume>, <fpage>4041</fpage>&#x02013;<lpage>4056</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2020.2967829</pub-id><pub-id pub-id-type="pmid">31995493</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hu</surname> <given-names>H.</given-names></name> <name><surname>Peng</surname> <given-names>R.</given-names></name> <name><surname>Tai</surname> <given-names>Y.-W.</given-names></name> <name><surname>Tang</surname> <given-names>C.-K.</given-names></name></person-group> (<year>2016</year>). <article-title>Network trimming: a data-driven neuron pruning approach towards efficient deep architectures</article-title>. <source>arXiv [Preprint]</source> arXiv: 1607.03250.</citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kang</surname> <given-names>L.</given-names></name> <name><surname>Ye</surname> <given-names>P.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Doermann</surname> <given-names>D.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;Convolutional neural networks for no-reference image quality assessment,&#x0201D;</article-title> in <source>2014 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Columbus, OH</publisher-loc>), <fpage>1733</fpage>&#x02013;<lpage>1740</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2014.224</pub-id><pub-id pub-id-type="pmid">33546412</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kang</surname> <given-names>L.</given-names></name> <name><surname>Ye</surname> <given-names>P.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Doermann</surname> <given-names>D.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Simultaneous estimation of image quality and distortion via multi-task convolutional neural networks,&#x0201D;</article-title> in <source>2015 IEEE International Conference on Image Processing (ICIP)</source> (<publisher-loc>Quebec City, QC</publisher-loc>), <fpage>2791</fpage>&#x02013;<lpage>2795</lpage>. <pub-id pub-id-type="doi">10.1109/ICIP.2015.7351311</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>J.</given-names></name> <name><surname>Lee</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>Fully deep blind image quality predictor</article-title>. <source>IEEE J. Select. Top. Signal Process</source>. <volume>11</volume>, <fpage>206</fpage>&#x02013;<lpage>220</lpage>. <pub-id pub-id-type="doi">10.1109/JSTSP.2016.2639328</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kottayil</surname> <given-names>N. K.</given-names></name> <name><surname>Cheng</surname> <given-names>I.</given-names></name> <name><surname>Dufaux</surname> <given-names>F.</given-names></name> <name><surname>Basu</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>A color intensity invariant low-level feature optimization framework for image quality assessment</article-title>. <source>Signal Image Video Process</source>. <volume>10</volume>, <fpage>1169</fpage>&#x02013;<lpage>1176</lpage>. <pub-id pub-id-type="doi">10.1007/s11760-016-0873-x</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Q.</given-names></name> <name><surname>Lin</surname> <given-names>W.</given-names></name> <name><surname>Fang</surname> <given-names>Y.</given-names></name></person-group> (<year>2016</year>). <article-title>No-reference quality assessment for multiply-distorted images in gradient domain</article-title>. <source>IEEE Signal Process. Lett</source>. <volume>23</volume>, <fpage>541</fpage>&#x02013;<lpage>545</lpage>. <pub-id pub-id-type="doi">10.1109/LSP.2016.2537321</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>T.-Y.</given-names></name> <name><surname>Dollar</surname> <given-names>P.</given-names></name> <name><surname>Girshick</surname> <given-names>R.</given-names></name> <name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Hariharan</surname> <given-names>B.</given-names></name> <name><surname>Belongie</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Feature pyramid networks for object detection,&#x0201D;</article-title> in <source>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Honolulu, HI</publisher-loc>), <fpage>936</fpage>&#x02013;<lpage>944</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.106</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>T.-Y.</given-names></name> <name><surname>Maire</surname> <given-names>M.</given-names></name> <name><surname>Belongie</surname> <given-names>S.</given-names></name> <name><surname>Hays</surname> <given-names>J.</given-names></name> <name><surname>Perona</surname> <given-names>P.</given-names></name> <name><surname>Ramanan</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>&#x0201C;Microsoft COCO: common objects in context,&#x0201D;</article-title> in <source>2014 Proceedings of European Conference on Computer Vision (ECCV)</source> (<publisher-loc>Zurich</publisher-loc>), <fpage>740</fpage>&#x02013;<lpage>755</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-10602-1_48</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ma</surname> <given-names>K.</given-names></name> <name><surname>Liu</surname> <given-names>W.</given-names></name> <name><surname>Zhang</surname> <given-names>K.</given-names></name> <name><surname>Duanmu</surname> <given-names>Z.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Zuo</surname> <given-names>W.</given-names></name></person-group> (<year>2017</year>). <article-title>End-to-end blind image quality assessment using deep neural networks</article-title>. <source>IEEE Trans. Image Process</source>. <volume>27</volume>, <fpage>1202</fpage>&#x02013;<lpage>1213</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2017.2774045</pub-id><pub-id pub-id-type="pmid">29220321</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mittal</surname> <given-names>A.</given-names></name> <name><surname>Moorthy</surname> <given-names>A. K.</given-names></name> <name><surname>Bovik</surname> <given-names>A. C.</given-names></name></person-group> (<year>2012</year>). <article-title>No-reference image quality assessment in the spatial domain</article-title>. <source>IEEE Trans. Image Process</source>. <volume>21</volume>, <fpage>4695</fpage>&#x02013;<lpage>4708</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2012.2214050</pub-id><pub-id pub-id-type="pmid">22910118</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moorthy</surname> <given-names>A. K.</given-names></name> <name><surname>Bovik</surname> <given-names>A. C.</given-names></name></person-group> (<year>2010</year>). <article-title>A two-step framework for constructing blind image quality indices</article-title>. <source>IEEE Signal Process. Lett</source>. <volume>17</volume>, <fpage>513</fpage>&#x02013;<lpage>516</lpage>. <pub-id pub-id-type="doi">10.1109/LSP.2010.2043888</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pan</surname> <given-names>S. J.</given-names></name> <name><surname>Yang</surname> <given-names>Q.</given-names></name></person-group> (<year>2010</year>). <article-title>A survey on transfer learning</article-title>. <source>IEEE Trans. Knowledge Data Eng</source>. <volume>22</volume>, <fpage>1345</fpage>&#x02013;<lpage>1359</lpage>. <pub-id pub-id-type="doi">10.1109/TKDE.2009.191</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Girshick</surname> <given-names>R.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>Faster R-CNN: towards real-time object detection with region proposal networks</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. <volume>39</volume>, <fpage>1137</fpage>&#x02013;<lpage>1149</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2016.2577031</pub-id><pub-id pub-id-type="pmid">27295650</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Russakovsky</surname> <given-names>O.</given-names></name> <name><surname>Deng</surname> <given-names>J.</given-names></name> <name><surname>Su</surname> <given-names>H.</given-names></name> <name><surname>Krause</surname> <given-names>J.</given-names></name> <name><surname>Satheesh</surname> <given-names>S.</given-names></name> <name><surname>Ma</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>ImageNet large scale visual recognition challenge</article-title>. <source>Int. J. Comput. Vision</source> <volume>115</volume>, <fpage>211</fpage>&#x02013;<lpage>252</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-015-0816-y</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Saad</surname> <given-names>M. A.</given-names></name> <name><surname>Bovik</surname> <given-names>A. C.</given-names></name> <name><surname>Charrier</surname> <given-names>C.</given-names></name></person-group> (<year>2012</year>). <article-title>Blind image quality assessment: a natural scene statistics approach in the DCT domain</article-title>. <source>IEEE Trans. Image Process</source>. <volume>21</volume>, <fpage>3339</fpage>&#x02013;<lpage>3352</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2012.2191563</pub-id><pub-id pub-id-type="pmid">22453635</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Selvaraju</surname> <given-names>R. R.</given-names></name> <name><surname>Cogswell</surname> <given-names>M.</given-names></name> <name><surname>Das</surname> <given-names>A.</given-names></name> <name><surname>Vedantam</surname> <given-names>R.</given-names></name> <name><surname>Parikh</surname> <given-names>D.</given-names></name> <name><surname>Batra</surname> <given-names>D.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Grad-CAM: Visual explanations from deep networks via gradient-based localization,&#x0201D;</article-title> in <source>2017 IEEE International Conference on Computer Vision (ICCV)</source> (<publisher-loc>Venice</publisher-loc>), <fpage>618</fpage>&#x02013;<lpage>626</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2017.74</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Simonyan</surname> <given-names>K.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Very deep convolutional networks for large-scale image recognition,&#x0201D;</article-title> in <source>2015 International Conference on Learning Representations (ICLR)</source> (<publisher-loc>San Diego, CA</publisher-loc>).</citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Su</surname> <given-names>S.</given-names></name> <name><surname>Yan</surname> <given-names>Q.</given-names></name> <name><surname>Zhu</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Ge</surname> <given-names>X.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>&#x0201C;Blindly assess image quality in the wild guided by a self-adaptive hyper network,&#x0201D;</article-title> in <source>2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>, <fpage>3664</fpage>&#x02013;<lpage>3673</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00372</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tan</surname> <given-names>M.</given-names></name> <name><surname>Le</surname> <given-names>Q. V.</given-names></name></person-group> (<year>2019</year>). <article-title>EfficientNet: rethinking model scaling for convolutional neural networks</article-title>. <source>arXiv [preprint]</source> arXiv: 1905.11946.</citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thomee</surname> <given-names>B.</given-names></name> <name><surname>Shamma</surname> <given-names>D. A.</given-names></name> <name><surname>Friedland</surname> <given-names>G.</given-names></name> <name><surname>Elizalde</surname> <given-names>B.</given-names></name> <name><surname>Ni</surname> <given-names>K.</given-names></name> <name><surname>Poland</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>YFCC100M: the new data in multimedia research</article-title>. <source>Commun. ACM</source> <volume>59</volume>, <fpage>64</fpage>&#x02013;<lpage>73</lpage>. <pub-id pub-id-type="doi">10.1145/2812802</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Varga</surname> <given-names>D.</given-names></name> <name><surname>Saupe</surname> <given-names>D.</given-names></name> <name><surname>Sziranyi</surname> <given-names>T.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;DeepRN: a content preserving deep architecture for blind image quality assessment,&#x0201D;</article-title> in <source>2018 IEEE International Conference on Multimedia and Expo (ICME)</source> (<publisher-loc>San Diego, CA</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1109/ICME.2018.8486528</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Virtanen</surname> <given-names>T.</given-names></name> <name><surname>Nuutinen</surname> <given-names>M.</given-names></name> <name><surname>Vaahteranoksa</surname> <given-names>M.</given-names></name> <name><surname>Oittinen</surname> <given-names>P.</given-names></name> <name><surname>Hakkinen</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>CID2013: a database for evaluating no-reference image quality assessment algorithms</article-title>. <source>IEEE Trans. Image Process</source>. <volume>24</volume>, <fpage>390</fpage>&#x02013;<lpage>402</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2014.2378061</pub-id><pub-id pub-id-type="pmid">25494511</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Guo</surname> <given-names>S.</given-names></name> <name><surname>Huang</surname> <given-names>W.</given-names></name> <name><surname>Qiao</surname> <given-names>Y.</given-names></name></person-group> (<year>2015</year>). <article-title>Places205-vggnet models for scene recognition</article-title>. <source>arXiv [Preprint]</source> arXiv: 1508.01667.</citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>J.</given-names></name> <name><surname>Ye</surname> <given-names>P.</given-names></name> <name><surname>Li</surname> <given-names>Q.</given-names></name> <name><surname>Du</surname> <given-names>H.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Doermann</surname> <given-names>D.</given-names></name></person-group> (<year>2016</year>). <article-title>Blind image quality assessment based on high order statistics aggregation</article-title>. <source>IEEE Trans. Image Process</source>. <volume>25</volume>, <fpage>4444</fpage>&#x02013;<lpage>4457</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2016.2585880</pub-id><pub-id pub-id-type="pmid">27362977</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yan</surname> <given-names>B.</given-names></name> <name><surname>Bare</surname> <given-names>B.</given-names></name> <name><surname>Tan</surname> <given-names>W.</given-names></name></person-group> (<year>2019</year>). <article-title>Naturalness-aware deep no-reference image quality assessment</article-title>. <source>IEEE Trans. Multimedia</source> <volume>21</volume>, <fpage>2603</fpage>&#x02013;<lpage>2615</lpage>. <pub-id pub-id-type="doi">10.1109/TMM.2019.2904879</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ye</surname> <given-names>P.</given-names></name> <name><surname>Kumar</surname> <given-names>J.</given-names></name> <name><surname>Kang</surname> <given-names>L.</given-names></name> <name><surname>Doermann</surname> <given-names>D.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Unsupervised feature learning framework for no-reference image quality assessment,&#x0201D;</article-title> in <source>2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Providence, RI</publisher-loc>), <fpage>1098</fpage>&#x02013;<lpage>1105</lpage>.</citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhai</surname> <given-names>G.</given-names></name> <name><surname>Zhu</surname> <given-names>Y.</given-names></name> <name><surname>Min</surname> <given-names>X.</given-names></name></person-group> (<year>2020</year>). <article-title>Comparative perceptual assessment of visual signals using free energy features</article-title>. <source>IEEE Trans. Multimedia</source>. <pub-id pub-id-type="doi">10.1109/TMM.2020.3029891</pub-id>. [Epub ahead of print].</citation>
</ref>
<ref id="B44">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Ding</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Ogunbona</surname> <given-names>P.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Importance weighted adversarial nets for partial domain adaptation,&#x0201D;</article-title> in <source>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>), <fpage>8156</fpage>&#x02013;<lpage>8164</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00851</pub-id></citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Min</surname> <given-names>X.</given-names></name> <name><surname>Zhu</surname> <given-names>Y.</given-names></name> <name><surname>Zhai</surname> <given-names>G.</given-names></name> <name><surname>Zhou</surname> <given-names>J.</given-names></name> <name><surname>Yang</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Hazdesnet: an end-to-end network for haze density prediction</article-title>. <source>IEEE Trans. Intell. Transport. Syst</source>. <pub-id pub-id-type="doi">10.1109/TITS.2020.3030673</pub-id>. [Epub ahead of print].</citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Q.</given-names></name> <name><surname>Wu</surname> <given-names>Y. N.</given-names></name> <name><surname>Zhu</surname> <given-names>S.-C.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Interpretable convolutional neural networks,&#x0201D;</article-title> in <source>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>, <fpage>8827</fpage>&#x02013;<lpage>8836</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00920</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>W.</given-names></name> <name><surname>Ma</surname> <given-names>K.</given-names></name> <name><surname>Yan</surname> <given-names>J.</given-names></name> <name><surname>Deng</surname> <given-names>D.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name></person-group> (<year>2020</year>). <article-title>Blind image quality assessment using a deep bilinear convolutional neural network</article-title>. <source>IEEE Trans. Circ. Syst. Video Technol</source>. <volume>30</volume>, <fpage>36</fpage>&#x02013;<lpage>47</lpage>. <pub-id pub-id-type="doi">10.1109/TCSVT.2018.2886771</pub-id></citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>B.</given-names></name> <name><surname>Bau</surname> <given-names>D.</given-names></name> <name><surname>Oliva</surname> <given-names>A.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name></person-group> (<year>2019</year>). <article-title>Interpreting deep visual representations via network dissection</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. <volume>41</volume>, <fpage>2131</fpage>&#x02013;<lpage>2145</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2018.2858759</pub-id><pub-id pub-id-type="pmid">30040625</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>B.</given-names></name> <name><surname>Khosla</surname> <given-names>A.</given-names></name> <name><surname>Lapedriza</surname> <given-names>A.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name> <name><surname>Oliva</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). <article-title>Places: an image database for deep scene understanding</article-title>. <source>J. Vis</source>. <volume>17</volume>:<fpage>296</fpage>. <pub-id pub-id-type="doi">10.1167/17.10.296</pub-id></citation>
</ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Wu</surname> <given-names>J.</given-names></name> <name><surname>Dong</surname> <given-names>W.</given-names></name> <name><surname>Shi</surname> <given-names>G.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;MetaIQA: deep meta-learning for no-reference image quality assessment,&#x0201D;</article-title> in <source>2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>, <fpage>14131</fpage>&#x02013;<lpage>14140</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.01415</pub-id></citation>
</ref>
</ref-list> 
</back>
</article>