<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurorobot.</journal-id>
<journal-title>Frontiers in Neurorobotics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurorobot.</abbrev-journal-title>
<issn pub-type="epub">1662-5218</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnbot.2022.841426</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Color Constancy <italic>via</italic> Multi-Scale Region-Weighed Network Guided by Semantics</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Wang</surname> <given-names>Fei</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1609099/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Wang</surname> <given-names>Wei</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Wu</surname> <given-names>Dan</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Gao</surname> <given-names>Guowang</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>School of Electronic Engineering, Xi&#x00027;an Shiyou University</institution>, <addr-line>Xi&#x00027;an</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>State Key Laboratory of Advanced Design and Manufacturing for Vehicle Body, Hunan University</institution>, <addr-line>Changsha</addr-line>, <country>China</country></aff>
<aff id="aff3"><sup>3</sup><institution>School of Telecommunications Engineering, Xidian University</institution>, <addr-line>Xi&#x00027;an</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Kok-Lim Alvin Yau, Tunku Abdul Rahman University, Malaysia</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Masataka Sawayama, Institut National de Recherche en Informatique et en Automatique (INRIA), France; Shaobing Gao, University of Electronic Science and Technology of China, China; Fengshun Lu, Beijing Aerohydrodynamic Frontier Research Center, China; Kannimuthu Subramanian, Karpagam Academy of Higher Education, India</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Fei Wang <email>200102&#x00040;xsyu.edu.cn</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>08</day>
<month>04</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>16</volume>
<elocation-id>841426</elocation-id>
<history>
<date date-type="received">
<day>22</day>
<month>12</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>08</day>
<month>03</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Wang, Wang, Wu and Gao.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Wang, Wang, Wu and Gao</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>In obtaining color constancy, estimating the illumination of a scene is the most important task. However, due to unknown light sources and the influence of the external imaging environment, the estimated illumination is prone to color ambiguity. In this article, a learning-based multi-scale region-weighed network guided by semantic features is proposed to estimate the illuminated color of the light source in a scene. Cued by the human brain&#x00027;s processing of color constancy, we use image semantics and scale information to guide the process of illumination estimation. First, we put the image and its semantics into the network, and then obtain the region weights of the image at different scales. After that, through a special weight-pooling layer (WPL), the illumination on each scale is estimated. The final illumination is calculated by weighting each scale. The results of extensive experiments on Color Checker and NUS 8-Camera datasets show that the proposed approach is superior to the current state-of-the-art methods in both efficiency and effectiveness.</p></abstract>
<kwd-group>
<kwd>color constancy</kwd>
<kwd>multi-scale</kwd>
<kwd>weight pooling layer</kwd>
<kwd>semantic</kwd>
<kwd>network</kwd>
</kwd-group>
<counts>
<fig-count count="9"/>
<table-count count="6"/>
<equation-count count="11"/>
<ref-count count="61"/>
<page-count count="15"/>
<word-count count="8783"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>The observed color of an object in an image (representing the observed values in RGB space) depends on the intrinsic color and light-source color. It is quite easy to distinguish the reflectance from the light-source color for human beings while endowing a computer with the same ability is difficult (Gilchrist, <xref ref-type="bibr" rid="B29">2006</xref>). For example, given a red object, how can one discern if it is a white object under red light or a red object under a white light? To assist a computer in solving this problem, it is necessary to separate the color of the light source, namely, the color constancy.<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref> The goal of computational color constancy is to preserve the perceptive colors of objects under different lighting conditions by removing the effect of color casts caused by the scene&#x00027;s illumination.</p>
<p>Color constancy is a fundamental research topic in the image-processing and computer-vision fields, and it has many applications in photographic technology, object recognition, object detection, image segmentation, and other version systems. Color casts caused by incorrectly applied computational color constancy can negatively impact the performance of image segmentation and classification (Afifi and Brown, <xref ref-type="bibr" rid="B3">2019</xref>; Xue et al., <xref ref-type="bibr" rid="B57">2021</xref>), thus, there is a rich body of work on this topic. Generally, methods for obtaining color constancy with image data are divided into two main categories: low-level-feature-based methods (Buchsbaum, <xref ref-type="bibr" rid="B14">1980</xref>; Brainard and Wandell, <xref ref-type="bibr" rid="B13">1986</xref>; Lee, <xref ref-type="bibr" rid="B42">1986</xref>; Wandell and Tominaga, <xref ref-type="bibr" rid="B54">1989</xref>; Nieves et al., <xref ref-type="bibr" rid="B45">2000</xref>; Krasilnikov et al., <xref ref-type="bibr" rid="B38">2002</xref>; Weijer et al., <xref ref-type="bibr" rid="B55">2007</xref>; Gehler et al., <xref ref-type="bibr" rid="B26">2008</xref>; Tan et al., <xref ref-type="bibr" rid="B51">2008</xref>; Toro, <xref ref-type="bibr" rid="B52">2008</xref>; Gijsenij et al., <xref ref-type="bibr" rid="B28">2011</xref>; Finlayson, <xref ref-type="bibr" rid="B17">2013</xref>; Gao et al., <xref ref-type="bibr" rid="B23">2013</xref>; Barron, <xref ref-type="bibr" rid="B9">2015</xref>; Bianco et al., <xref ref-type="bibr" rid="B11">2015</xref>, <xref ref-type="bibr" rid="B12">2017</xref>; Cheng et al., <xref ref-type="bibr" rid="B16">2015</xref>; Shi et al., <xref ref-type="bibr" rid="B48">2016</xref>; Xiao et al., <xref ref-type="bibr" rid="B56">2020</xref>; Yu et al., <xref ref-type="bibr" rid="B58">2020</xref>) and semantic-feature-based methods (Schroeder and Moser, <xref ref-type="bibr" rid="B47">2001</xref>; Spitzer and Semo, <xref ref-type="bibr" rid="B50">2002</xref>; Van De Weijer et al., <xref ref-type="bibr" rid="B53">2007</xref>; Bianco et al., <xref ref-type="bibr" rid="B10">2008</xref>; Lau, <xref ref-type="bibr" rid="B41">2008</xref>; Li et al., <xref ref-type="bibr" rid="B43">2008</xref>; Gao et al., <xref ref-type="bibr" rid="B25">2015</xref>; Afifi, <xref ref-type="bibr" rid="B2">2018</xref>).</p>
<p><italic>Low-level-features-based methods</italic> pay attention to the law of the color of the image itself, and they do not consider the image-content information. These methods consider the relationship between color and achromatic color statistics (Weijer et al., <xref ref-type="bibr" rid="B55">2007</xref>; Gehler et al., <xref ref-type="bibr" rid="B26">2008</xref>), inspired by the human visual system (Nieves et al., <xref ref-type="bibr" rid="B45">2000</xref>; Krasilnikov et al., <xref ref-type="bibr" rid="B38">2002</xref>; Gao et al., <xref ref-type="bibr" rid="B23">2013</xref>), spatial derivatives, and frequency information of scene illuminations on the image (Nayar et al., <xref ref-type="bibr" rid="B44">2007</xref>; Joze and Drew, <xref ref-type="bibr" rid="B35">2014</xref>), extract hand-crafted features from training data (Buchsbaum, <xref ref-type="bibr" rid="B14">1980</xref>; Brainard and Wandell, <xref ref-type="bibr" rid="B13">1986</xref>; Finlayson, <xref ref-type="bibr" rid="B17">2013</xref>; Cheng et al., <xref ref-type="bibr" rid="B16">2015</xref>), and learn features automatically by a convolutional neural network (CNN) from samples (Barron, <xref ref-type="bibr" rid="B9">2015</xref>; Bianco et al., <xref ref-type="bibr" rid="B11">2015</xref>, <xref ref-type="bibr" rid="B12">2017</xref>; Shi et al., <xref ref-type="bibr" rid="B48">2016</xref>; Xiao et al., <xref ref-type="bibr" rid="B56">2020</xref>; Yu et al., <xref ref-type="bibr" rid="B58">2020</xref>). Although these methods have achieved good results, especially the CNN-based methods, various methods are used to make the illumination estimation as accurate as possible, but in some complex situations, due to inflexibility, they cannot well solve the color ambiguity.</p>
<p><italic>Semantic-feature-based methods</italic> are more in line with human vision. When observing a scene, human beings have a certain psychological memory of the color of the object itself in the scene. Therefore, the content of the scene can play a certain guiding role in color constancy. Because previous attempts at semantic information extraction have not been accurate, there is relatively little research on this type of algorithm. Van De Weijer et al. (<xref ref-type="bibr" rid="B53">2007</xref>) proposed a color-constancy algorithm based on advanced visual information. The algorithm models the image into many semantic categories, such as sky, grassland, road, pedestrian, and vehicle, and it calculates multiple possible illumination values from these semantic categories. Each illumination is used to correct the image and calculate the semantic combination with the greatest probability. At this time, the illumination is the optimal scene illumination. Schroeder and Moser (<xref ref-type="bibr" rid="B47">2001</xref>) divided images into different categories, and then they learned different features for each type. Bianco et al. (<xref ref-type="bibr" rid="B10">2008</xref>) proposed an indoor and outdoor adaptive illumination estimation algorithm that uses a classification algorithm to divide the image into indoor and outdoor scenes, and then they estimated the illumination according to the parameters learned by training data. Afifi (<xref ref-type="bibr" rid="B2">2018</xref>) exploited the semantic information together with the color and spatial information of the input image, and they trained a CNN to estimate the illuminant color and gamma correction parameters. This is one of the most effective methods of this type.</p>
<p>However, these two methods may not find the optimal solution in some complex situations due to inflexibility. To summarize, several open problems remain unsolved in these approaches, which can be generally concluded to have two aspects.</p>
<list list-type="bullet">
<list-item><p><italic>Color ambiguity with only low-level features:</italic> Many of these methods (Buchsbaum, <xref ref-type="bibr" rid="B14">1980</xref>; Brainard and Wandell, <xref ref-type="bibr" rid="B13">1986</xref>; Lee, <xref ref-type="bibr" rid="B42">1986</xref>; Wandell and Tominaga, <xref ref-type="bibr" rid="B54">1989</xref>; Nieves et al., <xref ref-type="bibr" rid="B45">2000</xref>; Krasilnikov et al., <xref ref-type="bibr" rid="B38">2002</xref>; Weijer et al., <xref ref-type="bibr" rid="B55">2007</xref>; Gehler et al., <xref ref-type="bibr" rid="B26">2008</xref>; Tan et al., <xref ref-type="bibr" rid="B51">2008</xref>; Toro, <xref ref-type="bibr" rid="B52">2008</xref>; Gijsenij et al., <xref ref-type="bibr" rid="B28">2011</xref>; Finlayson, <xref ref-type="bibr" rid="B17">2013</xref>; Gao et al., <xref ref-type="bibr" rid="B23">2013</xref>; Cheng et al., <xref ref-type="bibr" rid="B16">2015</xref>) only focus on the color of the image itself, for images with large color deviation, it is difficult to accurately estimate the illumination. The development of CNNs has facilitated a qualitative leap in illumination estimation, but many CNN-based methods (Barron, <xref ref-type="bibr" rid="B9">2015</xref>; Shi et al., <xref ref-type="bibr" rid="B48">2016</xref>; Bianco et al., <xref ref-type="bibr" rid="B12">2017</xref>) are patches-based, which take the small sampled image patches as input and learn the corresponding local estimations subsequently pooled into a global result. The small patches contain less contextual information, which commonly leads to ambiguity in local estimation. When inferring the illumination color in a patch, it is often the case that the patch contains little or no semantic context benefiting its reflectance or illumination estimation. To solve this problem, Hu et al. (<xref ref-type="bibr" rid="B32">2017</xref>) used a global image as input, and they designed a confidence weight layer to learn the weight of each patch. Afifi and Brown (<xref ref-type="bibr" rid="B4">2020</xref>) proposed an end-to-end approach to learn the correct white balance, which consists of a single encoder and multiple decoders, mapping an input image to two additional white-balance settings corresponding to indoor and outdoor illuminations. Both methods achieved great success. However, due to the lack of attention to semantic information, large errors in illumination estimation exist in some scenes, which has also been found in our experiments.</p></list-item>
<list-item><p><italic>Inaccurate illumination with only semantics:</italic> Owing to the low accuracy of semantic segmentation, the early semantic-based color-constancy algorithm has great limitations. CNNs have greatly improved the accuracy of semantic segmentation. Afifi (<xref ref-type="bibr" rid="B2">2018</xref>) exploited the semantic information together with the color and spatial information of the input image, and they trained a CNN to estimate the illuminant color and gamma correction parameters. However, there is a possibility of error in segmentation, and incorrect segmentation will lead to errors in illumination estimation. In our experiments, we also verified that incorrect segmentation can lead to incorrect illumination estimates.</p></list-item>
</list>
<p>To address the aforementioned open problems, we did some experiments. In one experiment, we conducted, yellow banana and a red apple were placed into a scene, illuminated with different colors of light, and then different observers were allowed to view the results. It was found that the observers can correctly distinguish the color of the fruit because humans generally think that bananas are yellow and apples are red. It can be seen that objects with inherent colors can guide observers to estimate the lighting of the scene; i.e., the objects in the scene have a great effect on the color constancy of human vision. In addition, we conducted an experiment in which some objects were placed in the scene to allow the observer to observe them from different distances. It was found that the colors of certain areas in the scene observed at different distances were biased (this deviation is relatively small), but the bias did not have much impact on the overall color. In addition, in the present study, traditional algorithms were used to estimate the illumination of the image at different scales. It was found that the estimated illumination at different scales exhibits some deviations, but they were all close to the actual illumination. It can be thought that scale has a certain influence on the color constancy of human vision.</p>
<p>Estimating multiple illuminations from one image at multiple scales is also in line with a verification conclusion obtained in Shi et al. (<xref ref-type="bibr" rid="B48">2016</xref>), namely, that multiple hypothetical illuminations from one image can help improve the accuracy of illumination estimation. Inspired by that, we propose herein a learning-based <bold>M</bold>ulti-scale <bold>R</bold>egion-<bold>w</bold>eighed <bold>N</bold>etwork guided by <bold>S</bold>emantics (MSRWNS) to estimate the illuminated color of the light source in a scene. First, the semantic context of an image is extracted, the image and its semantics are put into the network, and through a series of convolution layers, the region weights of the image at different scales are obtained. Then, through the weight-pooling layer (WPL), the illumination estimation on each scale is obtained. The global illumination is calculated by weighting on each scale.</p>
<p>The MSRWNS network differs from the existing methods and has three contributions, which follow.</p>
<list list-type="bullet">
<list-item><p>It estimates multiple global illuminations at different feature scales, and it obtains the final lighting by simple weighting.</p></list-item>
<list-item><p>Different from previous semantic-based methods, while using semantic guidance, a new region-WPL is used. The network layer simultaneously learns the contribution and local illumination of different regions in the image at each scale. It can, thus, effectively solve the illumination estimation error caused by the semantic segmentation error.</p></list-item>
<list-item><p>A large strip is used in the convolution to replace the max pooling layer in the network, which improves the speed of light estimation without reducing accuracy.</p></list-item>
</list>
<p>The rest of this article is organized as follows. In section 2, the structure of the proposed network and training strategy is presented, together with the related experimental content in section 3. Conclusions are given in section 4.</p>
</sec>
<sec id="s2">
<title>2. Multi-Scale Region-Weighed Network Guided by Semantics</title>
<p>Following the widely accepted simplified diagonal model (Finlayson et al., <xref ref-type="bibr" rid="B18">1994</xref>; Funt and Lewis, <xref ref-type="bibr" rid="B22">2000</xref>), the color of the light source is represented as</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mi>r</mml:mi><mml:mo>,</mml:mo><mml:mi>g</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>I</italic><sub><italic>c</italic></sub> &#x0003D; {<italic>I</italic><sub><italic>r</italic></sub>, <italic>I</italic><sub><italic>g</italic></sub>, <italic>I</italic><sub><italic>b</italic></sub>} is the color image under an unknown light source, <italic>R</italic><sub><italic>c</italic></sub> &#x0003D; {<italic>R</italic><sub><italic>r</italic></sub>, <italic>R</italic><sub><italic>g</italic></sub>, <italic>R</italic><sub><italic>b</italic></sub>} the color image recorded by a white-light source, and <italic>E</italic><sub><italic>c</italic></sub> &#x0003D; {<italic>E</italic><sub><italic>r</italic></sub>, <italic>E</italic><sub><italic>g</italic></sub>, <italic>E</italic><sub><italic>b</italic></sub>} the light source needed to be estimated from <italic>I</italic><sub><italic>c</italic></sub>.</p>
<p>A new color-space model has been used by color-constancy methods (Finlayson et al., <xref ref-type="bibr" rid="B19">2004</xref>; Barron, <xref ref-type="bibr" rid="B9">2015</xref>; Shi et al., <xref ref-type="bibr" rid="B48">2016</xref>) in recent years and has certain advantages, i.e., <italic>log</italic> &#x02212; <italic>uv</italic> space. <xref ref-type="fn" rid="fn0002"><sup>2</sup></xref> The calculation method proceeds as follows:</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M2"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>R</mml:mi><mml:mo>/</mml:mo><mml:mi>G</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>B</mml:mi><mml:mo>/</mml:mo><mml:mi>G</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>After estimating the light, it can be converted back to RGB space through a very simple formula:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M3"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>G</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>/</mml:mo><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mi>B</mml:mi><mml:mo>=</mml:mo><mml:mi>e</mml:mi><mml:mi>x</mml:mi><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:mi>z</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>z</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mo class="qopname">exp</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:mo class="qopname">exp</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msqrt><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where (<italic>L</italic><sub><italic>u</italic></sub>, <italic>L</italic><sub><italic>v</italic></sub>) is the image in <italic>log</italic> &#x02212; <italic>uv</italic> color space. (<italic>R, G, B</italic>) is the image in RGB color space.</p>
<sec>
<title>2.1. Problem Formulation</title>
<p>Generally, we only know the image <italic>I</italic><sub><italic>c</italic></sub> under an unknown light source <italic>E</italic><sub><italic>c</italic></sub> that must be estimated. The goal of color constancy is to estimate <italic>E</italic><sub><italic>c</italic></sub> from <italic>I</italic><sub><italic>c</italic></sub> and then compute it as <italic>E</italic><sub><italic>c</italic></sub> &#x0003D; <italic>I</italic><sub><italic>c</italic></sub>/<italic>R</italic><sub><italic>c</italic></sub>. How do we estimate <italic>E</italic><sub><italic>c</italic></sub> from <italic>I</italic><sub><italic>c</italic></sub>? To address this problem, we formulate color constancy as a regression problem.</p>
<p>First, we obtain the semantic context of the image, and for this, we use PSPNet (Zhao et al., <xref ref-type="bibr" rid="B59">2016</xref>), defined as <italic>I</italic><sub><italic>s</italic></sub>, and then convert <italic>I</italic><sub><italic>c</italic></sub> from RGB space to <italic>log</italic> &#x02212; <italic>uv</italic> space to obtain (<italic>I</italic><sub><italic>u</italic></sub>, <italic>I</italic><sub><italic>v</italic></sub>). Combining these three channels into a new three-channel image <italic>I</italic><sub><italic>n</italic></sub>, the aim is to find a mapping <italic>f</italic><sub>&#x003B8;</sub> such that <italic>f</italic><sub>&#x003B8;</sub>(<italic>I</italic><sub><italic>n</italic></sub>) &#x0003D; <italic>P</italic><sub><italic>uv</italic></sub>, where <italic>P</italic><sub><italic>uv</italic></sub> represents the light value in <italic>log</italic> &#x02212; <italic>uv</italic> space.</p>
<p>As mentioned earlier, the objects in the scene have a great effect on the color constancy of human vision (Van De Weijer et al., <xref ref-type="bibr" rid="B53">2007</xref>; Gao et al., <xref ref-type="bibr" rid="B24">2019</xref>). Therefore, the designed color-constancy algorithm should imitate the human visual system, i.e., the mapping <italic>f</italic><sub>&#x003B8;</sub> should be able to be based on semantic information and is used to support the larger contribution area and suppress the smaller contribution area in the image. Therefore, two aspects must be considered in the model: First, one must find a way to estimate the illumination of each area in the image, and, second, one must use an adaptive algorithm to integrate the illumination of these multiple areas into a global illumination. Supposing that <italic>R</italic> &#x0003D; <italic>R</italic><sub>1</sub>, <italic>R</italic><sub>2</sub>, ..., <italic>R</italic><sub><italic>n</italic></sub> represents <italic>n</italic> non-overlapping regions in the image <italic>I</italic><sub><italic>c</italic></sub>, <inline-formula><mml:math id="M4"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> represents the estimated scene illumination of the <italic>i</italic>th area <italic>R</italic><sub><italic>i</italic></sub>. Therefore, the mapping <italic>f</italic><sub>&#x003B8;</sub> can be expressed as follows:</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M5"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mtext class="textrm" mathvariant="normal">0</mml:mtext></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mtext class="textrm" mathvariant="normal">- 1</mml:mtext></mml:mrow></mml:munderover></mml:mstyle><mml:mi>w</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msubsup><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>w</italic>(<italic>R</italic><sub><italic>i</italic></sub>) represents the contribution of each area to the illumination estimation, i.e., the weight. In other words, if <italic>R</italic><sub><italic>i</italic></sub> contains the semantic context information, the corresponding <italic>w</italic>(<italic>R</italic><sub><italic>i</italic></sub>) has a higher weight value.</p>
</sec>
<sec>
<title>2.2. Network Architecture</title>
<p>It can be seen from Equation (4) that it is necessary to design a network structure <italic>f</italic><sub>&#x003B8;</sub> to be able to calculate <italic>w</italic>(<italic>R</italic><sub><italic>i</italic></sub>) and <inline-formula><mml:math id="M6"><mml:msubsup><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> in each area. The network structure is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>The architecture of Multi-scale Region-weighed Network guided by Semantics (MSRWNS) is trained to estimate the illuminant and weights of a given image in each region.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-841426-g0001.tif"/>
</fig>
<p>We learn <italic>P</italic><sub><italic>uv</italic></sub> at different scales from the intermediate features with scales of 1/16, 1/32, and 1/64. Defining the superscript <italic>j</italic> to represent the scale, <inline-formula><mml:math id="M7"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>u</mml:mi><mml:mi>v</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> then represents the estimated illumination in <italic>log</italic> &#x02212; <italic>uv</italic> space under the <italic>j</italic>th scale, which is converted back to RGB space according to Equation (3) to obtain <inline-formula><mml:math id="M8"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, where <italic>c</italic> &#x0003D; <italic>R, G, B</italic>. Finally, the final illumination <italic>P</italic><sub><italic>c</italic></sub> in RGB color space is obtained by simple calculation of the obtained illumination on different scales, and the formula is:</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M9"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">c</mml:mtext></mml:mrow><mml:mrow><mml:mtext class="textrm" mathvariant="normal">j</mml:mtext></mml:mrow></mml:msubsup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>C</italic><sub><italic>j</italic></sub> represents the weight of the illumination obtained at each scale. In this article, <italic>j</italic> &#x0003D; 1, 2, 3. In the final illumination calculation, it is assumed that the estimated illumination at different scales contributes the same to the scene illumination, namely, <italic>C</italic><sub><italic>j</italic></sub> &#x0003D; 1/3, <italic>j</italic> &#x0003D; 1, 2, 3.</p>
</sec>
<sec>
<title>2.3. Weight-Pooling Layer</title>
<p>In most of the previous methods, the extracted features are directly calculated through several fully connected layers to obtain a global illumination. However, it can be seen from the results of the earlier literature that the effect of illumination estimation is not significantly improved. Referring to Hu et al. (<xref ref-type="bibr" rid="B32">2017</xref>), we used a custom network layer, called a WPL, the main function of which is to converge the regional illumination into a global illumination, and at the same time, learn the weight <italic>w</italic>(<italic>R</italic><sub><italic>i</italic></sub>) of each area; <italic>R</italic><sub><italic>i</italic></sub> represents the <italic>i</italic>th area. The <italic>WPL</italic> on each scale is expressed as follows:</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M10"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mtext class="textrm" mathvariant="normal">0</mml:mtext></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mtext class="textrm" mathvariant="normal">- 1</mml:mtext></mml:mrow></mml:munderover></mml:mstyle><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msup><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>P</italic><sup><italic>j</italic></sup> represents the output of the WPL layer in the <italic>j</italic>th scale, <inline-formula><mml:math id="M11"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> the contribution of each region that must be learned on the <italic>j</italic>th scale, <inline-formula><mml:math id="M12"><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> the illumination of the <italic>i</italic> area on the <italic>j</italic>th scale that must be learned, and <italic>n</italic> represents the number of areas. In this study, each scale is <inline-formula><mml:math id="M13"><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>16</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>16</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>32</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>32</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>64</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>64</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula>.</p>
</sec>
<sec>
<title>2.4. Illumination Fusion for Multiple Scales</title>
<p>We made a simple attempt to determine how to select illumination at multiple scales, and we used several samples to train the 3 classification problems, hoping to obtain the probability of illumination at different scales. However, the training process model is difficult to converge and the effect is not ideal. Finally, for the sake of simplicity, it is assumed that each scale has the same contribution to the illumination estimation, so the average value of 3-scale illumination is taken as the final illumination in this section<xref ref-type="fn" rid="fn0003"><sup>3</sup></xref>, and the results also show that the average value is higher than that of a single scale on a variety of datasets.</p>
</sec>
<sec>
<title>2.5. Loss Function</title>
<p>At the time of training optimization, Euclidian loss is utilized for the network, defined</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M14"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E8"><label>(8)</label><mml:math id="M15"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msup><mml:mrow><mml:mo>&#x02016;</mml:mo><mml:msubsup><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>Loss</italic><sub><italic>j</italic></sub> is the loss function of the <italic>j</italic>th scale, <italic>E</italic><sup><italic>e</italic></sup> is the illumination estimated by the network, <italic>E</italic><sup><italic>t</italic></sup> is the ground truth illumination, and <italic>N</italic> represents the batch of the training samples. The loss is minimized using stochastic gradient descent with standard back-propagation.</p>
</sec>
<sec>
<title>2.6. Discussion of Network Structure</title>
<p>Either shallower (i.e., Shi et al. <xref ref-type="bibr" rid="B48">2016</xref>) or deeper networks (i.e., VGG-net Simonyan and Zisserman <xref ref-type="bibr" rid="B49">2014</xref>) could replace the pre-feature extraction in the proposed system. However, due to the color-constancy problem, the best network for feature extraction should have enough capacity to distinguish ambiguities and should be sensitive to different illuminants. We tried several common networks, such as AlexNet (Krizhevsky et al., <xref ref-type="bibr" rid="B39">2012</xref>), VggNet16 (Simonyan and Zisserman, <xref ref-type="bibr" rid="B49">2014</xref>), and VggNet19 (Simonyan and Zisserman, <xref ref-type="bibr" rid="B49">2014</xref>), and they all achieved good results. Finally, to improve the computational efficiency of the network, we simplified AlexNet (Krizhevsky et al., <xref ref-type="bibr" rid="B39">2012</xref>) and removed all of its pooling layers. The test results showed that the efficiency increased by approximately 4 and the accuracy by an average of 5.2%.</p>
</sec>
</sec>
<sec id="s3">
<title>3. Experimental Results</title>
<sec>
<title>3.1. Datasets</title>
<p>Because semantic information is needed in this study, we mainly used a semantic segmentation dataset, namely, the ADE20k dataset (Zhou et al., <xref ref-type="bibr" rid="B61">2016</xref>). Meanwhile, we used PSPNET (Zhao et al., <xref ref-type="bibr" rid="B59">2016</xref>; Zhou et al., <xref ref-type="bibr" rid="B60">2017</xref>)<xref ref-type="fn" rid="fn0004"><sup>4</sup></xref> to segment the semantic information for Color Checker (Gehler et al., <xref ref-type="bibr" rid="B26">2008</xref>; Zhou et al., <xref ref-type="bibr" rid="B60">2017</xref>) and NUS 8-Camera datasets (Cheng et al., <xref ref-type="bibr" rid="B15">2014</xref>).</p>
<p>During the training process, 100 images with accurate semantic segmentation were manually selected from the Color Checker dataset (Gehler et al., <xref ref-type="bibr" rid="B26">2008</xref>), and 200 images were extracted from the NUS 8-Camera dataset (Cheng et al., <xref ref-type="bibr" rid="B15">2014</xref>).</p>
<p>In addition, since the ADE20k dataset (Zhou et al., <xref ref-type="bibr" rid="B60">2017</xref>) does not provide illumination information, it is assumed that the images in the ADE20k dataset are corrected white-balanced images, and we, therefore, visually selected 500 images with normal color. Different lights were then rendered according to the following equations:</p>
<disp-formula id="E9"><label>(9)</label><mml:math id="M16"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E10"><label>(10)</label><mml:math id="M17"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none none none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:msub><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where (<italic>r</italic><sub><italic>i</italic></sub>, <italic>g</italic><sub><italic>i</italic></sub>, <italic>b</italic><sub><italic>i</italic></sub>) represents the simulated scene illumination. After simulation, more than 2,000 training images with illumination labels and accurate semantic information were obtained from the ADE20k (Zhou et al., <xref ref-type="bibr" rid="B60">2017</xref>) dataset. These 2,000 images were cut, mirrored horizontally and vertically, and rotated from (&#x02212;30<sup><italic>o</italic></sup>, 30<sup><italic>o</italic></sup>), 90<sup><italic>o</italic></sup>, 180<sup><italic>o</italic></sup>, and a total of approximately 20,000 pieces of data were obtained. Similarly, the images selected from the Color Checker (Gehler et al., <xref ref-type="bibr" rid="B26">2008</xref>) and NUS 8-Camera datasets (Cheng et al., <xref ref-type="bibr" rid="B15">2014</xref>) were processed in the same way to obtain approximately 8,000 images. All 28,000 images were randomly cropped and normalized to 512 &#x000D7; 512 as network input. As in the previous study, 3-fold cross-validation was used for all of the datasets, and each one run was used for training, one for validation, and one for testing.</p>
</sec>
<sec>
<title>3.2. Metrics</title>
<p>Color-constancy algorithms are often evaluated using a distance measure, such as Euclidean distance (Land, <xref ref-type="bibr" rid="B40">1977</xref>; Buchsbaum, <xref ref-type="bibr" rid="B14">1980</xref>), perceptual distances (Gijsenij et al., <xref ref-type="bibr" rid="B27">2009</xref>), reproduction angular error (Finlayson et al., <xref ref-type="bibr" rid="B21">2016</xref>), and angular error (Hordley and Finlayson, <xref ref-type="bibr" rid="B31">2004</xref>; Cheng et al., <xref ref-type="bibr" rid="B15">2014</xref>; Bianco et al., <xref ref-type="bibr" rid="B11">2015</xref>; Shi et al., <xref ref-type="bibr" rid="B48">2016</xref>). Within these metrics, the angular error is the most widely used in this field, and most of the existing works (Cheng et al., <xref ref-type="bibr" rid="B15">2014</xref>; Bianco et al., <xref ref-type="bibr" rid="B11">2015</xref>; Shi et al., <xref ref-type="bibr" rid="B48">2016</xref>) that tested using the Color Checker (Gehler et al., <xref ref-type="bibr" rid="B26">2008</xref>), NUS 8-Camera (Cheng et al., <xref ref-type="bibr" rid="B15">2014</xref>), and ADE20k datasetd (Zhou et al., <xref ref-type="bibr" rid="B60">2017</xref>) reported their performance in terms of the angular error in five indexes: <italic>Mean, Median</italic>, and <italic>TriMean</italic> of all errors; mean of the lowest 25% of errors (<italic>Best</italic> 25%); and mean of the highest 25% of errors (<italic>Worst</italic> 25%). Hence, in the present study, we also used these indexes. In particular, although we aimed to estimate the single illuminant in this study, we also tested performance on the popular outdoor multi-illuminant dataset (Arjan et al., <xref ref-type="bibr" rid="B5">2012</xref>). The angular error between the estimated illuminant <italic>E</italic><sub><italic>e</italic></sub> and ground-truth illuminant <inline-formula><mml:math id="M18"><mml:msubsup><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> is computed for each image as follows:</p>
<disp-formula id="E11"><label>(11)</label><mml:math id="M19"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>e</mml:mi><mml:mo>=</mml:mo><mml:mo class="qopname">arccos</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo><mml:msubsup><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mo>&#x02016;</mml:mo><mml:msub><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02016;</mml:mo><mml:mo>.</mml:mo><mml:mo>&#x02016;</mml:mo><mml:msubsup><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0002A;</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x02016;</mml:mo></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>The less the value of <italic>e</italic> is the better performance of the method.</p>
</sec>
<sec>
<title>3.3. Implementation Parameters</title>
<p>In this subsection, the parameter settings for training our final model are given.</p>
<p><bold>Feature-extraction-network selection:</bold> Different network structures, such as AlexNet (Krizhevsky et al., <xref ref-type="bibr" rid="B39">2012</xref>), VGGNet-19 (Simonyan and Zisserman, <xref ref-type="bibr" rid="B49">2014</xref>), and SqueezeNet (Iandola et al., <xref ref-type="bibr" rid="B33">2016</xref>), were used to test performance. The comparison diagram is shown in <xref ref-type="fig" rid="F2">Figure 2B</xref>, from which it can be seen that, although VGGNet-16 (Simonyan and Zisserman, <xref ref-type="bibr" rid="B49">2014</xref>) and VGGNet-19 (Simonyan and Zisserman, <xref ref-type="bibr" rid="B49">2014</xref>) network structures have a better effect than other networks, they take more time. Finally, considering effect and efficiency, the structure in <xref ref-type="fig" rid="F1">Figure 1</xref> is used in this study.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p><bold>(A&#x02013;D)</bold> Performance under different parameters.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-841426-g0002.tif"/>
</fig>
<p><bold>Network input and output:</bold> We compared the performance of the model trained with (<italic>I</italic><sub><italic>u</italic></sub>, <italic>I</italic><sub><italic>v</italic></sub>, <italic>I</italic><sub><italic>t</italic></sub>) three-channel image input and (<italic>I</italic><sub><italic>u</italic></sub>, <italic>I</italic><sub><italic>v</italic></sub>) two-channel input, and we tested the performance of different resolution images as network input. The comparison results are shown in <xref ref-type="table" rid="T1">Table 1</xref>. Considering effect and efficiency, the effect is the best when the network input image resolution is 512 and semantic information is used at the same time.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Accuracy of different network input sizes and whether semantics are used or not in the Color Checker dataset.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Size/Method</bold></th>
<th valign="top" align="center"><bold>Mean</bold></th>
<th valign="top" align="center"><bold>Median</bold></th>
<th valign="top" align="center"><bold>TriMean</bold></th>
<th valign="top" align="center"><bold>Best 25%</bold></th>
<th valign="top" align="center"><bold>Worst 25%</bold></th>
<th valign="top" align="center"><bold>Speed(ms)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic>256,T</italic></td>
<td valign="top" align="center">1.67</td>
<td valign="top" align="center">1.36</td>
<td valign="top" align="center">1.53</td>
<td valign="top" align="center">0.45</td>
<td valign="top" align="center">4.09</td>
<td valign="top" align="center">27</td>
</tr>
<tr>
<td valign="top" align="left"><italic>256,None</italic></td>
<td valign="top" align="center">1.81</td>
<td valign="top" align="center">1.44</td>
<td valign="top" align="center">1.62</td>
<td valign="top" align="center">0.55</td>
<td valign="top" align="center">4.31</td>
<td valign="top" align="center">23</td>
</tr>
<tr>
<td valign="top" align="left"><italic>512,T</italic></td>
<td valign="top" align="center" style="color:#ee1c23">1.64</td>
<td valign="top" align="center" style="color:#2e3092">1.17</td>
<td valign="top" align="center" style="color:#ee1c23">1.28</td>
<td valign="top" align="center" style="color:#ee1c23">0.31</td>
<td valign="top" align="center" style="color:#ee1c23">3.82</td>
<td valign="top" align="center">34</td>
</tr>
<tr>
<td valign="top" align="left"><italic>512,None</italic></td>
<td valign="top" align="center">1.66</td>
<td valign="top" align="center">1.33</td>
<td valign="top" align="center">1.48</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">4.01</td>
<td valign="top" align="center">30</td>
</tr>
<tr>
<td valign="top" align="left"><italic>128,T</italic></td>
<td valign="top" align="center">1.74</td>
<td valign="top" align="center">1.33</td>
<td valign="top" align="center">1.49</td>
<td valign="top" align="center">0.55</td>
<td valign="top" align="center">4.17</td>
<td valign="top" align="center">19</td>
</tr>
<tr>
<td valign="top" align="left"><italic>128,None</italic></td>
<td valign="top" align="center">1.76</td>
<td valign="top" align="center">1.35</td>
<td valign="top" align="center">1.51</td>
<td valign="top" align="center">0.55</td>
<td valign="top" align="center">4.27</td>
<td valign="top" align="center" style="color:#ee1c23">15</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>256, 512, and 128 are defined as different input sizes. T, None defined as the use of semantics or not. Red indicates best accuracy</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>For output, we tested the effects of 1&#x02013;5 scales outputs on illumination estimation, where 1 scale uses <inline-formula><mml:math id="M20"><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>64</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>64</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula>, 2 scales uses <inline-formula><mml:math id="M21"><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>32</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>32</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>64</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>64</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula>, 3 scales use <inline-formula><mml:math id="M22"><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>16</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>16</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>32</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>32</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>64</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>64</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula>, 4 scales use <inline-formula><mml:math id="M23"><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>16</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>16</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>32</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>32</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>64</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>64</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula>, and 5 scales use <inline-formula><mml:math id="M24"><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>8</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>16</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>16</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>32</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>32</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>64</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>64</mml:mn></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mfrac><mml:mrow><mml:mi>W</mml:mi></mml:mrow><mml:mrow><mml:mn>128</mml:mn></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mn>128</mml:mn></mml:mrow></mml:mfrac></mml:math></inline-formula>. The curves are shown in <xref ref-type="fig" rid="F2">Figures 2C,D</xref>. It can be seen that the effect is best when the scale is 4, but the time consumption is more than doubled when the scale is 3. Considering comprehensive effect and efficiency, we used 3 scales in the network. <xref ref-type="fig" rid="F3">Figure 3</xref> shows the intermediate results of estimating illumination at different 3 scales.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Local illumination estimation of network output at different scales. Left to right: <bold>(A)</bold> original color-biased image; <bold>(B)</bold> Local illumination at scale 3; <bold>(C)</bold> Local illumination at scale 2; <bold>(D)</bold> Local illumination at scale 1; <bold>(E)</bold> Final correction results. It can be seen from the figure that under different scales, due to different area sizes, the estimated illumination is also different, but the overall color is basically the same.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-841426-g0003.tif"/>
</fig>
<p><bold>Batch size and learning:</bold> For optimization, Adam (Kingma and Ba, <xref ref-type="bibr" rid="B37">2014</xref>) was employed with a batch size of 64, and a basic learning rate of 0.0001 was set for training. We trained all of the experiments over 4,000 epochs (2,50,000 iterations with batch size 64). The average angular error is calculated in the Color Checker dataset for every 20 epochs. The curve is shown in <xref ref-type="fig" rid="F2">Figure 2A</xref>. <xref ref-type="fig" rid="F4">Figure 4</xref> shows the resulting image and local illumination after different numbers of epochs.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Effect of training at different epochs. Left to right: <bold>(A)</bold> original image and normal image; <bold>(B)</bold> output image and local illumination at 100 epochs; <bold>(C)</bold> output image and local illumination at 500 epochs; <bold>(D)</bold> output image and local illumination at 2,000 epochs; <bold>(E)</bold> output image and local illumination 4,000 epochs. It can be seen that with an increasing number of epochs, the corrected image tends to the real result. In addition, it can be seen from the red circle that the area in which people are located and that at the junction of the wall and sofa are different in local color.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-841426-g0004.tif"/>
</fig>
</sec>
<sec>
<title>3.4. Comparison With State-of-the-Art Methods</title>
<p>To evaluate the performance of the proposed method and the influence of the WPL layer on the method. We trained two models. One used the mean value when the local region converges to the global, which is defined as <italic>MSRWNS-AVG</italic>, the other model used the <italic>WPL</italic> layer, which is defined as <italic>MSRWNS</italic>. In addition, we also estimated the illumination effect at each scale, defined as <italic>MSRWNS-1, MSRWNS-2</italic>, and <italic>MSRWNS-3</italic>. Several visualizations of processing outputs obtained using the proposed method are presented in <xref ref-type="fig" rid="F5">Figure 5</xref>.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Effect comparison between using weight-pooling layer (WPL) layer or not using it. For each group, <bold>(A)</bold> original image and normal image; <bold>(B)</bold> local illumination and image with WPL layer; <bold>(C)</bold> weight image and result without WPL; <bold>(D)</bold> semantic and global illumination (in global illumination, the left-hand side is the light estimated after using WPL and the right-hand side is that estimated without WPL). It can be seen that in an image with rich scenic content, the effect of using the WPL layer is significantly better than that of not using it. From the weighted image (the image after the weight is normalized, for which the gray value is high, the weight is large, and vice versa) and the local illumination image, different objects in the scene have different weights, and some contribute significantly. It can be seen from the estimated illumination and semantics that the illumination is also different in different semantic parts, and the approximate shape of the objects in the scene can be seen from the illumination image.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-841426-g0005.tif"/>
</fig>
<p>The quantitative performance comparison on the Color Checker dataset is presented in <xref ref-type="table" rid="T2">Table 2</xref> and the results on the NUS 8-Camera dataset in <xref ref-type="table" rid="T3">Table 3</xref>. The performance comparisons on the ADE20k (Zhou et al., <xref ref-type="bibr" rid="B60">2017</xref>), SFU Lab dataset (Barnard et al., <xref ref-type="bibr" rid="B8">2010</xref>), and SFU Gray-Ball dataset (200, <xref ref-type="bibr" rid="B1">2003</xref>) are shown in <xref ref-type="table" rid="T4">Tables 4</xref>&#x02013;<xref ref-type="table" rid="T6">6</xref>, respectively. Most CNN-based methods compare the effects on only two datasets, Color Checker dataset and NUS 8-Camera dataset. In order to make the comparison results consistent, the data in <xref ref-type="table" rid="T2">Tables 2</xref>, <xref ref-type="table" rid="T3">3</xref> are from Shi et al. (<xref ref-type="bibr" rid="B48">2016</xref>), while others were trained by us with the same training samples mentioned in the datasets section.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Performance comparison on Color Checker dataset (Gehler et al., <xref ref-type="bibr" rid="B26">2008</xref>).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Mean</bold></th>
<th valign="top" align="center"><bold>Median</bold></th>
<th valign="top" align="center"><bold>TriMean</bold></th>
<th valign="top" align="center"><bold>Best 25%</bold></th>
<th valign="top" align="center"><bold>Worst 25%</bold></th>
<th valign="top" align="center"><bold>95th percentile</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic>White-Patch (Brainard and Wandell, <xref ref-type="bibr" rid="B13">1986</xref>)</italic></td>
<td valign="top" align="center">7.55</td>
<td valign="top" align="center">5.68</td>
<td valign="top" align="center">6.35</td>
<td valign="top" align="center">1.45</td>
<td valign="top" align="center">16.12</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Edge-based Gamut (Barnard, <xref ref-type="bibr" rid="B6">2000</xref>)</italic></td>
<td valign="top" align="center">6.52</td>
<td valign="top" align="center">5.04</td>
<td valign="top" align="center">5.43</td>
<td valign="top" align="center">1.90</td>
<td valign="top" align="center">13.58</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Gray-World (Buchsbaum, <xref ref-type="bibr" rid="B14">1980</xref>)</italic></td>
<td valign="top" align="center">6.36</td>
<td valign="top" align="center">6.28</td>
<td valign="top" align="center">6.28</td>
<td valign="top" align="center">2.33</td>
<td valign="top" align="center">10.58</td>
<td valign="top" align="center">11.3</td>
</tr>
<tr>
<td valign="top" align="left"><italic>1st-order Gray-Edge (Weijer et al., <xref ref-type="bibr" rid="B55">2007</xref>)</italic></td>
<td valign="top" align="center">5.33</td>
<td valign="top" align="center">4.52</td>
<td valign="top" align="center">4.73</td>
<td valign="top" align="center">1.86</td>
<td valign="top" align="center">10.03</td>
<td valign="top" align="center">11.0</td>
</tr>
<tr>
<td valign="top" align="left"><italic>2nd-order Gray-Edge (Weijer et al., <xref ref-type="bibr" rid="B55">2007</xref>)</italic></td>
<td valign="top" align="center">5.13</td>
<td valign="top" align="center">4.44</td>
<td valign="top" align="center">4.62</td>
<td valign="top" align="center">2.11</td>
<td valign="top" align="center">9.26</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Shades-of-Gray (Finlayson and Trezzi, <xref ref-type="bibr" rid="B20">2004</xref>)</italic></td>
<td valign="top" align="center">4.93</td>
<td valign="top" align="center">4.01</td>
<td valign="top" align="center">4.23</td>
<td valign="top" align="center">1.14</td>
<td valign="top" align="center">10.20</td>
<td valign="top" align="center">11.9</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Bayesian (Gehler et al., <xref ref-type="bibr" rid="B26">2008</xref>)</italic></td>
<td valign="top" align="center">4.82</td>
<td valign="top" align="center">3.46</td>
<td valign="top" align="center">3.88</td>
<td valign="top" align="center">1.26</td>
<td valign="top" align="center">10.49</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left"><italic>General Gray-World (Barnard et al., <xref ref-type="bibr" rid="B7">2002</xref>)</italic></td>
<td valign="top" align="center">4.66</td>
<td valign="top" align="center">3.48</td>
<td valign="top" align="center">3.81</td>
<td valign="top" align="center">1.00</td>
<td valign="top" align="center">10.09</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Intersection-based Gamut (Gehler et al., <xref ref-type="bibr" rid="B26">2008</xref>)</italic></td>
<td valign="top" align="center">4.20</td>
<td valign="top" align="center">2.39</td>
<td valign="top" align="center">2.93</td>
<td valign="top" align="center">0.51</td>
<td valign="top" align="center">10.70</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Pixel-based Gamut (Gehler et al., <xref ref-type="bibr" rid="B26">2008</xref>)</italic></td>
<td valign="top" align="center">4.20</td>
<td valign="top" align="center">2.33</td>
<td valign="top" align="center">2.91</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">10.72</td>
<td valign="top" align="center">14.1</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Natural Image Statistics (?)</italic></td>
<td valign="top" align="center">4.19</td>
<td valign="top" align="center">3.13</td>
<td valign="top" align="center">3.45</td>
<td valign="top" align="center">1.00</td>
<td valign="top" align="center">9.22</td>
<td valign="top" align="center">11.7</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Bright Pixels (Joze et al., <xref ref-type="bibr" rid="B36">2012</xref>)</italic></td>
<td valign="top" align="center">3.98</td>
<td valign="top" align="center">2.61</td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Spatio-spectral (GenPrior) (Hirakawa et al., <xref ref-type="bibr" rid="B30">2012</xref>)</italic></td>
<td valign="top" align="center">3.59</td>
<td valign="top" align="center">2.96</td>
<td valign="top" align="center">3.10</td>
<td valign="top" align="center">0.95</td>
<td valign="top" align="center">7.61</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Cheng et al. (<xref ref-type="bibr" rid="B15">2014</xref>)</italic></td>
<td valign="top" align="center">3.52</td>
<td valign="top" align="center">2.14</td>
<td valign="top" align="center">2.47</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">8.74</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Corrected-Moment (19 Color) (Finlayson, <xref ref-type="bibr" rid="B17">2013</xref>)</italic></td>
<td valign="top" align="center">3.50</td>
<td valign="top" align="center">2.60</td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">8.6</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Corrected-Moment (19 Color)&#x0002A; (Finlayson, <xref ref-type="bibr" rid="B17">2013</xref>)</italic></td>
<td valign="top" align="center">2.96</td>
<td valign="top" align="center">2.15</td>
<td valign="top" align="center">2.37</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center">6.69</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Corrected-Moment (19 Edge) (Finlayson, <xref ref-type="bibr" rid="B17">2013</xref>)</italic></td>
<td valign="top" align="center">2.82</td>
<td valign="top" align="center">2.00</td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">6.9</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Corrected-Moment (19 Edge)&#x0002A; (Finlayson, <xref ref-type="bibr" rid="B17">2013</xref>)</italic></td>
<td valign="top" align="center">3.12</td>
<td valign="top" align="center">2.38</td>
<td valign="top" align="center">2.59</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">6.46</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Regression Tree (Cheng et al., <xref ref-type="bibr" rid="B16">2015</xref>)</italic></td>
<td valign="top" align="center">2.42</td>
<td valign="top" align="center">1.65</td>
<td valign="top" align="center">1.75</td>
<td valign="top" align="center">0.38</td>
<td valign="top" align="center">5.87</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left"><italic>CNN (Bianco et al., <xref ref-type="bibr" rid="B11">2015</xref>)</italic></td>
<td valign="top" align="center">2.36</td>
<td valign="top" align="center">1.98</td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">-</td>
</tr>
<tr>
<td valign="top" align="left"><italic>CCC (Barron, <xref ref-type="bibr" rid="B9">2015</xref>)</italic></td>
<td valign="top" align="center">1.95</td>
<td valign="top" align="center">1.22</td>
<td valign="top" align="center">1.38</td>
<td valign="top" align="center">0.35</td>
<td valign="top" align="center">4.76</td>
<td valign="top" align="center">5.85</td>
</tr>
<tr>
<td valign="top" align="left"><italic>DS-Net (Shi et al., <xref ref-type="bibr" rid="B48">2016</xref>)</italic></td>
<td valign="top" align="center">1.90</td>
<td valign="top" align="center">1.12</td>
<td valign="top" align="center">1.33</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center">4.84</td>
<td valign="top" align="center">5.99</td>
</tr>
<tr>
<td valign="top" align="left"><italic>FC4 (Hu et al., <xref ref-type="bibr" rid="B32">2017</xref>)</italic></td>
<td valign="top" align="center">1.77</td>
<td valign="top" align="center" style="color:#ee1c23">1.11</td>
<td valign="top" align="center">1.29</td>
<td valign="top" align="center">0.34</td>
<td valign="top" align="center">4.29</td>
<td valign="top" align="center">5.44</td>
</tr>
<tr>
<td valign="top" align="left"><italic>MSRWNS-AVG</italic></td>
<td valign="top" align="center">1.93</td>
<td valign="top" align="center">1.38</td>
<td valign="top" align="center">1.42</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">4.32</td>
<td valign="top" align="center">4.20</td>
</tr>
<tr>
<td valign="top" align="left"><italic>MSRWNS-1</italic></td>
<td valign="top" align="center">1.72</td>
<td valign="top" align="center">1.16</td>
<td valign="top" align="center">1.33</td>
<td valign="top" align="center" style="color:#2e3092">0.32</td>
<td valign="top" align="center" style="color:#2e3092">3.79</td>
<td valign="top" align="center">4.36</td>
</tr>
<tr>
<td valign="top" align="left"><italic>MSRWNS-2</italic></td>
<td valign="top" align="center" style="color:#2e3092">1.68</td>
<td valign="top" align="center">1.13</td>
<td valign="top" align="center" style="color:#2e3092">1.28</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center">3.84</td>
<td valign="top" align="center">4.44</td>
</tr>
<tr>
<td valign="top" align="left"><italic>MSRWNS-3</italic></td>
<td valign="top" align="center">1.71</td>
<td valign="top" align="center">1.13</td>
<td valign="top" align="center">1.31</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center">3.82</td>
<td valign="top" align="center" style="color:#2e3092">4.18</td>
</tr>
<tr>
<td valign="top" align="left"><italic>MSRWNS</italic></td>
<td valign="top" align="center" style="color:#ee1c23">1.64</td>
<td valign="top" align="center" style="color:#2e3092">1.13</td>
<td valign="top" align="center" style="color:#ee1c23">1.28</td>
<td valign="top" align="center" style="color:#ee1c23">0.31</td>
<td valign="top" align="center" style="color:#ee1c23">3.78</td>
<td valign="top" align="center" style="color:#ee1c23">4.07</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>For each metric, red indicates the best performance and blue indicates the second-best performance</italic>.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Performance comparison on NUS 8-Camera dataset (Cheng et al., <xref ref-type="bibr" rid="B15">2014</xref>).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Mean</bold></th>
<th valign="top" align="center"><bold>Median</bold></th>
<th valign="top" align="center"><bold>TriMean</bold></th>
<th valign="top" align="center"><bold>Best 25%</bold></th>
<th valign="top" align="center"><bold>Worst 25%</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic>White-Patch (Brainard and Wandell, <xref ref-type="bibr" rid="B13">1986</xref>)</italic></td>
<td valign="top" align="center">10.62</td>
<td valign="top" align="center">10.58</td>
<td valign="top" align="center">10.49</td>
<td valign="top" align="center">1.86</td>
<td valign="top" align="center">19.45</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Edge-based Gamut (Barnard, <xref ref-type="bibr" rid="B6">2000</xref>)</italic></td>
<td valign="top" align="center">8.43</td>
<td valign="top" align="center">7.05</td>
<td valign="top" align="center">7.37</td>
<td valign="top" align="center">2.41</td>
<td valign="top" align="center">16.08</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Pixel-Based Gamut (Gehler et al., <xref ref-type="bibr" rid="B26">2008</xref>)</italic></td>
<td valign="top" align="center">7.70</td>
<td valign="top" align="center">6.71</td>
<td valign="top" align="center">6.90</td>
<td valign="top" align="center">2.51</td>
<td valign="top" align="center">14.05</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Intersection-based Gamut (Gehler et al., <xref ref-type="bibr" rid="B26">2008</xref>)</italic></td>
<td valign="top" align="center">7.20</td>
<td valign="top" align="center">5.96</td>
<td valign="top" align="center">6.28</td>
<td valign="top" align="center">2.20</td>
<td valign="top" align="center">13.61</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Gray-World (Buchsbaum, <xref ref-type="bibr" rid="B14">1980</xref>)</italic></td>
<td valign="top" align="center">4.14</td>
<td valign="top" align="center">3.20</td>
<td valign="top" align="center">3.39</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">9.00</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Bayesian (Gehler et al., <xref ref-type="bibr" rid="B26">2008</xref>)</italic></td>
<td valign="top" align="center">3.67</td>
<td valign="top" align="center">2.73</td>
<td valign="top" align="center">2.91</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">8.21</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Natural Image Statistics (?)</italic></td>
<td valign="top" align="center">3.71</td>
<td valign="top" align="center">2.60</td>
<td valign="top" align="center">2.84</td>
<td valign="top" align="center">0.79</td>
<td valign="top" align="center">8.47</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Shades-of-Gray (Finlayson and Trezzi, <xref ref-type="bibr" rid="B20">2004</xref>)</italic></td>
<td valign="top" align="center">3.40</td>
<td valign="top" align="center">2.57</td>
<td valign="top" align="center">2.73</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">7.41</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Spatio-spectral (ML) (Hirakawa et al., <xref ref-type="bibr" rid="B30">2012</xref>)</italic></td>
<td valign="top" align="center">3.11</td>
<td valign="top" align="center">2.49</td>
<td valign="top" align="center">2.60</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">6.59</td>
</tr>
<tr>
<td valign="top" align="left"><italic>2nd-order Gray-Edge (Weijer et al., <xref ref-type="bibr" rid="B55">2007</xref>)</italic></td>
<td valign="top" align="center">3.20</td>
<td valign="top" align="center">2.26</td>
<td valign="top" align="center">2.44</td>
<td valign="top" align="center">0.75</td>
<td valign="top" align="center">7.27</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Bright Pixels (Joze et al., <xref ref-type="bibr" rid="B36">2012</xref>)</italic></td>
<td valign="top" align="center">3.17</td>
<td valign="top" align="center">2.41</td>
<td valign="top" align="center">2.55</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">7.02</td>
</tr>
<tr>
<td valign="top" align="left"><italic>1st-order Gray-Edge (Weijer et al., <xref ref-type="bibr" rid="B55">2007</xref>)</italic></td>
<td valign="top" align="center">3.20</td>
<td valign="top" align="center">2.22</td>
<td valign="top" align="center">2.43</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">7.36</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Spatio-spectral (GenPrior) (Hirakawa et al., <xref ref-type="bibr" rid="B30">2012</xref>)</italic></td>
<td valign="top" align="center">2.96</td>
<td valign="top" align="center">2.33</td>
<td valign="top" align="center">2.47</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">6.18</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Corrected-Moment (19 Edge)&#x0002A;(Finlayson, <xref ref-type="bibr" rid="B17">2013</xref>)</italic></td>
<td valign="top" align="center">3.03</td>
<td valign="top" align="center">2.11</td>
<td valign="top" align="center">2.25</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">7.08</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Corrected-Moment (19 Color)&#x0002A;(Finlayson, <xref ref-type="bibr" rid="B17">2013</xref>)</italic></td>
<td valign="top" align="center">3.05</td>
<td valign="top" align="center">1.90</td>
<td valign="top" align="center">2.13</td>
<td valign="top" align="center">0.65</td>
<td valign="top" align="center">7.41</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Cheng et al. (<xref ref-type="bibr" rid="B15">2014</xref>)</italic></td>
<td valign="top" align="center">2.96</td>
<td valign="top" align="center">2.04</td>
<td valign="top" align="center">2.24</td>
<td valign="top" align="center">0.62</td>
<td valign="top" align="center">6.61</td>
</tr>
<tr>
<td valign="top" align="left"><italic>CCC (Barron, <xref ref-type="bibr" rid="B9">2015</xref>)</italic></td>
<td valign="top" align="center">2.38</td>
<td valign="top" align="center">1.48</td>
<td valign="top" align="center">1.69</td>
<td valign="top" align="center">0.45</td>
<td valign="top" align="center">5.85</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Regression Tree (Cheng et al., <xref ref-type="bibr" rid="B16">2015</xref>)</italic></td>
<td valign="top" align="center">2.36</td>
<td valign="top" align="center">1.59</td>
<td valign="top" align="center">1.74</td>
<td valign="top" align="center">0.49</td>
<td valign="top" align="center">5.54</td>
</tr>
<tr>
<td valign="top" align="left"><italic>DS-Net (Shi et al., <xref ref-type="bibr" rid="B48">2016</xref>)</italic></td>
<td valign="top" align="center">2.24</td>
<td valign="top" align="center">1.46</td>
<td valign="top" align="center">1.68</td>
<td valign="top" align="center">0.48</td>
<td valign="top" align="center">6.08</td>
</tr>
<tr>
<td valign="top" align="left"><italic>FC4 (Hu et al., <xref ref-type="bibr" rid="B32">2017</xref>)</italic></td>
<td valign="top" align="center">2.12</td>
<td valign="top" align="center">1.53</td>
<td valign="top" align="center">1.67</td>
<td valign="top" align="center">0.48</td>
<td valign="top" align="center" style="color:#2e3092">4.78</td>
</tr>
<tr>
<td valign="top" align="left"><italic>MSRWNS-AVG</italic></td>
<td valign="top" align="center">2.13</td>
<td valign="top" align="center">1.51</td>
<td valign="top" align="center">1.72</td>
<td valign="top" align="center">0.58</td>
<td valign="top" align="center">5.44</td>
</tr>
<tr>
<td valign="top" align="left"><italic>MSRWNS-1</italic></td>
<td valign="top" align="center" style="color:#2e3092">2.12</td>
<td valign="top" align="center" style="color:#2e3092">1.45</td>
<td valign="top" align="center" style="color:#2e3092">1.64</td>
<td valign="top" align="center" style="color:#2e3092">0.46</td>
<td valign="top" align="center">5.12</td>
</tr>
<tr>
<td valign="top" align="left"><italic>MSRWNS-2</italic></td>
<td valign="top" align="center">2.11</td>
<td valign="top" align="center">1.45</td>
<td valign="top" align="center">1.66</td>
<td valign="top" align="center">0.51</td>
<td valign="top" align="center">5.29</td>
</tr>
<tr>
<td valign="top" align="left"><italic>MSRWNS-3</italic></td>
<td valign="top" align="center">2.11</td>
<td valign="top" align="center">1.46</td>
<td valign="top" align="center">1.67</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">5.33</td>
</tr>
<tr>
<td valign="top" align="left"><italic>MSRWNS</italic></td>
<td valign="top" align="center" style="color:#ee1c23">2.11</td>
<td valign="top" align="center" style="color:#ee1c23">1.45</td>
<td valign="top" align="center" style="color:#ee1c23">1.64</td>
<td valign="top" align="center" style="color:#ee1c23">0.45</td>
<td valign="top" align="center" style="color:#ee1c23">4.77</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>For each metric, red indicates the best performance and blue indicates the second-best performance</italic>.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Performance comparison on ADE20k dataset (Zhou et al., <xref ref-type="bibr" rid="B60">2017</xref>).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Mean</bold></th>
<th valign="top" align="center"><bold>Median</bold></th>
<th valign="top" align="center"><bold>TriMean</bold></th>
<th valign="top" align="center"><bold>Best 25%</bold></th>
<th valign="top" align="center"><bold>Worst 25%</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic>CCC (disc&#x0002B;ext) (Barron, <xref ref-type="bibr" rid="B9">2015</xref>)</italic></td>
<td valign="top" align="center">2.14</td>
<td valign="top" align="center">1.66</td>
<td valign="top" align="center">1.82</td>
<td valign="top" align="center">0.32</td>
<td valign="top" align="center">4.24</td>
</tr>
<tr>
<td valign="top" align="left"><italic>CNN (Bianco et al., <xref ref-type="bibr" rid="B11">2015</xref>)</italic></td>
<td valign="top" align="center">1.96</td>
<td valign="top" align="center">1.32</td>
<td valign="top" align="center">1.14</td>
<td valign="top" align="center" style="color:#2e3092">0.23</td>
<td valign="top" align="center">3.94</td>
</tr>
<tr>
<td valign="top" align="left"><italic>DS-Net (Shi et al., <xref ref-type="bibr" rid="B48">2016</xref>)</italic></td>
<td valign="top" align="center">1.68</td>
<td valign="top" align="center" style="color:#2e3092">0.96</td>
<td valign="top" align="center">1.06</td>
<td valign="top" align="center">0.26</td>
<td valign="top" align="center">3.86</td>
</tr>
<tr>
<td valign="top" align="left"><italic>FC4 (Hu et al., <xref ref-type="bibr" rid="B32">2017</xref>)</italic></td>
<td valign="top" align="center" style="color:#2e3092">1.56</td>
<td valign="top" align="center">1.32</td>
<td valign="top" align="center" style="color:#2e3092">1.02</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center">3.86</td>
</tr>
<tr>
<td valign="top" align="left"><italic>MSRWNS-AVG</italic></td>
<td valign="top" align="center">1.66</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">1.44</td>
<td valign="top" align="center">0.32</td>
<td valign="top" align="center" style="color:#2e3092">3.66</td>
</tr>
<tr>
<td valign="top" align="left"><italic>MSRWNS</italic></td>
<td valign="top" align="center" style="color:#ee1c23">1.18</td>
<td valign="top" align="center" style="color:#ee1c23">0.61</td>
<td valign="top" align="center" style="color:#ee1c23">0.83</td>
<td valign="top" align="center" style="color:#ee1c23">0.11</td>
<td valign="top" align="center" style="color:#ee1c23">2.87</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>For each metric, red indicates the best performance and blue indicates the second-best performance</italic>.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Performance comparison on SFU Lab dataset (Barnard et al., <xref ref-type="bibr" rid="B8">2010</xref>).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Mean</bold></th>
<th valign="top" align="center"><bold>Median</bold></th>
<th valign="top" align="center"><bold>TriMean</bold></th>
<th valign="top" align="center"><bold>Best 25%</bold></th>
<th valign="top" align="center"><bold>Worst 25%</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic>CCC (disc&#x0002B;ext) (Barron, <xref ref-type="bibr" rid="B9">2015</xref>)</italic></td>
<td valign="top" align="center">3.77</td>
<td valign="top" align="center">2.19</td>
<td valign="top" align="center">2.21</td>
<td valign="top" align="center">0.42</td>
<td valign="top" align="center">8.87</td>
</tr>
<tr>
<td valign="top" align="left"><italic>CNN (Bianco et al., <xref ref-type="bibr" rid="B11">2015</xref>)</italic></td>
<td valign="top" align="center">3.18</td>
<td valign="top" align="center">2.31</td>
<td valign="top" align="center">2.40</td>
<td valign="top" align="center">0.39</td>
<td valign="top" align="center">6.98</td>
</tr>
<tr>
<td valign="top" align="left"><italic>DS-Net (Shi et al., <xref ref-type="bibr" rid="B48">2016</xref>)</italic></td>
<td valign="top" align="center" style="color:#2e3092">2.93</td>
<td valign="top" align="center">2.11</td>
<td valign="top" align="center">2.23</td>
<td valign="top" align="center">0.34</td>
<td valign="top" align="center">5.87</td>
</tr>
<tr>
<td valign="top" align="left"><italic>FC4 (Hu et al., <xref ref-type="bibr" rid="B32">2017</xref>)</italic></td>
<td valign="top" align="center">2.99</td>
<td valign="top" align="center" style="color:#2e3092">1.78</td>
<td valign="top" align="center" style="color:#2e3092">2.11</td>
<td valign="top" align="center">0.33</td>
<td valign="top" align="center" style="color:#2e3092">4.62</td>
</tr>
<tr>
<td valign="top" align="left"><italic>MSRWNS-AVG</italic></td>
<td valign="top" align="center">3.15</td>
<td valign="top" align="center">1.88</td>
<td valign="top" align="center">2.01</td>
<td valign="top" align="center" style="color:#2e3092">0.32</td>
<td valign="top" align="center">4.88</td>
</tr>
<tr>
<td valign="top" align="left"><italic>MSRWNS</italic></td>
<td valign="top" align="center" style="color:#ee1c23">2.82</td>
<td valign="top" align="center" style="color:#ee1c23">1.71</td>
<td valign="top" align="center" style="color:#ee1c23">1.85</td>
<td valign="top" align="center" style="color:#ee1c23">0.26</td>
<td valign="top" align="center" style="color:#ee1c23">4.65</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>For each metric, red indicates the best performance and blue the second-best performance</italic>.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>Performance comparison on SFU Gray-Ball dataset (200, <xref ref-type="bibr" rid="B1">2003</xref>).</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Mean</bold></th>
<th valign="top" align="center"><bold>Median</bold></th>
<th valign="top" align="center"><bold>TriMean</bold></th>
<th valign="top" align="center"><bold>Best 25%</bold></th>
<th valign="top" align="center"><bold>Worst 25%</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic>CCC (disc&#x0002B;ext) (Barron, <xref ref-type="bibr" rid="B9">2015</xref>)</italic></td>
<td valign="top" align="center">3.31</td>
<td valign="top" align="center">1.66</td>
<td valign="top" align="center">1.82</td>
<td valign="top" align="center">0.32</td>
<td valign="top" align="center">4.24</td>
</tr>
<tr>
<td valign="top" align="left"><italic>CNN (Bianco et al., <xref ref-type="bibr" rid="B11">2015</xref>)</italic></td>
<td valign="top" align="center">2.96</td>
<td valign="top" align="center">1.32</td>
<td valign="top" align="center">1.14</td>
<td valign="top" align="center" style="color:#2e3092">0.23</td>
<td valign="top" align="center">3.94</td>
</tr>
<tr>
<td valign="top" align="left"><italic>DS-Net (Shi et al., <xref ref-type="bibr" rid="B48">2016</xref>)</italic></td>
<td valign="top" align="center">2.41</td>
<td valign="top" align="center" style="color:#2e3092">0.96</td>
<td valign="top" align="center" style="color:#2e3092">1.06</td>
<td valign="top" align="center">0.26</td>
<td valign="top" align="center">3.86</td>
</tr>
<tr>
<td valign="top" align="left"><italic>FC4 (Hu et al., <xref ref-type="bibr" rid="B32">2017</xref>)</italic></td>
<td valign="top" align="center">2.33</td>
<td valign="top" align="center">1.12</td>
<td valign="top" align="center">1.46</td>
<td valign="top" align="center">0.41</td>
<td valign="top" align="center" style="color:#2e3092">3.76</td>
</tr>
<tr>
<td valign="top" align="left"><italic>MSRWNS-AVG</italic></td>
<td valign="top" align="center" style="color:#2e3092">2.18</td>
<td valign="top" align="center">1.12</td>
<td valign="top" align="center">1.34</td>
<td valign="top" align="center">0.26</td>
<td valign="top" align="center">3.88</td>
</tr>
<tr>
<td valign="top" align="left"><italic>MSRWNS</italic></td>
<td valign="top" align="center" style="color:#ee1c23">1.83</td>
<td valign="top" align="center" style="color:#ee1c23">0.82</td>
<td valign="top" align="center" style="color:#2e3092">0.94</td>
<td valign="top" align="center" style="color:#ee1c23">0.20</td>
<td valign="top" align="center" style="color:#ee1c23">3.65</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>For each metric, red indicates the best performance and blue indicates the second-best performance</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>In addition, we compared our method with most of the existing works; typical works include the Deep Specialized Network (DS-Net) (Shi et al., <xref ref-type="bibr" rid="B48">2016</xref>) and Fully Convolutional Color Constancy with confidence-weighted pooling (FC4) (Hu et al., <xref ref-type="bibr" rid="B32">2017</xref>). Several visualizations of testing outputs obtained using the proposed method are presented in <xref ref-type="fig" rid="F6">Figures 6</xref>, <xref ref-type="fig" rid="F7">7</xref>.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Visual comparison results. <bold>(A)</bold> Original image; <bold>(B)</bold> result obtained by CNN (Bianco et al., <xref ref-type="bibr" rid="B11">2015</xref>); <bold>(C)</bold> result obtained by DS-Net (Shi et al., <xref ref-type="bibr" rid="B48">2016</xref>); <bold>(D)</bold> result obtained by FC4 (Hu et al., <xref ref-type="bibr" rid="B32">2017</xref>); <bold>(E)</bold> result obtained by proposed method; <bold>(F)</bold> ground truth. Regardless of quantification or visual effects, the proposed method shows better performance, especially in the first, second, and fourth lines, because there is a large area in the scene that can be accurately segmented, and the correction result of the method proposed in this article is very close to the real image. In the second line, because the color of the sofa and that of the light are relatively close. Although the network model considers the contribution of different regions, it is difficult to eliminate the color cast caused by the similar color of the light and the surface of the object. From the corrected image look, the image has a slight red tint. In the fourth line of the image, because the objects in the scene are too singular, the red objects on the left and the white objects on the right have greater contributions and the color of the objects on the left is biased in the result.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-841426-g0006.tif"/>
</fig>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Visual comparison between the proposed method with the auto-white-balance (AWB) function of the Canon D7100 camera used. Left to right: <bold>(A)</bold> original image; <bold>(B)</bold> results by D7100 camera with AWB; <bold>(C)</bold> results by CNN; <bold>(D)</bold> results by DS-Net; <bold>(E)</bold> results by present study.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-841426-g0007.tif"/>
</fig>
<p>From <xref ref-type="table" rid="T2">Tables 2</xref>, <xref ref-type="table" rid="T3">3</xref>, it can be seen that the mean error of the proposed method is reduced by 12.3% on the Color Checker dataset and by 5.8% compared to DS-Net (Shi et al., <xref ref-type="bibr" rid="B48">2016</xref>), and reduced by 7.3% on the Color Checker dataset and reduced by 0.5% compared to FC4 (Hu et al., <xref ref-type="bibr" rid="B32">2017</xref>). In addition, using the WPL layer under a single scale gives a better result than that obtained using the mean. The effect of using the mean of three scales on most indicators is better than the result of using a single scale.</p>
<p>In particular, the proposed method shows the best performance on the ADE20k (Zhou et al., <xref ref-type="bibr" rid="B60">2017</xref>) dataset. The mean angular error is lower than that of DS-Net (Shi et al., <xref ref-type="bibr" rid="B48">2016</xref>) by 29.7% (from 1.68 to 1.18), and the mean error of the worst 25% was reduced by 25.6% (from 3.86 to 2.87) compared to DS-Net (Shi et al., <xref ref-type="bibr" rid="B48">2016</xref>). Other indicators have also been reduced to a certain extent. The reason for these results is that the semantic segmentation model used is trained based on the ADE20k (Zhou et al., <xref ref-type="bibr" rid="B60">2017</xref>) dataset. The accuracy of the semantics on this dataset is high, and there are parts of training data in the ADE20k dataset (Zhou et al., <xref ref-type="bibr" rid="B60">2017</xref>).</p>
<p>In addition, we also provided several natural examples captured by a Canon D7100 without auto-white-balance <xref ref-type="fn" rid="fn0005"><sup>5</sup></xref>, as shown in <xref ref-type="fig" rid="F7">Figure 7</xref>, and we obtained several images of natural scenes with more accurate colors from the Internet, and then performed some random color casting. Results are shown in <xref ref-type="fig" rid="F8">Figure 8</xref>.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Result in a natural scene. From left to right. <bold>(A)</bold> Original image; <bold>(B)</bold> random color-cast image; <bold>(C)</bold> corrected image; <bold>(D)</bold> local-light image; <bold>(E)</bold> final-light color. It can be seen from the first row that, although there is a phenomenon of highlight overflow in some areas of the sky (due to random color casts leading to pixel overflow in some areas), the overall color is close to the original image; the second row is due to the random color cast. The resulting color cast is small, and the corrected image is basically the same as the original image. In the third row, it can be seen that, although the corrected image is different from the original image, the visual perception is more realistic.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-841426-g0008.tif"/>
</fig>
<p>It can be seen From <xref ref-type="fig" rid="F7">Figure 7</xref> that on the image taken indoors (the first line in <xref ref-type="fig" rid="F7">Figure 7</xref>) the image corrected by a CNN (Bianco et al., <xref ref-type="bibr" rid="B11">2015</xref>) is obviously yellowish, and that corrected by DS-Net is slightly greenish. The white-balance effect of the camera and our result is similar and look relatively natural. In outdoor scenes, it can be seen from the second line that the image corrected by a CNN (Bianco et al., <xref ref-type="bibr" rid="B11">2015</xref>) has a very obvious color cast. The other types are visually more natural. Our results are seen in the sky part of the image, in which the clouds are more realistic. The scene in the third row exhibits little difference in visual effects. This may be due to sufficient sunlight in the shooting scene, and the objects in the scene receive very uniform illumination. In this way, any method can obtain better results more accurately. From <xref ref-type="fig" rid="F8">Figure 8</xref>, it can be found from the first row that, although there is a phenomenon of highlight overflow in some areas of the sky (due to random color casts leading to pixel overflow in some areas), the overall color is close to that of the real image; the second row is due to the random color cast. The resulting color cast is small and the corrected image is basically the same as the original image. In the third row, it can be seen that, although the corrected image is different from the original image, the visual perception is better.</p>
</sec>
<sec>
<title>3.5. Efficiency</title>
<p>The code used to test the efficiency of the proposed method is based on Tensorflow (Rampasek and Goldenberg, <xref ref-type="bibr" rid="B46">2016</xref>) and training took approximately 5 h, after which the loss tended to stabilize. In the testing phase, we converted the model to that of Caffe (Jia et al., <xref ref-type="bibr" rid="B34">2014</xref>), implemented the WPL layer with C&#x0002B;&#x0002B; under Caffe, and finally used C&#x0002B;&#x0002B; code for testing. An average image took 34 ms on a CPU and only 12 ms on a GPU (the time does not include semantic segmentation) <xref ref-type="fn" rid="fn0006"><sup>6</sup></xref>.</p>
</sec>
<sec>
<title>3.6. Adaptation for Multi-Illuminant</title>
<p>As mentioned in this article, the proposed method aims to solve the color constancy under a single illuminant, and we only compare our algorithm with existing single illuminant based methods. In addition, after the WPL layer, we can get the local illumination of the regions, it can estimate the multi-illuminant sources in different local regions, shown in <xref ref-type="fig" rid="F9">Figure 9</xref>, where the images are taken from the popular outdoor multi-illuminant dataset (Arjan et al., <xref ref-type="bibr" rid="B5">2012</xref>). However, there has a large deviation between the estimated illumination and the real multi illumination, we have analyzed the reasons and found that there are big errors in semantics. We will solve this problem in future research.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p>The evolution of illumination estimated. Left to right: <bold>(A)</bold> original image; <bold>(B)</bold> the estimated local illumination; <bold>(C)</bold> ground-truth illumination; <bold>(D)</bold> the corrected image by our method; <bold>(E)</bold> the ground-truth image.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-16-841426-g0009.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="conclusions" id="s4">
<title>4. Conclusion</title>
<p>In this article, we proposed a learning based multi-scale region weighed network guided by semantics (MSRWNS) to estimate the illuminated color of the light source in a scene. Cued by the human brain&#x00027;s processing of color constancy, we used image semantics and scale information to guide the process of illumination estimation. First, we put an image and its semantic mask into the network and, through a series of convolution layers, the region weights of the image at different scales were obtained. Then, through a WPL, the illumination estimation on each scale was obtained. Finally, we obtained the best estimation using the weighting on each scale, and state-of-the-art performance was achieved on three of the largest color-constancy datasets, i.e., the Color Checker, NUS 8-Camera, and ADE20k datasets. This study should prove applicable in the exploration of multi-scale and semantically directed networks for other fusion tasks in computer vision. In this study, we aim to solve the color-constancy problem with a single light source, however, there are multiple light sources in the real world, in our future research, we will try to solve the problem of multiple illuminations. In addition, it is time-consuming to obtain semantics, in our future work, we will try to use semantic information only in the training phase, not in the illumination estimation phase.</p>
</sec>
<sec sec-type="data-availability" id="s5">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author/s.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>FW is responsible for conceptualization, investigation, data curation, and writing. WW is responsible for formal analysis, investigation, and methodology. DW is responsible for formal analysis, investigation, and validation. GG is responsible for data curation and investigation. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec sec-type="funding-information" id="s7">
<title>Funding</title>
<p>This project was supported by the Science Fund of State Key Laboratory of Advanced Design and Manufacturing for Vehicle Body (No. 32015013).</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s8">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book">(<year>2003</year>). <article-title>&#x0201C;A large image database for color constancy research,&#x0201D;</article-title> in <source>Color and Imaging Conference</source> (<publisher-loc>Scottsdale, AZ</publisher-loc>).</citation>
</ref>
<ref id="B2">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Afifi</surname> <given-names>M.</given-names></name></person-group> (<year>2018</year>). <article-title>Semantic white balance: Semantic color constancy using convolutional neural network</article-title>. <source>arXiv [Preprint].</source> arXiv: 1802.00153. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/pdf/1802.00153.pdf">https://arxiv.org/pdf/1802.00153.pdf</ext-link></citation>
</ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Afifi</surname> <given-names>M.</given-names></name> <name><surname>Brown</surname> <given-names>M. S.</given-names></name></person-group> (<year>2019</year>). <article-title>&#x0201C;What else can fool deep learning? addressing color constancy errors on deep neural network performance,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF International Conference on Computer Vision</source> (<publisher-loc>Seoul</publisher-loc>), <fpage>243</fpage>&#x02013;<lpage>252</lpage>.</citation>
</ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Afifi</surname> <given-names>M.</given-names></name> <name><surname>Brown</surname> <given-names>M. S.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Deep white-balance editing,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Seattle, WA</publisher-loc>).</citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Arjan</surname> <given-names>G.</given-names></name> <name><surname>Rui</surname> <given-names>L.</given-names></name> <name><surname>Theo</surname> <given-names>G.</given-names></name></person-group> (<year>2012</year>). <article-title>Color constancy for multiple light sources</article-title>. <source>IEEE Trans. Image Process.</source> <volume>21</volume>, <fpage>697</fpage>. <pub-id pub-id-type="doi">10.1109/TIP.2011.2165219</pub-id><pub-id pub-id-type="pmid">21859624</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Barnard</surname> <given-names>K.</given-names></name></person-group> (<year>2000</year>). <article-title>&#x0201C;Improvements to gamut mapping colour constancy algorithms,&#x0201D;</article-title> in <source>Proc. European Conference on Computer Vision</source> (<publisher-loc>Dublin</publisher-loc>), <fpage>390</fpage>&#x02013;<lpage>403</lpage>.</citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barnard</surname> <given-names>K.</given-names></name> <name><surname>Martin</surname> <given-names>L.</given-names></name> <name><surname>Coath</surname> <given-names>A.</given-names></name> <name><surname>Funt</surname> <given-names>B. V.</given-names></name></person-group> (<year>2002</year>). <article-title>A comparison of computational color constancy algorithms. II. experiments with image data</article-title>. <source>IEEE Trans. Image Process.</source> <volume>11</volume>, <fpage>985</fpage>&#x02013;<lpage>996</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2002.802529</pub-id><pub-id pub-id-type="pmid">18249721</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Barnard</surname> <given-names>K.</given-names></name> <name><surname>Martin</surname> <given-names>L.</given-names></name> <name><surname>Funt</surname> <given-names>B.</given-names></name> <name><surname>Coath</surname> <given-names>A.</given-names></name></person-group> (<year>2010</year>). <article-title>A data set for color research</article-title>. <source>Color Res. Appl.</source> <volume>27</volume>, <fpage>148</fpage>&#x02013;<lpage>152</lpage>.</citation>
</ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Barron</surname> <given-names>J. T.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Convolutional color constancy,&#x0201D;</article-title> in <source>Proc. IEEE International Conference on Computer Vision</source> (<publisher-loc>Santiago</publisher-loc>), <fpage>379</fpage>&#x02013;<lpage>387</lpage>.</citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bianco</surname> <given-names>S.</given-names></name> <name><surname>Ciocca</surname> <given-names>G.</given-names></name> <name><surname>Cusano</surname> <given-names>C.</given-names></name> <name><surname>Schettini</surname> <given-names>R.</given-names></name></person-group> (<year>2008</year>). <article-title>Improving color constancy using indoor-outdoor image classification</article-title>. <source>IEEE Trans. Image Process.</source> <volume>17</volume>, <fpage>2381</fpage>&#x02013;<lpage>2392</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2008.2006661</pub-id><pub-id pub-id-type="pmid">19004710</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Bianco</surname> <given-names>S.</given-names></name> <name><surname>Cusano</surname> <given-names>C.</given-names></name> <name><surname>Schettini</surname> <given-names>R.</given-names></name></person-group> (<year>2015</year>). <article-title>Color constancy using CNNs</article-title>. <source>arXiv [Preprint].</source> arXiv: 1504. 04548. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/pdf/1504.04548.pdf">https://arxiv.org/pdf/1504.04548.pdf</ext-link></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bianco</surname> <given-names>S.</given-names></name> <name><surname>Cusano</surname> <given-names>C.</given-names></name> <name><surname>Schettini</surname> <given-names>R.</given-names></name></person-group> (<year>2017</year>). <article-title>Single and multiple illuminant estimation using convolutional neural networks</article-title>. <source>IEEE Trans. Image Process.</source> <volume>26</volume>, <fpage>4347</fpage>. <pub-id pub-id-type="doi">10.1109/TIP.2017.2713044</pub-id><pub-id pub-id-type="pmid">28600246</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brainard</surname> <given-names>D. H.</given-names></name> <name><surname>Wandell</surname> <given-names>B. A.</given-names></name></person-group> (<year>1986</year>). <article-title>Analysis of the retinex theory of color vision</article-title>. <source>J. Opt. Soc. America Opt. Image Sci.</source> <volume>3</volume>, <fpage>1651</fpage>. <pub-id pub-id-type="pmid">3772627</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Buchsbaum</surname> <given-names>G.</given-names></name></person-group> (<year>1980</year>). <article-title>A spatial processor model for object colour perception</article-title>. <source>J. Frankl. Inst.</source> <volume>310</volume>, <fpage>1</fpage>&#x02013;<lpage>26</lpage>.</citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cheng</surname> <given-names>D.</given-names></name> <name><surname>Prasad</surname> <given-names>D. K.</given-names></name> <name><surname>Brown</surname> <given-names>M. S.</given-names></name></person-group> (<year>2014</year>). <article-title>Illuminant estimation for color constancy: why spatial-domain methods work and the role of the color distribution</article-title>. <source>J. Opt. Soc. America Opt. Image Sci. Vis.</source> <volume>31</volume>, <fpage>1049</fpage>. <pub-id pub-id-type="doi">10.1364/JOSAA.31.001049</pub-id><pub-id pub-id-type="pmid">24979637</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cheng</surname> <given-names>D.</given-names></name> <name><surname>Price</surname> <given-names>B.</given-names></name> <name><surname>Cohen</surname> <given-names>S.</given-names></name> <name><surname>Brown</surname> <given-names>M. S.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Effective learning-based illuminant estimation using simple features,&#x0201D;</article-title> in <source>Proc. IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Boston, MA</publisher-loc>), <fpage>1000</fpage>&#x02013;<lpage>1008</lpage>.</citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Finlayson</surname> <given-names>G. D.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;Corrected-moment illuminant estimation,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Sydney, NSW</publisher-loc>), <fpage>1904</fpage>&#x02013;<lpage>1911</lpage>.</citation>
</ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Finlayson</surname> <given-names>G. D.</given-names></name> <name><surname>Drew</surname> <given-names>M. S.</given-names></name> <name><surname>Funt</surname> <given-names>B. V.</given-names></name></person-group> (<year>1994</year>). <article-title>Spectral sharpening: sensor transformations for improved color constancy</article-title>. <source>J. Opt. Soc. America Opt. Image Sci. Vis.</source> <volume>11</volume>, <fpage>1553</fpage>&#x02013;<lpage>63</lpage>. <pub-id pub-id-type="doi">10.1364/josaa.11.001553</pub-id><pub-id pub-id-type="pmid">8006721</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Finlayson</surname> <given-names>G. D.</given-names></name> <name><surname>Drew</surname> <given-names>M. S.</given-names></name> <name><surname>Lu</surname> <given-names>C.</given-names></name></person-group> (<year>2004</year>). <source>Intrinsic Images by Entropy Minimization</source>. <publisher-loc>Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>.</citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Finlayson</surname> <given-names>G. D.</given-names></name> <name><surname>Trezzi</surname> <given-names>E.</given-names></name></person-group> (<year>2004</year>). <article-title>&#x0201C;Shades of gray and colour constancy,&#x0201D;</article-title> in <source>Proc. Color and Imaging Conference</source> (<publisher-loc>Scottsdale, AZ</publisher-loc>), <fpage>37</fpage>&#x02013;<lpage>41</lpage>.</citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Finlayson</surname> <given-names>G. D.</given-names></name> <name><surname>Zakizadeh</surname> <given-names>R.</given-names></name> <name><surname>Gijsenij</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>The reproduction angular error for evaluating the performance of illuminant estimation algorithms</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>39</volume>, <fpage>1482</fpage>&#x02013;<lpage>1488</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2016.2582171</pub-id><pub-id pub-id-type="pmid">27333601</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Funt</surname> <given-names>B. V.</given-names></name> <name><surname>Lewis</surname> <given-names>B. C.</given-names></name></person-group> (<year>2000</year>). <article-title>Diagonal versus affine transformations for color correction</article-title>. <source>J. Opt. Soc. America Opt. Image Sci. Vis.</source> <volume>17</volume>, <fpage>2108</fpage>&#x02013;<lpage>2112</lpage>. <pub-id pub-id-type="doi">10.1364/josaa.17.002108</pub-id><pub-id pub-id-type="pmid">11059611</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>S.</given-names></name> <name><surname>Yang</surname> <given-names>K.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;A color constancy model with double-opponency mechanisms,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Sydney, NSW</publisher-loc>), <fpage>929</fpage>&#x02013;<lpage>936</lpage>.</citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>S.-B.</given-names></name> <name><surname>Ren</surname> <given-names>Y.-Z.</given-names></name> <name><surname>Zhang</surname> <given-names>M.</given-names></name> <name><surname>Li</surname> <given-names>Y.-J.</given-names></name></person-group> (<year>2019</year>). <article-title>Combining bottom-up and top-down visual mechanisms for color constancy under varying illumination</article-title>. <source>IEEE Trans. Image Process.</source> <volume>28</volume>, <fpage>4387</fpage>&#x02013;<lpage>4400</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2019.2908783</pub-id><pub-id pub-id-type="pmid">30946665</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>S. B.</given-names></name> <name><surname>Yang</surname> <given-names>K. F.</given-names></name> <name><surname>Li</surname> <given-names>C. Y.</given-names></name> <name><surname>Li</surname> <given-names>Y. J.</given-names></name></person-group> (<year>2015</year>). <article-title>Color constancy using double-opponency</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>37</volume>, <fpage>1973</fpage>&#x02013;<lpage>1985</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2015.2396053</pub-id><pub-id pub-id-type="pmid">26353182</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gehler</surname> <given-names>P. V.</given-names></name> <name><surname>Rother</surname> <given-names>C.</given-names></name> <name><surname>Blake</surname> <given-names>A.</given-names></name> <name><surname>Minka</surname> <given-names>T.</given-names></name> <name><surname>Sharp</surname> <given-names>T.</given-names></name></person-group> (<year>2008</year>). <article-title>&#x0201C;Bayesian color constancy revisited,&#x0201D;</article-title> in <source>Proc. IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Anchorage, AK</publisher-loc>), <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gijsenij</surname> <given-names>A.</given-names></name> <name><surname>Gevers</surname> <given-names>T.</given-names></name> <name><surname>Lucassen</surname> <given-names>M. P.</given-names></name></person-group> (<year>2009</year>). <article-title>Perceptual analysis of distance measures for color constancy algorithms</article-title>. <source>J. Opt. Soc. America Opt. Image Sci. Vis.</source> <volume>26</volume>, <fpage>2243</fpage>. <pub-id pub-id-type="doi">10.1364/JOSAA.26.002243</pub-id><pub-id pub-id-type="pmid">19798406</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gijsenij</surname> <given-names>A.</given-names></name> <name><surname>Gevers</surname> <given-names>T.</given-names></name> <name><surname>Van</surname> <given-names>d. W. J.</given-names></name></person-group> (<year>2011</year>). <article-title>Computational color constancy: survey and experiments</article-title>. <source>IEEE Trans. Image Process.</source> <volume>20</volume>, <fpage>2475</fpage>&#x02013;<lpage>2489</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2011.2118224</pub-id><pub-id pub-id-type="pmid">21342844</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gilchrist</surname> <given-names>A.</given-names></name></person-group> (<year>2006</year>). <source>Seeing Black and White</source> <publisher-name>Oxford University Press</publisher-name>.</citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hirakawa</surname> <given-names>K.</given-names></name> <name><surname>Chakrabarti</surname> <given-names>A.</given-names></name> <name><surname>Zickler</surname> <given-names>T.</given-names></name></person-group> (<year>2012</year>). <article-title>Color constancy with spatio-spectral statistics</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>34</volume>, <fpage>1509</fpage>&#x02013;<lpage>1519</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2011.252</pub-id><pub-id pub-id-type="pmid">22745000</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hordley</surname> <given-names>S. D.</given-names></name> <name><surname>Finlayson</surname> <given-names>G. D.</given-names></name></person-group> (<year>2004</year>). <article-title>&#x0201C;Re-evaluating colour constancy algorithms,&#x0201D;</article-title> in <source>Proceedings of the 17th International Conference on Pattern Recognition</source>, Vol. 1 (IEEE), <fpage>76</fpage>&#x02013;<lpage>79</lpage>.</citation>
</ref>
<ref id="B32">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hu</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>B.</given-names></name> <name><surname>Lin</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Fc4: fully convolutional color constancy with confidence-weighted pooling,&#x0201D;</article-title> in <source>IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Honolulu, HI</publisher-loc>), <fpage>330</fpage>&#x02013;<lpage>339</lpage>.</citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Iandola</surname> <given-names>F. N.</given-names></name> <name><surname>Han</surname> <given-names>S.</given-names></name> <name><surname>Moskewicz</surname> <given-names>M. W.</given-names></name> <name><surname>Ashraf</surname> <given-names>K.</given-names></name> <name><surname>Dally</surname> <given-names>W. J.</given-names></name> <name><surname>Keutzer</surname> <given-names>K.</given-names></name></person-group> (<year>2016</year>). <article-title>Squeezenet: alexnet-level accuracy with 50x fewer parameters and &#x0003C;0.5 mb model size</article-title>. <source>arXiv preprint</source> arXiv:1602.07360.</citation>
</ref>
<ref id="B34">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Jia</surname> <given-names>Y.</given-names></name> <name><surname>Shelhamer</surname> <given-names>E.</given-names></name> <name><surname>Donahue</surname> <given-names>J.</given-names></name> <name><surname>Karayev</surname> <given-names>S.</given-names></name> <name><surname>Long</surname> <given-names>J.</given-names></name> <name><surname>Girshick</surname> <given-names>R.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>&#x0201C;Caffe: convolutional architecture for fast feature embedding,&#x0201D;</article-title> in <source>Proceedings of the 22nd ACM International Conference on Multimedia</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>675</fpage>&#x02013;<lpage>678</lpage>. <pub-id pub-id-type="pmid">32210685</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Joze</surname> <given-names>H. R. V.</given-names></name> <name><surname>Drew</surname> <given-names>M. S.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;White patch gamut mapping colour constancy,&#x0201D;</article-title> in <source>Proc. IEEE International Conference on Image Processing</source> (<publisher-loc>Orlando, FL</publisher-loc>), <fpage>801</fpage>&#x02013;<lpage>804</lpage>.</citation>
</ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Joze</surname> <given-names>H. R. V.</given-names></name> <name><surname>Drew</surname> <given-names>M. S.</given-names></name> <name><surname>Finlayson</surname> <given-names>G. D.</given-names></name> <name><surname>Rey</surname> <given-names>P. A. T.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;The role of bright pixels in 416 illumination estimation,&#x0201D;</article-title> in <source>Color and Imaging Conference, Vol. 2012</source> (<publisher-loc>Society for Imaging Science and Technology</publisher-loc>), <fpage>41</fpage>&#x02013;<lpage>46</lpage>.</citation>
</ref>
<ref id="B37">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D. P.</given-names></name> <name><surname>Ba</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>Adam: A method for stochastic optimization</article-title>. <source>arXiv [Preprint]</source>. arXiv:1412.6980. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/pdf/1412.6980.pdf">https://arxiv.org/pdf/1412.6980.pdf</ext-link></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krasilnikov</surname> <given-names>N. N.</given-names></name> <name><surname>Krasilnikova</surname> <given-names>O. I.</given-names></name> <name><surname>Shelepin</surname> <given-names>Y. E.</given-names></name></person-group> (<year>2002</year>). <article-title>Mathematical model of the color constancy of the human visual system</article-title>. <source>J. Opt. Technol. C Opticheskii Zhurnal</source> <volume>69</volume>, <fpage>102</fpage>&#x02013;<lpage>107</lpage>. <pub-id pub-id-type="doi">10.1364/JOT.69.000327</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krizhevsky</surname> <given-names>A.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Imagenet classification with deep convolutional neural networks,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems</source>, <fpage>25</fpage>.</citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Land</surname> <given-names>E. H.</given-names></name></person-group> (<year>1977</year>). <article-title>The retinex theory of color vision</article-title>. <source>Sci. Am.</source> <volume>237</volume>, <fpage>108</fpage>.</citation>
</ref>
<ref id="B41">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lau</surname> <given-names>H. Y.</given-names></name></person-group> (<year>2008</year>). <source>Neural inspired color constancy model based on double opponent neurons.</source> <publisher-name>Hong Kong University of Science and Technology</publisher-name>, <publisher-loc>Hong Kong, China</publisher-loc>.</citation>
</ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lee</surname> <given-names>H. C.</given-names></name></person-group> (<year>1986</year>). <article-title>Method for computing the scene-illuminant chromaticity from specular highlights</article-title>. <source>J. Opt. Soc. America Opt. Image Sci. Vis.</source> <volume>3</volume>, <fpage>1694</fpage>&#x02013;<lpage>1699</lpage>. <pub-id pub-id-type="pmid">3772631</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>B.</given-names></name> <name><surname>Xu</surname> <given-names>D.</given-names></name> <name><surname>Wang</surname> <given-names>J. H.</given-names></name> <name><surname>Lu</surname> <given-names>R.</given-names></name></person-group> (<year>2008</year>). <article-title>&#x0201C;Color constancy based on image similarity,&#x0201D;</article-title> in <source>Transactions on Information and Systems E91-D</source>, <fpage>375</fpage>&#x02013;<lpage>378</lpage>.</citation>
</ref>
<ref id="B44">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Nayar</surname> <given-names>S. K.</given-names></name> <name><surname>Krishnan</surname> <given-names>G.</given-names></name> <name><surname>Grossberg</surname> <given-names>M. D.</given-names></name> <name><surname>Raskar</surname> <given-names>R.</given-names></name></person-group> (<year>2007</year>). <source>Method for separating direct and global illumination in a scene. US Patent App. 11/624,016.</source> <publisher-loc>Washington, DC</publisher-loc>: <publisher-name>U.S. Patent and Trademark Office</publisher-name>.</citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nieves</surname> <given-names>J. L.</given-names></name> <name><surname>Garc&#x000ED;a-Beltr&#x000E1;n</surname> <given-names>A.</given-names></name> <name><surname>Romero</surname> <given-names>J.</given-names></name></person-group> (<year>2000</year>). <article-title>Response of the human visual system to variable illuminant conditions: an analysis of opponent-colour mechanisms in colour constancy</article-title>. <source>Ophthalmic Physiol. Opt. J. Brit. Coll. Ophthalmic Opticians</source> <volume>20</volume>, <fpage>44</fpage>. <pub-id pub-id-type="doi">10.1046/j.1475-1313.2000.00471.x</pub-id><pub-id pub-id-type="pmid">10884929</pub-id></citation></ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rampasek</surname> <given-names>L.</given-names></name> <name><surname>Goldenberg</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>Tensorflow: biology&#x00027;s gateway to deep learning?</article-title> <source>Cell Syst.</source> <volume>2</volume>, <fpage>12</fpage>&#x02013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1016/j.cels.2016.01.009</pub-id><pub-id pub-id-type="pmid">27136685</pub-id></citation></ref>
<ref id="B47">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Schroeder</surname> <given-names>M.</given-names></name> <name><surname>Moser</surname> <given-names>S.</given-names></name></person-group> (<year>2001</year>). <article-title>&#x0201C;Automatic color correction based on generic content based image analysis,&#x0201D;</article-title> in <source>Color and Imaging Conference</source> (<publisher-loc>Scottsdale, AZ</publisher-loc>), <fpage>41</fpage>&#x02013;<lpage>45</lpage>.</citation>
</ref>
<ref id="B48">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Shi</surname> <given-names>W.</given-names></name> <name><surname>Chen</surname> <given-names>C. L.</given-names></name> <name><surname>Tang</surname> <given-names>X.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Deep specialized network for illuminant estimation,&#x0201D;</article-title> in <source>Proc. European Conference on Computer Vision</source> (<publisher-loc>Amsterdam</publisher-loc>), <fpage>371</fpage>&#x02013;<lpage>387</lpage>.</citation>
</ref>
<ref id="B49">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Simonyan</surname> <given-names>K.</given-names></name> <name><surname>Zisserman</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Very deep convolutional networks for large-scale image recognition</article-title>. <source>arXiv [preprint].</source> arXiv:1409.1556. Available online at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/pdf/1409.1556.pdf">https://arxiv.org/pdf/1409.1556.pdf</ext-link></citation>
</ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Spitzer</surname> <given-names>H.</given-names></name> <name><surname>Semo</surname> <given-names>S.</given-names></name></person-group> (<year>2002</year>). <article-title>Color constancy: a biological model and its application for still and video images</article-title>. <source>Pattern Recognit.</source> <volume>35</volume>, <fpage>1645</fpage>&#x02013;<lpage>1659</lpage>. <pub-id pub-id-type="doi">10.1016/S0031-3203(01)00160-1</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tan</surname> <given-names>R. T.</given-names></name> <name><surname>Nishino</surname> <given-names>K.</given-names></name> <name><surname>Ikeuchi</surname> <given-names>K.</given-names></name></person-group> (<year>2008</year>). <source>Color Constancy Through Inverse-Intensity Chromaticity Space.</source> <publisher-loc>Boston, MA</publisher-loc>: <publisher-name>Springer US</publisher-name>. <pub-id pub-id-type="pmid">15005396</pub-id></citation></ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Toro</surname> <given-names>J.</given-names></name></person-group> (<year>2008</year>). <article-title>Dichromatic illumination estimation without pre-segmentation</article-title>. <source>Pattern Recognit. Lett.</source> <volume>29</volume>, <fpage>871</fpage>&#x02013;<lpage>877</lpage>. <pub-id pub-id-type="doi">10.1016/j.patrec.2008.01.004</pub-id></citation>
</ref>
<ref id="B53">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Van De Weijer</surname> <given-names>J.</given-names></name> <name><surname>Schmid</surname> <given-names>C.</given-names></name> <name><surname>Verbeek</surname> <given-names>J.</given-names></name></person-group> (<year>2007</year>). <article-title>&#x0201C;Using high-level visual information for color constancy,&#x0201D;</article-title> in <source>2007 IEEE 11th International Conference on Computer Vision</source> (<publisher-loc>Rio de Janeiro</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wandell</surname> <given-names>B. A.</given-names></name> <name><surname>Tominaga</surname> <given-names>S.</given-names></name></person-group> (<year>1989</year>). <article-title>Standard surface-reflectance model and illuminant estimation</article-title>. <source>J. Opt. Soc. America A</source> <volume>6</volume>, <fpage>576</fpage>&#x02013;<lpage>584</lpage>.</citation>
</ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Weijer</surname> <given-names>J. V. D.</given-names></name> <name><surname>Gevers</surname> <given-names>T.</given-names></name> <name><surname>Gijsenij</surname> <given-names>A.</given-names></name></person-group> (<year>2007</year>). <article-title>Edge-based color constancy</article-title>. <source>IEEE Trans. Image Process.</source> <volume>16</volume>, <fpage>2207</fpage>&#x02013;<lpage>2214</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2007.901808</pub-id><pub-id pub-id-type="pmid">17784594</pub-id></citation></ref>
<ref id="B56">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xiao</surname> <given-names>J.</given-names></name> <name><surname>Gu</surname> <given-names>S.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Multi-domain learning for accurate and few-shot color constancy,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Seattle, WA</publisher-loc>), <fpage>3258</fpage>&#x02013;<lpage>3267</lpage>.</citation>
</ref>
<ref id="B57">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xue</surname> <given-names>S.</given-names></name> <name><surname>Gao</surname> <given-names>S.</given-names></name> <name><surname>Tan</surname> <given-names>M.</given-names></name> <name><surname>He</surname> <given-names>Z.</given-names></name> <name><surname>He</surname> <given-names>L.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;How does color constancy affect target recognition and instance segmentation?&#x0201D;</article-title> in <source>Proceedings of the 29th ACM International Conference on Multimedia</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>5537</fpage>&#x02013;<lpage>5545</lpage>.</citation>
</ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>H.</given-names></name> <name><surname>Chen</surname> <given-names>K.</given-names></name> <name><surname>Wang</surname> <given-names>K.</given-names></name> <name><surname>Qian</surname> <given-names>Y.</given-names></name> <name><surname>Jia</surname> <given-names>K.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Cascading convolutional color constancy,&#x0201D;</article-title> in <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>, vol. 34 (New York, NY), <fpage>12725</fpage>&#x02013;<lpage>12732</lpage>.</citation>
</ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>H.</given-names></name> <name><surname>Shi</surname> <given-names>J.</given-names></name> <name><surname>Qi</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Jia</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Pyramid scene parsing network,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>, <fpage>2881</fpage>&#x02013;<lpage>2890</lpage>. <pub-id pub-id-type="pmid">33390119</pub-id></citation></ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>B.</given-names></name> <name><surname>Hang</surname> <given-names>Z.</given-names></name> <name><surname>Fernandez</surname> <given-names>F. X. P.</given-names></name> <name><surname>Fidler</surname> <given-names>S.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Scene parsing through ade20k dataset,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>, <fpage>633</fpage>&#x02013;<lpage>641</lpage>.</citation>
</ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>B.</given-names></name> <name><surname>Zhao</surname> <given-names>H.</given-names></name> <name><surname>Puig</surname> <given-names>X.</given-names></name> <name><surname>Fidler</surname> <given-names>S.</given-names></name> <name><surname>Barriuso</surname> <given-names>A.</given-names></name> <name><surname>Torralba</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>Semantic understanding of scenes through the ade20k dataset</article-title>.</citation>
</ref>
</ref-list>
<fn-group>
<fn id="fn0001"><p><sup>1</sup>In this study, we aim to solve the color-constancy problem with a single light source.</p></fn>
<fn id="fn0002"><p><sup>2</sup>As reported in Finlayson et al. (<xref ref-type="bibr" rid="B19">2004</xref>); Barron (<xref ref-type="bibr" rid="B9">2015</xref>), <italic>log</italic> &#x02212; <italic>uv</italic> is better than RGB, since, first, there are two variables instead of three; and second, the multiplicative constraint of the illumination estimation model is converted to the linear constraint.</p></fn>
<fn id="fn0003"><p><sup>3</sup>In this work, the average value of multiple-scale illumination is used.</p></fn>
<fn id="fn0004"><p><sup>4</sup>We did not train the scene parsing network ourselves, the model used is downloaded from <ext-link ext-link-type="uri" xlink:href="https://github.com/hszhao/PSPNet">https://github.com/hszhao/PSPNet</ext-link>.</p></fn>
<fn id="fn0005"><p><sup>5</sup>The sensor of this camera does not have a low-pass filter, and the color filter comprises several Bayer filters that overlap each other.</p></fn>
<fn id="fn0006"><p><sup>6</sup>experimental hardware platform: i7 7700k, 32 GB memory, gtx1080ti. If the test-image resolution was greater than 512 &#x000D7; 512, it was resized to 512 &#x000D7; 512</p></fn>
</fn-group>
</back>
</article>