<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="systematic-review">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Bioeng. Biotechnol.</journal-id>
<journal-title>Frontiers in Bioengineering and Biotechnology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Bioeng. Biotechnol.</abbrev-journal-title>
<issn pub-type="epub">2296-4185</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fbioe.2018.00033</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Bioengineering and Biotechnology</subject>
<subj-group>
<subject>Systematic Review</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A Comparative Survey of Methods for Remote Heart Rate Detection From Frontal Face Videos</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Wang</surname> <given-names>Chen</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="fn001">&#x0002A;</xref>
<uri xlink:href="https://frontiersin.org/people/u/415814"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Pun</surname> <given-names>Thierry</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="https://frontiersin.org/people/u/177996"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Chanel</surname> <given-names>Guillaume</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="https://frontiersin.org/people/u/182475"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Computer Vision and Multimedia Laboratory, Computer Science Department, University of Geneva</institution>, <addr-line>Geneva</addr-line>, <country>Switzerland</country></aff>
<aff id="aff2"><sup>2</sup><institution>Swiss Center for Affective Sciences, Campus Biotech, University of Geneva</institution>, <addr-line>Geneva</addr-line>, <country>Switzerland</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Danilo Emilio De Rossi, Universit&#x000E0; degli Studi di Pisa, Italy</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Gholamreza Anbarjafari, University of Tartu, Estonia; Andrea Bonarini, Politecnico di Milano, Italy</p></fn>
<corresp id="fn001">&#x0002A;Correspondence: Chen Wang, <email>chen.wang&#x00040;unige.ch</email></corresp>
<fn fn-type="other" id="fn002"><p>Specialty section: This article was submitted to Bionics and Biomimetics, a section of the journal Frontiers in Bioengineering and Biotechnology</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>01</day>
<month>05</month>
<year>2018</year>
</pub-date>
<pub-date pub-type="collection">
<year>2018</year>
</pub-date>
<volume>6</volume>
<elocation-id>33</elocation-id>
<history>
<date date-type="received">
<day>21</day>
<month>07</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>13</day>
<month>03</month>
<year>2018</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2018 Wang, Pun and Chanel.</copyright-statement>
<copyright-year>2018</copyright-year>
<copyright-holder>Wang, Pun and Chanel</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Remotely measuring physiological activity can provide substantial benefits for both the medical and the affective computing applications. Recent research has proposed different methodologies for the unobtrusive detection of heart rate (HR) using human face recordings. These methods are based on subtle color changes or motions of the face due to cardiovascular activities, which are invisible to human eyes but can be captured by digital cameras. Several approaches have been proposed such as signal processing and machine learning. However, these methods are compared with different datasets, and there is consequently no consensus on method performance. In this article, we describe and evaluate several methods defined in literature, from 2008 until present day, for the remote detection of HR using human face recordings. The general HR processing pipeline is divided into three stages: face video processing, face blood volume pulse (BVP) signal extraction, and HR computation. Approaches presented in the paper are classified and grouped according to each stage. At each stage, algorithms are analyzed and compared based on their performance using the public database MAHNOB-HCI. Results found in this article are limited on MAHNOB-HCI dataset. Results show that extracted face skin area contains more BVP information. Blind source separation and peak detection methods are more robust with head motions for estimating HR.</p>
</abstract>
<kwd-group>
<kwd>heart rate</kwd>
<kwd>remote sensing</kwd>
<kwd>physiological signals</kwd>
<kwd>photoplethysmography</kwd>
<kwd>human&#x02013;computer interaction</kwd>
</kwd-group>
<contract-num rid="cn01">200021E - 164326</contract-num>
<contract-sponsor id="cn01">Schweizerischer Nationalfonds zur F&#x000F6;rderung der Wissenschaftlichen Forschung<named-content content-type="fundref-id">10.13039/501100001711</named-content></contract-sponsor>
<counts>
<fig-count count="7"/>
<table-count count="3"/>
<equation-count count="0"/>
<ref-count count="73"/>
<page-count count="16"/>
<word-count count="11936"/>
</counts>
</article-meta>
</front>
<body>
<sec id="S1" sec-type="introduction">
<title>Introduction</title>
<p>Heart rate (HR) is a measure of physiological activity and it can indicate a person&#x02019;s health and affective status (Malik, <xref ref-type="bibr" rid="B31">1996</xref>; Armony and Vuilleumier, <xref ref-type="bibr" rid="B3">2013</xref>). Physical exercise, mental stress, and medicines all influence on cardiac activities. Consequently, HR information can be used in a wide range of applications, such as medical diagnosis, fitness assessment, and emotion recognition. Traditional methods of measuring HR rely on electronic or optical sensors. The majority of these methods require skin-contact, such as electrocardiograms (ECGs), sphygmomanometry and pulse oximetry, and the later giving a photoplethysmogram (PPG). Among all cardiac pulse measurements, the current gold standard is the usage of ECG (Dawson et al., <xref ref-type="bibr" rid="B13">2010</xref>), which places adhesive gel electrodes on the participants&#x02019; limbs or chest surface. Another, widely applied contact method, is to compute the blood volume pulse (BVP) from a PPG captured by an oximeter emitting and measuring light at proper wavelengths (Allen, <xref ref-type="bibr" rid="B2">2007</xref>). However, the skin-contact measurements can be considered as inconvenient, unpractical, and may cause uncomfortable feelings.</p>
<p>In the past decade, researchers have focused on remote (i.e., contactless) detection methods, which are mainly based on computer vision techniques. Using human faces as physiological measurement resources was first proposed in 2007 (Pavlidis et al., <xref ref-type="bibr" rid="B38">2007</xref>). According to Pavlidis et al. (<xref ref-type="bibr" rid="B38">2007</xref>), the face area facilitated observation as it featured a thin layer of tissue. With facial thermal imaging, HR can be detected based on bioheat models (Garbey et al., <xref ref-type="bibr" rid="B16">2007</xref>; Pavlidis et al., <xref ref-type="bibr" rid="B38">2007</xref>). After that, the PPG technique, which is non-invasive and optical, was used for detecting HR. The method is often implemented with dedicated light sources such as red lights or infrared lights (Allen, <xref ref-type="bibr" rid="B2">2007</xref>; Jeanne et al., <xref ref-type="bibr" rid="B21">2013</xref>).</p>
<p>In 2008, Verkruysse et al. (<xref ref-type="bibr" rid="B58">2008</xref>) showed the possibility of using PPG under ambient light to estimate HR from videos of human face. Then in 2010, Poh et al. (<xref ref-type="bibr" rid="B39">2010</xref>) developed a framework for automatic HR detection using the color of human face recordings obtained from a standard camera. This framework was widely adopted and modified in Poh et al. (<xref ref-type="bibr" rid="B40">2011</xref>), Pursche et al. (<xref ref-type="bibr" rid="B41">2012</xref>), and Kwon et al. (<xref ref-type="bibr" rid="B25">2012</xref>). For all those methods, the core idea is to recover the heartbeat signal using blind source separation (BSS) on the temporal changes of face color. Later in 2013, another method for estimating HR based on subtle head motions (Balakrishnan et al., <xref ref-type="bibr" rid="B4">2013</xref>, Rubinstein, <xref ref-type="bibr" rid="B44">2013</xref>) was proposed. Besides, researchers (Li et al., <xref ref-type="bibr" rid="B29">2014</xref>; Stricker et al., <xref ref-type="bibr" rid="B50">2014</xref>; Xu et al., <xref ref-type="bibr" rid="B67">2014</xref>) investigated the estimation of HR directly by applying diverging noise reduction algorithms and optical modeling methods. Alternatively, the usage of manifold learning methods mapping multi-dimensional face video data into one-dimensional space has been studied to reveal the HR signal as well.</p>
<p>As shown above, remote HR detection has been an active field of research for the past decade and produced different strategies using diverse processing methods and models. However, many implementations were evaluated on different datasets and it is consequently difficult to compare them. Furthermore, no survey paper has been conducted, with the objective of gathering, classifying, and analyzing the existing work within this domain. The objectives of this article are first to fill this gap by presenting a general pipeline composed of several steps and how the different state-of-the-art methods can be classified based on the pipeline (presenting in Section &#x0201C;<xref ref-type="sec" rid="S2">Remote Methods for HR Detection</xref>&#x0201D;). The second objective is to evaluate the mainstream methods at each step of the pipeline to finally obtain a full implementation with the best performance (presenting in Section &#x0201C;<xref ref-type="sec" rid="S3">Comparative Analysis</xref>&#x0201D;). This objective is achieved by testing the methods on a unique set of data: the MAHNOB-HCI database. Given the methods&#x02019; popularity, this analysis is limited to color intensity-based methods.</p>
</sec>
<sec id="S2" sec-type="methods">
<title>Remote Methods for HR Detection</title>
<p>To the best of our knowledge, we are unaware of existing reviews that touch upon this topic. To access the method performance, this article investigates several methods, which were published in international conferences and journals from 2008 until 2017. This time period was selected because 2008 was the year when the remote HR detection was first proposed. Methods that require no skin-contact and no specific light sources were exclusively taken into account because they are more likely to be applied outside the laboratory.</p>
<p>The existing remote methods for obtaining HR from face videos can be classified as either color intensity-based methods or motion-based methods. Currently, the intensity-based methods are the most popular (Poh et al., <xref ref-type="bibr" rid="B40">2011</xref>; Kwon et al., <xref ref-type="bibr" rid="B25">2012</xref>; Pursche et al., <xref ref-type="bibr" rid="B41">2012</xref>; etc.) shown in Table <xref ref-type="table" rid="T1">1</xref>. Intensity-based methods come from PPG signals captured by digital cameras. Blood absorbs light more than the surrounding tissues and variations in blood volume affect light transmission and reflectance (Verkruysse et al., <xref ref-type="bibr" rid="B58">2008</xref>). That leads to the subtle color changes on human skin, which is invisible to human eyes but recorded by cameras. Diverse optical models are applied to extract the intensity of color changes caused by pulse. As shown in Figure <xref ref-type="fig" rid="F1">1</xref>, hemoglobin and oxyhemoglobin both have high ability of absorption in the green color range and low in the red color range. But all three color channels contain PPG information (Verkruysse et al., <xref ref-type="bibr" rid="B58">2008</xref>). More detailed information on PPG-based methods can be found in the work of Allen (<xref ref-type="bibr" rid="B2">2007</xref>) and Sun et al. (<xref ref-type="bibr" rid="B52">2012</xref>).</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Classification of state-of-the-art methods.</p></caption>
<table frame="hsides" rules="groups">
<tbody>
<tr>
<td align="left" valign="top" rowspan="5">Dimensionality reduction</td>
<td align="left" valign="top" rowspan="3">Blind source separation</td>
<td align="left" valign="top">Independent component analysis</td>
<td align="left" valign="top">Poh et al. (<xref ref-type="bibr" rid="B39">2010</xref>), Poh et al. (<xref ref-type="bibr" rid="B40">2011</xref>), Pursche et al. (<xref ref-type="bibr" rid="B41">2012</xref>), Kwon et al. (<xref ref-type="bibr" rid="B25">2012</xref>), Lewandowska et al. (<xref ref-type="bibr" rid="B28">2011</xref>), Sahindrakar et al. (<xref ref-type="bibr" rid="B45">2011</xref>), Datcu et al. (<xref ref-type="bibr" rid="B12">2013</xref>), Jensen and Hannemose (<xref ref-type="bibr" rid="B22">2014</xref>), Yu et al. (<xref ref-type="bibr" rid="B71">2015</xref>), Lam and Yoshinori (<xref ref-type="bibr" rid="B27">2015</xref>), Kumar et al. (<xref ref-type="bibr" rid="B24">2015</xref>), and McDuff et al. (<xref ref-type="bibr" rid="B32">2017</xref>)</td>
</tr>
<tr>
<td align="left" valign="top" colspan="2"><hr/></td>
</tr>
<tr>
<td align="left" valign="top">Principle component analysis</td>
<td align="left" valign="top">Lewandowska et al. (<xref ref-type="bibr" rid="B28">2011</xref>), Wei et al. (<xref ref-type="bibr" rid="B63">2012</xref>), Rubinstein (<xref ref-type="bibr" rid="B44">2013</xref>), Irani et al. (<xref ref-type="bibr" rid="B20">2014</xref>), Balakrishnan et al. (<xref ref-type="bibr" rid="B4">2013</xref>), and Chen et al. (<xref ref-type="bibr" rid="B10">2017</xref>)</td>
</tr>
<tr>
<td align="left" valign="top" colspan="3"><hr/></td>
</tr>
<tr>
<td align="left" valign="top" colspan="2">Other dimensionality methods</td>
<td align="left" valign="top">Wei et al. (<xref ref-type="bibr" rid="B63">2012</xref>), Rubinstein (<xref ref-type="bibr" rid="B44">2013</xref>), and Tran et al. (<xref ref-type="bibr" rid="B56">2015</xref>)</td>
</tr>
<tr>
<td align="left" valign="top" colspan="4"><hr/></td>
</tr>
<tr>
<td align="left" valign="top" rowspan="3">Optical modeling</td>
<td align="left" valign="top" colspan="2">Green channel</td>
<td align="left" valign="top">Verkruysse et al. (<xref ref-type="bibr" rid="B58">2008</xref>), Pursche et al. (<xref ref-type="bibr" rid="B41">2012</xref>), Stricker et al. (<xref ref-type="bibr" rid="B50">2014</xref>), Li et al. (<xref ref-type="bibr" rid="B29">2014</xref>), Zaunseder et al. (<xref ref-type="bibr" rid="B72">2014</xref>), Muender et al. (<xref ref-type="bibr" rid="B36">2016</xref>), Mestha et al. (<xref ref-type="bibr" rid="B33">2014</xref>), Kumar et al. (<xref ref-type="bibr" rid="B24">2015</xref>), and Moreno et al. (<xref ref-type="bibr" rid="B35">2015</xref>)</td>
</tr>
<tr>
<td align="left" valign="top" colspan="3"><hr/></td>
</tr>
<tr>
<td align="left" valign="top" colspan="2">Other optical modeling methods</td>
<td align="left" valign="top">Pursche et al. (<xref ref-type="bibr" rid="B41">2012</xref>), Stricker et al. (<xref ref-type="bibr" rid="B50">2014</xref>), Li et al. (<xref ref-type="bibr" rid="B29">2014</xref>), Zaunseder et al. (<xref ref-type="bibr" rid="B72">2014</xref>), Muender et al. (<xref ref-type="bibr" rid="B36">2016</xref>), Mestha et al. (<xref ref-type="bibr" rid="B33">2014</xref>), Kumar et al. (<xref ref-type="bibr" rid="B24">2015</xref>), and Moreno et al. (<xref ref-type="bibr" rid="B35">2015</xref>)</td>
</tr>
<tr>
<td align="left" valign="top" colspan="4"><hr/></td>
</tr>
<tr>
<td align="left" valign="top" colspan="3">Motion-based methods</td>
<td align="left" valign="top">Balakrishnan et al. (<xref ref-type="bibr" rid="B4">2013</xref>), Rubinstein (<xref ref-type="bibr" rid="B44">2013</xref>), and Irani et al. (<xref ref-type="bibr" rid="B20">2014</xref>)</td>
</tr>
<tr>
<td align="left" valign="top" colspan="4"><hr/></td>
</tr>
<tr>
<td align="left" valign="top" colspan="3">Machine learning</td>
<td align="left" valign="top">Monkaresi et al. (<xref ref-type="bibr" rid="B34">2014</xref>), Tarassenko et al. (<xref ref-type="bibr" rid="B53">2014</xref>), Osman et al. (<xref ref-type="bibr" rid="B37">2015</xref>), and Villarroel et al. (<xref ref-type="bibr" rid="B60">2017</xref>)</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Hemoglobin (green) and oxyhemoglobin (blue) absorption spectra (Jensen and Hannemose, <xref ref-type="bibr" rid="B22">2014</xref>).</p></caption>
<graphic xlink:href="fbioe-06-00033-g001.tif"/>
</fig>
<p>Head motions caused by pulse are mixed together with other involuntary and voluntary head movements. Subtle upright head motions in the vertical direction are mainly caused by pulse activities, while the bobbing movements are caused by respiration (Da et al., <xref ref-type="bibr" rid="B11">2011</xref>; Balakrishnan et al., <xref ref-type="bibr" rid="B4">2013</xref>). Motion-based methods for detecting HR stemmed from ballistocardiogram (Starr et al., <xref ref-type="bibr" rid="B49">1939</xref>). Ballistocardiographic head movement is obtained by lying a participant on a low-friction platform from which displacements are measured to get cardiac information. In Da et al. (<xref ref-type="bibr" rid="B11">2011</xref>), head motion was measured by accelerometers to monitor HR. Balakrishnan et al. (<xref ref-type="bibr" rid="B4">2013</xref>) proposed to detect HR remotely from face videos through head motions. The basic approach consists of tracking features from a person&#x02019;s head, filtering out the velocity of interest, and then extracting the periodic signal caused by heartbeats.</p>
<p>Both subtle color changes and head motions can be easily &#x0201C;hidden&#x0201D; by other signals during recording. The accuracy of HR estimation is influenced by the participants&#x02019; movements, complex facial features (face shape, hair, glasses, beards, etc.), facial expressions, camera noise and distortion, and changing light conditions. Many papers in this field use strictly controlled experiment settings to eliminate the influential factors. Besides well-controlled conditions, algorithms for noise reduction and signal recovery are applied to retrieve HR information. For intensity-based methods, averaging the pixel values inside a region of interest (ROI) is often applied to overcome sensor and quantization noise. Subsequently, temporal filters are adopted to extract the signal of interest (Poh et al., <xref ref-type="bibr" rid="B39">2010</xref>; Wu et al., <xref ref-type="bibr" rid="B66">2012</xref>). As for motion-based approaches, similar algorithms are used such as face tracking and noise reduction.</p>
<p>To categorize existing methods, we divide the HR detection procedure into three stages based on the implementation sequence: face video processing, face BVP signal extraction, and HR computation (Figure <xref ref-type="fig" rid="F2">2</xref>). Face video processing aims to detect faces, improve the motion robustness, reduce quantization errors, and prepare the featured signals for further BVP signal extraction. There are more algorithm variations at this stage than at BVP signal extraction and HR computation. For BVP signal extraction, temporal filtering, component analysis, and other approaches are used to recover HR information from noisy signals. The HR computation stage aims to compute HR from the cardiac signal obtained from the previous stage. At this stage, the methods can be grouped into time domain analysis and frequency domain analysis. For the time domain processing, peak detection is diffusely applied to get the inter-beat interval (IBI) from which HR is computed. In frequency domain, the power spectral density is mostly used, where the dominant frequency is taken as HR. HR computation can become complex for applications including buffer handling functions to present HR results after a certain time period (Stricker et al., <xref ref-type="bibr" rid="B50">2014</xref>).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>General schematic diagram for remote heart rate (HR) detection from face videos.</p></caption>
<graphic xlink:href="fbioe-06-00033-g002.tif"/>
</fig>
<sec id="S2-1">
<title>Experiment Setting</title>
<p>Only a few papers used public datasets for remote HR estimation from face videos (Li et al., <xref ref-type="bibr" rid="B29">2014</xref>; Werner et al., <xref ref-type="bibr" rid="B64">2014</xref>; Lam and Yoshinori, <xref ref-type="bibr" rid="B27">2015</xref>; Tulyakov et al., <xref ref-type="bibr" rid="B57">2016</xref>). While other researchers gathered their own datasets, where experiment settings vary substantially from camera settings, lighting situations to ground truth HR measurements as shown in Appendix I in Supplementary Material. The experimental setting often consists of placing a stable digital video camera in front of the participant under a controlled lighting condition. Furthermore, a ground truth HR measurement is also collected using a more traditional method. Figure <xref ref-type="fig" rid="F3">3</xref> shows an example of the experimental setting. Finger BVP serves as the ground truth (13 out of 42 papers), while the face recordings are captured by the built-in camera of a laptop computer.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Experimental setup (Poh et al., <xref ref-type="bibr" rid="B39">2010</xref>).</p></caption>
<graphic xlink:href="fbioe-06-00033-g003.tif"/>
</fig>
<p>The digital cameras used for capturing videos are mainly commercial cameras like web cameras or portable device cameras. Following the Nyquist&#x02013;Shannon sampling theorem, it is possible to capture HR signals at a frame rate of eight frames per second (fps), under the hypothesis that the human heartbeat frequency lies between 0.4 and 4&#x02009;Hz. According to Sun et al. (<xref ref-type="bibr" rid="B51">2013</xref>), a frame rate between 15 and 30 fps is sufficient for HR detection. Among existing research, captured video frame rate differs from 15 (Poh et al., <xref ref-type="bibr" rid="B39">2010</xref>) to 100 fps (Zaunseder et al., <xref ref-type="bibr" rid="B72">2014</xref>). 30 fps, however, is the most often used within literature (Kwon et al., <xref ref-type="bibr" rid="B25">2012</xref>; Pursche et al., <xref ref-type="bibr" rid="B41">2012</xref>; Wei et al., <xref ref-type="bibr" rid="B63">2012</xref>; etc.). It is important to note that for the majority of commercial digital cameras, the frame rate is not fixed. Sudden movements or illumination changes may force to drop or interpolate frames depending on the camera used. Frame rate is also closely related with frame resolution. Generally, cameras capture higher resolution frames at relatively low frame rates and <italic>vice versa</italic>. Both resolution and frame rate influence the HR estimation performance and computation load directly. By examining the table in Appendix I in Supplementary Material, it can be seen that the majority of research tends to utilize the video graphic array standard with a video resolution of 640&#x02009;&#x000D7;&#x02009;480 pixels per frame.</p>
<p>Illumination is strictly controlled for some experiments (De Haan and Vincent, <xref ref-type="bibr" rid="B14">2013</xref>; Mestha et al., <xref ref-type="bibr" rid="B33">2014</xref>; etc.) with specified fluorescent lights and no natural sunlight. There is also research using indirect sunlight only or fluorescent lights as supplementary. The distance between the tester and the camera highly depends on the lens properties. The distance is commonly set at 1.5&#x02009;m to capture the entire face while minimizing the quantity of visualized background. In addition, the duration of recorded videos varies as well. Many face recordings are short-term with an approximate duration of 1&#x02009;min. A setting description of references can be found in Appendix I in Supplementary Material.</p>
</sec>
<sec id="S2-2">
<title>Face Video Processing</title>
<sec id="S2-2-1">
<title>ROI Selection</title>
<p>Facial ROI selection is used to obtain blood circulation features and get the raw BVP signal, which highly influences the following HR detection steps. First, it affects the tracking directly since a commonly applied tracking method uses first frame ROI (Poh et al., <xref ref-type="bibr" rid="B39">2010</xref>; Kwon et al., <xref ref-type="bibr" rid="B25">2012</xref>; Pursche et al., <xref ref-type="bibr" rid="B41">2012</xref>; etc.). Second, the selected ROI regions are regarded as the source of cardiac information. The pixel values inside a ROI are used for intensity-based methods, while feature point locations inside a ROI are used for motion-based methods. For intensity-based methods, when the selected region of the face is too large, the HR signal may be hidden in background noise. On the other hand, if the selected ROI is too small, the quantization noise caused by the camera may not be fully attenuated by the averaging of pixels intensity inside the ROI. For motion-based methods, significantly more computation time is required for a larger ROI. But there might not be enough feature points for effective motion tracking when the ROI is too small.</p>
<p>This step is similar for both intensity-based methods and motion-based methods (Figure <xref ref-type="fig" rid="F3">3</xref>). We classify the methods into two groups: box ROI detection and model-based ROI detection. Box ROI is the general area of the face regulated by a rectangle sometimes coupled with skin detection. While model-based ROI detection extracts the accurate face contours.</p>
<p>The easiest way of implementing a box ROI extraction is to manually select the desired area, such as the largest facial area with a rectangle as the bounding box on the first frame. This solution is applied for motion-based methods when the video resource contains solely hidden facial features, such as being covered by masks, or if the participants back is turned toward the camera (Balakrishnan et al., <xref ref-type="bibr" rid="B4">2013</xref>). It is simple but highly subjective. Among automatic box face detection methods, the face detector proposed by Viola and Jones (<xref ref-type="bibr" rid="B61">2001</xref>) is often applied for HR detection (Poh et al., <xref ref-type="bibr" rid="B39">2010</xref>; Balakrishnan et al., <xref ref-type="bibr" rid="B4">2013</xref>; Irani et al., <xref ref-type="bibr" rid="B20">2014</xref>; etc.). This method works rapidly and achieves reasonable detection accuracy which is 93.9% tested on MIT&#x02009;&#x0002B;&#x02009;CMU frontal face test set (Viola and Jones, <xref ref-type="bibr" rid="B61">2001</xref>). To remove face edges and background area of the box ROI only part of the detected face area is used. According to Poh et al. (<xref ref-type="bibr" rid="B39">2010</xref>), 60% of width and full height of detected facial area are used. While Mestha et al. (<xref ref-type="bibr" rid="B33">2014</xref>) use the middle 50% of the rectangle&#x02019;s width and 90% of its height. Some papers suggested to divide the roughly detected face region into a coarse grid with multiple ROIs, with the aim of removing the effect of head movements and facial expressions (Verkruysse et al., <xref ref-type="bibr" rid="B58">2008</xref>; Sun et al., <xref ref-type="bibr" rid="B51">2013</xref>; Kumar et al., <xref ref-type="bibr" rid="B24">2015</xref>; Moreno et al., <xref ref-type="bibr" rid="B35">2015</xref>). Skin detection methods are usually applied with other face detection solutions such as box ROI approach for HR estimation. Further details on this specific group can be consulted in the work of Vezhnevets et al. (<xref ref-type="bibr" rid="B59">2003</xref>), Xu et al. (<xref ref-type="bibr" rid="B67">2014</xref>), Sahindrakar et al. (<xref ref-type="bibr" rid="B45">2011</xref>), and Kakumanu et al. (<xref ref-type="bibr" rid="B23">2007</xref>).</p>
<p>Model-based approaches have been applied with accurate localization and tracking of facial landmarks (Werner et al., <xref ref-type="bibr" rid="B64">2014</xref>; Lam and Yoshinori, <xref ref-type="bibr" rid="B27">2015</xref>; Tulyakov et al., <xref ref-type="bibr" rid="B57">2016</xref>). Datcu et al. (<xref ref-type="bibr" rid="B12">2013</xref>) uses a statistical method called active appearance model to handle shape and texture variation. In Stricker et al. (<xref ref-type="bibr" rid="B50">2014</xref>), deformable model fitting by regularized landmark mean-shift (Saragih et al., <xref ref-type="bibr" rid="B46">2011</xref>) is applied. Li et al. (<xref ref-type="bibr" rid="B29">2014</xref>) applies a similar model named discriminative response map fitting with 66 facial landmarks inside the face region which is detected <italic>a priori</italic> by a Viola Jones face detector (Viola and Jones, <xref ref-type="bibr" rid="B61">2001</xref>). Tulyakov et al. (<xref ref-type="bibr" rid="B57">2016</xref>) uses the facial landmark fitting tracker&#x02014;Intraface (De la Torre et al., <xref ref-type="bibr" rid="B15">2015</xref>). Alternatively to previously explored detection methods used in HR estimation, several other approaches exist such as OpenFace (Baltru&#x00161;aitis et al., <xref ref-type="bibr" rid="B5">2016</xref>), which can detect and track facial landmarks. Each area selected per frame is dynamic when using these model-based methods. Overall this makes the selection process more robust as it can vary in time based on the features themselves, increasing its efficiency when handling head motions and facial expressions. These algorithms, however, are more computationally expensive and time-consuming than box ROI detection. Further details on face detection methods can be found in Hjelm&#x000E5;s and Low (<xref ref-type="bibr" rid="B17">2001</xref>), Vezhnevets et al. (<xref ref-type="bibr" rid="B59">2003</xref>), Rother et al. (<xref ref-type="bibr" rid="B42">2004</xref>), and Baltru&#x00161;aitis et al. (<xref ref-type="bibr" rid="B5">2016</xref>).</p>
</sec>
<sec id="S2-2-2">
<title>Color Channel Decomposition</title>
<p>This step is specific for intensity-based methods. The basic idea is that the pixel intensity captured by a digital camera can be decomposed into the illumination intensity and reflectance of the skin. However, several approaches are proposed to relate the pixel value to the PPG signal. This can lead to various choices of color channel decomposition and combination. For example, Huelsbusch and Blazek (<xref ref-type="bibr" rid="B18">2002</xref>) separated the noise from the PPG signal by building a linear combination of two color channels to achieve motion robustness. An in-depth description of optical modeling can be found within the literature referenced in Table <xref ref-type="table" rid="T1">1</xref>.</p>
<p>Color channels are based on the color models. There are mainly three color models applied in HR detection: Red-Green-Blue (RGB), Hue-Saturation-Intensity (HSI), and YCbCr where Y stands for luminance component and Cb and Cr refer to blue-difference and red-difference chroma components, respectively. The HSI model decouples the intensity component from the hue and saturation that carry color information of a color image. The skin-color lies in a certain range of H ([0 50]) and S ([0.23 0.68]) channels and the illumination changes information is separated in I channel. With each heartbeat, there is a clear drop in hue channel but its amplitude is very small. For HSI model, only the H channel can be used for BVP signal extraction. It is motion sensitive but performs better than RGB model without head motions. According to Sahindrakar et al. (<xref ref-type="bibr" rid="B45">2011</xref>), YCbCr produced better results in detecting HR than HSI with limited rotation and no transition. Among these three models, the most robust model is still RGB.</p>
<p>Among current detection methods, the main color space is still RGB, though some research criticizes that it intermixes the color and intensity information. According to Verkruysse et al. (<xref ref-type="bibr" rid="B58">2008</xref>), Stricker et al. (<xref ref-type="bibr" rid="B50">2014</xref>), and Ruben (<xref ref-type="bibr" rid="B43">2015</xref>), all channels contain PPG information, but the green channel gives the strongest signal-to-noise ratio (SNR). Consequently, the green channel has been the most popularly used for extracting HR (Verkruysse et al., <xref ref-type="bibr" rid="B58">2008</xref>; Li et al., <xref ref-type="bibr" rid="B29">2014</xref>; Zaunseder et al., <xref ref-type="bibr" rid="B72">2014</xref>; Chen et al., <xref ref-type="bibr" rid="B9">2015</xref>; etc.). However, Lewandowska et al. (<xref ref-type="bibr" rid="B28">2011</xref>) showed that the combination of the R and G channels contain the majority of cardiac information. Several research papers have also investigated the usage of all three color channels in conjunction with BSS for the BVP signal extraction (Poh et al., <xref ref-type="bibr" rid="B40">2011</xref>; Kwon et al., <xref ref-type="bibr" rid="B25">2012</xref>; Pursche et al., <xref ref-type="bibr" rid="B41">2012</xref>; etc.).</p>
</sec>
<sec id="S2-2-3">
<title>Raw Featured Signal</title>
<p>Intensity-based methods use the intensity changes along the time as raw signal containing BVP information, while motion-based methods use the vertical component of the trajectories instead.</p>
<p>The spatial average is commonly employed in the majority of intensity-based methods, which aims to increase the SNR of PPG signals and enhance the subtle color changes (Verkruysse et al., <xref ref-type="bibr" rid="B58">2008</xref>). Depending on the color channel selection, all pixels of the corresponding color channel within the ROI area are averaged at each frame. For a RGB video with <italic>n</italic> frames, the signal after spatial average can be expressed as a vector: <italic>X</italic>(<italic>j</italic>)&#x02009;&#x0003D;&#x02009;(<italic>x</italic><sub>1</sub>(<italic>j</italic>), <italic>x</italic><sub>2</sub>(<italic>j</italic>), &#x02026;, <italic>x<sub>n</sub></italic>(<italic>j</italic>)), <italic>j</italic>&#x02009;&#x0003D;&#x02009;1, 2, 3 where <italic>j</italic> stands for the color channels. This method is simple and efficient to get raw featured signals for intensity-based methods. Several research papers used the spatial-averaged signal directly, as shown in Table <xref ref-type="table" rid="T1">1</xref>. On the other hand, some works (De Haan and Vincent, <xref ref-type="bibr" rid="B14">2013</xref>; Tulyakov et al., <xref ref-type="bibr" rid="B57">2016</xref>) apply optical models and use the chrominance features for HR estimation, which takes light transmission and reflection on skin into consideration.</p>
<p>For motion-based methods, the location of time-series <italic>x<sub>k</sub></italic>(<italic>n</italic>), <italic>y<sub>k</sub></italic>(<italic>n</italic>) for each feature point <italic>k</italic> on frame <italic>n</italic> is tracked. Only the vertical component <italic>y<sub>k</sub></italic>(<italic>n</italic>) is taken to extract the trajectory from each feature point. The longitudinal trajectories are then used as raw featured signals.</p>
</sec>
</sec>
<sec id="S2-3">
<title>Face BVP Signal Extraction</title>
<p>Now that the feature signal has been obtained from face videos, the heartbeat can be effectively extracted. This section is divided into two subsections exploring noise reduction and dimensionality reduction methods.</p>
<sec id="S2-3-1">
<title>Noise Reduction</title>
<p>As previously explored, color and motion changes caused by the cardiac activities are often noisy. Thus this step is applied on the raw signals to remove such changes in light and tracking errors. For intensity-based methods, the light variations are recorded together with intensity changes caused by blood pulses (Verkruysse et al., <xref ref-type="bibr" rid="B58">2008</xref>; Li et al., <xref ref-type="bibr" rid="B29">2014</xref>; Zaunseder et al., <xref ref-type="bibr" rid="B72">2014</xref>; etc.). For motion-based methods, trackers capture trajectories that are not solely caused by heartbeats, thus it is necessary to apply noise reduction as well (Balakrishnan et al., <xref ref-type="bibr" rid="B4">2013</xref>; Irani et al., <xref ref-type="bibr" rid="B20">2014</xref>). We present noise reduction methods based on two categories: temporal filtering and background noise estimation. Temporal filtering contains a series of filters that remove irrelevant information and keep trajectories and color frequencies that are of interest. Background noise estimation uses the background to estimate the noise caused by light changes.</p>
<p>For temporal filtering, various temporal filters are applied to exclude and amplify low-amplitude changes revealing hidden information (Poh et al., <xref ref-type="bibr" rid="B39">2010</xref>; Wang, <xref ref-type="bibr" rid="B62">2017</xref>). It contains detrending, moving-average, and bandpass filters, which are often applied to reduce irrelevant noise (Li et al., <xref ref-type="bibr" rid="B29">2014</xref>). A detrending filter aims to reduce slow and non-stationary trends of signals (Li et al., <xref ref-type="bibr" rid="B29">2014</xref>). After applying a detrending filter, the low frequencies of the raw signal are reduced drastically. This method is as effective as a high-pass, low-cutoff filter with substantially less latency. The moving-average filter removes random noise with temporal average of consecutive frames. It can efficiently smooth the trajectories and sudden color changes caused by light or motions. Additional methods such as a bandpass filter can also be used to remove irrelevant frequencies. Bandpass filters can be Butterworth or other FIR bandpass filters in the literature. It can be a Butterworth filter (Balakrishnan et al., <xref ref-type="bibr" rid="B4">2013</xref>; Irani et al., <xref ref-type="bibr" rid="B20">2014</xref>; Osman et al., <xref ref-type="bibr" rid="B37">2015</xref>; etc.) or other FIR bandpass filters (Li et al., <xref ref-type="bibr" rid="B29">2014</xref>) with cutoff frequency of normal HR. The cutoff frequency could be 0.7&#x02013;4 (Villarroel et al., <xref ref-type="bibr" rid="B60">2017</xref>), 0.25&#x02013;2 (Wei et al., <xref ref-type="bibr" rid="B63">2012</xref>), or other values. The parameter setting for these three types of filters differs from papers. For example, Ruben (<xref ref-type="bibr" rid="B43">2015</xref>) applies a fourth order bandpass zero-phase Butterworth filter, while Irani et al. (<xref ref-type="bibr" rid="B20">2014</xref>) employs an eighth order Butterworth filter to flat pass band maximally. More temporal filtering information applied for HR detection can be found in the work of Yu et al. (<xref ref-type="bibr" rid="B70">2014</xref>) and Tarvainen et al. (<xref ref-type="bibr" rid="B54">2002</xref>).</p>
<p>Background noise estimation methods target intensity changes and are only suitable under some situations. It is based on the assumptions that (a) both the ROI and background share the same light source and (b) the background is static and relatively monotone (Li et al., <xref ref-type="bibr" rid="B29">2014</xref>). Under these assumptions, the intensity changes in the background are caused by illumination only and are correlated with the light noise in the HR signal extracted from face recordings. Adaptive filters are applied on noised HR signal and background signal to remove the noise (Chan and Zhang, <xref ref-type="bibr" rid="B8">2002</xref>; Cennini et al., <xref ref-type="bibr" rid="B7">2010</xref>; Li et al., <xref ref-type="bibr" rid="B29">2014</xref>).</p>
<p>Once filtered, signals can be used either directly for post-processing or for further signal extraction (dimensionality reduction). If the signal is used directly, the green channel is mostly used since it contains the stronger PPG signal (Verkruysse et al., <xref ref-type="bibr" rid="B58">2008</xref>). Under the second case, all the signal channels are kept.</p>
</sec>
<sec id="S2-3-2">
<title>Dimensionality Reduction Methods</title>
<p>The BVP signal is a periodic one-dimensional signal in the time domain. Dimensionality reduction algorithms are used to reduce the dimensionality from raw signals in order to more clearly reveal BVP information. The main idea is to find a mapping between higher dimensional space, such as three-dimensional RGB color spaces and one-dimensional space uncovering cardiac information. Dimensionality reduction contains classic linear algorithms, e.g., BSS methods [e.g., independent component analysis (ICA) and principle component analysis (PCA)], linear discriminant analysis, and manifold learning methods such as Isomap, Laplacian Eigenmap (LE), and locally linear embedding (Zhang and Zha, <xref ref-type="bibr" rid="B73">2004</xref>). Wei et al. (<xref ref-type="bibr" rid="B63">2012</xref>) tested nine commonly used dimensionality reduction methods on RGB color channels and the result demonstrated that LE performs best for extracting BVP information on their dataset.</p>
<sec id="S2-3-2-1">
<title>Blind Source Separation</title>
<p>After Poh&#x02019;s publication in 2010, the mainstream technique to recover the BVP signal has been BSS which assumes that the observed signals (in our case the featured signals) are a mixture of source signals (BVP and noise). The goal of BSS is to recover the sources signals without or with a little prior information about their properties. The most popular BSS methods for HR detection from face video are ICA (Hyv&#x000E4;rinen and Oja, <xref ref-type="bibr" rid="B19">2000</xref>) and PCA (Wold et al., <xref ref-type="bibr" rid="B65">1987</xref>; Abdi and Williams, <xref ref-type="bibr" rid="B1">2010</xref>).</p>
<p>Independent component analysis is based on the assumption that all the sources are mutually independent. The basic principle is to maximize the statistical independence of all observed components in order to find the underlying components (Liao and Carin, <xref ref-type="bibr" rid="B30">2002</xref>; Yang, <xref ref-type="bibr" rid="B68">2006</xref>). For cardiac pulse detection, the observed signals are captured by camera color sensors, which are mixed with the heartbeat signals. Among various ICA algorithms, Joint Approximate Diagonalization of Eigen-matrices (JADE) (Cardoso, <xref ref-type="bibr" rid="B6">1999</xref>) is popular for HR detection since it is numerically efficient in computation (Poh et al., <xref ref-type="bibr" rid="B39">2010</xref>, <xref ref-type="bibr" rid="B40">2011</xref>; Kwon et al., <xref ref-type="bibr" rid="B25">2012</xref>; Pursche et al., <xref ref-type="bibr" rid="B41">2012</xref>; etc.). JADE is a high-order measures of independence for ICA. Further details on the JADE algorithm can be found in the work of Hyv&#x000E4;rinen and Oja (<xref ref-type="bibr" rid="B19">2000</xref>), while methods for optimizing JADE is further described by Kumar et al. (<xref ref-type="bibr" rid="B24">2015</xref>).</p>
<p>Principle component analysis can be used to extract both the intensity-based pulse signal and the head longitude trajectories caused by pulse (Lewandowska et al., <xref ref-type="bibr" rid="B28">2011</xref>; Balakrishnan et al., <xref ref-type="bibr" rid="B4">2013</xref>; Rubinstein, <xref ref-type="bibr" rid="B44">2013</xref>). For motion-based method, the frequency spectra of the PCA components with the highest periodicity is selected quantified from spectral power, meanwhile the component with maximum variance is selected for intensity-based method as BVP signal (Lewandowska et al., <xref ref-type="bibr" rid="B28">2011</xref>). Compared with ICA, PCA has lower computation complexity. PCA is concerned with finding the directions along which the data have maximum variance in addition to the relative importance of these directions. For HR detection, the goal of applying PCA is to extract the cardiac pulse information from the head motions or pixel intensity changes, represented it into principal components consisting of a new set of orthogonal variables (Abdi and Williams, <xref ref-type="bibr" rid="B1">2010</xref>; Balakrishnan et al., <xref ref-type="bibr" rid="B4">2013</xref>). Mathematically, PCA depends on the Eigen-decomposition of positive semi-definite matrices and the singular value decomposition of rectangular matrices (Wold et al., <xref ref-type="bibr" rid="B65">1987</xref>).</p>
</sec>
</sec>
</sec>
<sec id="S2-4">
<title>HR Computation</title>
<p>Once the BVP signal is effectively extracted, a post-processing procedure follows. HR can be estimated from time domain analysis (peak detection methods) or frequency domain analysis. Signals can be transformed to the frequency domain using standard methods, such as the Fast Fourier Transform (FFT) and discrete cosine transformation (DCT) methods. Currently, supervised learning methods are only applied at this stage for both the time and frequency domains.</p>
<p>Frequency domain algorithms are the most common post-processing methods within the literature. The extracted HR signal is converted to the frequency domain either by FFT (Poh et al., <xref ref-type="bibr" rid="B39">2010</xref>; Pursche et al., <xref ref-type="bibr" rid="B41">2012</xref>; Yu et al., <xref ref-type="bibr" rid="B69">2013</xref>; etc.) or DCT (Irani et al., <xref ref-type="bibr" rid="B20">2014</xref>). Using these methods, there is an assumption that the HR is the most periodic signal and thus, has the highest power of the spectrum within the frequency band corresponding to normal human HR. The drawback is that it can only compute the HR over a certain period instead of detecting instantaneous HR changes.</p>
<p>Peak detection methods (Poh et al., <xref ref-type="bibr" rid="B40">2011</xref>; Li et al., <xref ref-type="bibr" rid="B29">2014</xref>) detect the peak of HR signals in the time domain directly. With the detected peaks, IBI can be calculated. The IBI intervals are then averaged and the HR is computed from the average IBI. IBI allows for the beat-to-beat assessment of HR, however, it is quite sensitive to noise. To achieve more reliable results, a sliding window of short-time period is often implemented to average the HR result over the whole video (Li et al., <xref ref-type="bibr" rid="B29">2014</xref>; Ruben, <xref ref-type="bibr" rid="B43">2015</xref>).</p>
<p>Supervised learning methods have also been investigated as a potential solution for HR calculation (Monkaresi et al., <xref ref-type="bibr" rid="B34">2014</xref>; Tarassenko et al., <xref ref-type="bibr" rid="B53">2014</xref>; Osman et al., <xref ref-type="bibr" rid="B37">2015</xref>). Monkaresi et al. (<xref ref-type="bibr" rid="B34">2014</xref>) and Tarassenko et al. (<xref ref-type="bibr" rid="B53">2014</xref>) both use supervised learning method for power spectrum analysis. With features extracted from PSD, auto-regression or <italic>k</italic>-nearest neighbor classifier is used to predict the HR signal with a degree of accuracy. Osman et al. (<xref ref-type="bibr" rid="B37">2015</xref>) extract the first order derivative of green channels as the features. A feature at time <italic>t</italic> is positively labeled, if there is a ground truth BVP peak lies within a certain time tolerance and <italic>vice versa</italic>. These data are then used to train a support vector machine algorithm, capable of predicting IBI.</p>
</sec>
<sec id="S2-5">
<title>Discussion</title>
<p>Developed methods tend to be strongly tied into the dataset and the specific experiment protocol they were designed for. Unfortunately, this means that they neither generalize nor adapt to other datasets or scenarios, especially real-life situations. This section analyses the limitation aspect from setting to results.</p>
<sec id="S2-5-1">
<title>Experiment Setting</title>
<p>As shown in Section &#x0201C;<xref ref-type="sec" rid="S2-5-1">Experiment Setting</xref>,&#x0201D; experimental settings can vary significantly. For most experiments, both the participants&#x02019; behaviors and the environment are well-controlled which are not applicable for practical use. Non-grid motions like facial expressions are more difficult to handle compared with grid motions, e.g., head rotate horizontally or vertically. Detecting HR from face videos with spontaneous facial expressions is valuable for further study such as long-term monitoring and affective computing.</p>
<p>So far, there is no research designed specifically to test each influential factor of experimental setting such as video resolution, frame rate, and illumination changes. Consequently, it is difficult to distinguish whether the study results are affected by the experimental setting or the implemented approach. For example, low video resolution will lead to a limited number of pixels in the face area, which may be insufficient to extract the BVP signal. A high frame rate provides more information but may increase the computational load.</p>
<p>Furthermore, most of the self-collected datasets are not publicly accessible, complicating further investigations as researchers must continuously collect and construct new datasets, which can be time consuming. In fact, it complicates evaluation as methods are tested cross-dataset. A few studies used the open dataset MAHNOB-HCI (Li et al., <xref ref-type="bibr" rid="B29">2014</xref>; Lam and Yoshinori, <xref ref-type="bibr" rid="B27">2015</xref>; Tulyakov et al., <xref ref-type="bibr" rid="B57">2016</xref>). However, their results for the same method are not consistent with each other (Li et al., <xref ref-type="bibr" rid="B29">2014</xref>; Lam and Yoshinori, <xref ref-type="bibr" rid="B27">2015</xref>). This is probably due to differences in the implementation.</p>
<p>Besides, participants with darker skin tone are also rarely used within both self-collected and open datasets. The higher amount of melanin present in darker skin tones absorbs a significant amount of incident light, and thus degrades the quality of the camera-based PPG signals, making the system ineffective for extracting vital signs (Kumar et al., <xref ref-type="bibr" rid="B24">2015</xref>). Equipment limitations also exist. For example, the majority of digital cameras are unable to hold stable frame rates during recording, which can impact HR computation. Video compression is another factor that can influence performance (Ruben, <xref ref-type="bibr" rid="B43">2015</xref>). According to McDuff et al. (<xref ref-type="bibr" rid="B32">2017</xref>), even with a low constant rate factor to compress videos, the signal-to-noise ratio degrades considerably in face BVP signals.</p>
</sec>
<sec id="S2-5-2">
<title>Method Discussion</title>
<p>The ROI definition is important since it contains the raw BVP signals. However, as mentioned in Section &#x0201C;<xref ref-type="sec" rid="S2-2-1">ROI Selection</xref>,&#x0201D; there is still no consensus on which ROIs are the most relevant for HR computation. For example, Pursche et al. (<xref ref-type="bibr" rid="B41">2012</xref>) claim that the center of the face region provides better PPG information compared with other facial parts. By contrast, Lewandowska et al. (<xref ref-type="bibr" rid="B28">2011</xref>) and Stricker et al. (<xref ref-type="bibr" rid="B50">2014</xref>) assert that the forehead can represent the whole facial region although it can be unreliable if covered with hair. In Datcu et al. (<xref ref-type="bibr" rid="B12">2013</xref>), cheek and forehead are suggested as the most reliable parts containing the strongest PPG signals. Moreno et al. (<xref ref-type="bibr" rid="B35">2015</xref>) believe that forehead, cheeks, and mouth area provide more accurate heartbeat signals, in comparison with other parts such as nose and eyes. While Irani et al. (<xref ref-type="bibr" rid="B20">2014</xref>) state that the forehead and area around the nose are more reliable. These differences are caused specifically by the datasets used. For example, with videos containing head motions and facial expressions (Lewandowska et al., <xref ref-type="bibr" rid="B28">2011</xref>; Stricker et al., <xref ref-type="bibr" rid="B50">2014</xref>), the eyes and mouth areas tend to be less stable than the forehead area, since they are influenced by facial muscles.</p>
<p>The most popular method to extract the face BVP signal is by far ICA, as previously detailed in Section &#x0201C;<xref ref-type="sec" rid="S2-3">Face BVP Signal Extraction</xref>.&#x0201D; It works experimentally, nevertheless, it also has some limitations. Based on the Beer&#x02013;Lambert law, reflected light intensity through facial tissue varies nonlinearly with distance (Wei et al., <xref ref-type="bibr" rid="B63">2012</xref>; Xu et al., <xref ref-type="bibr" rid="B67">2014</xref>), while both PCA and ICA assume that the observed signal is a linear combination of several sources. Supporting the possibility that ICA is not the best option, Kwon et al. (<xref ref-type="bibr" rid="B25">2012</xref>) demonstrated that ICA performance is slightly lower compared with simple green channel detection methods. Besides, according to Mestha et al. (<xref ref-type="bibr" rid="B33">2014</xref>) and Sahindrakar et al. (<xref ref-type="bibr" rid="B45">2011</xref>), ICA needs at least 30&#x02009;s to be accurately estimated and cannot handle large head motions. The noise estimation method proposed by Li et al. (<xref ref-type="bibr" rid="B29">2014</xref>) can only be applied with monotone backgrounds. Also, it is not suitable for real-time HR prediction with its adaptive filter, which needs a certain time duration to guarantee estimation accuracy. For stable frontal face videos, the green channel tends to be a good solution, and provides a low computational complexity method for the extraction of BVP signals. Particularly noisy sources of the face can often be mitigated through the usage of methods like PCA or ICA, and can often achieve better results than using the green channel exclusively (Poh et al., <xref ref-type="bibr" rid="B40">2011</xref>; Li et al., <xref ref-type="bibr" rid="B29">2014</xref>).</p>
<p>At the HR computation stage, the frequency domain methods are not capable to detect instantaneous heartbeat changes, and are not as robust as time domain methods according to Poh et al. (<xref ref-type="bibr" rid="B40">2011</xref>). Supervised learning methods are mainly (three out of four papers) applied at this stage so far. There is one recent paper apply auto-regression to extract BVP signals after the video processing (Villarroel et al., <xref ref-type="bibr" rid="B60">2017</xref>), but there is no end-to-end usage of supervised learning methods. Future research might thus focus on the development of machine learning methods trained to take raw videos as inputs and compute the HR information as outputs. Without inter-processing stage, the performance of supervised learning methods may be improved significantly.</p>
</sec>
<sec id="S2-5-3">
<title>HR Estimation</title>
<p>The HR discussed in this article is actually the average of HR during a certain time interval (e.g., 30&#x02009;s or 1&#x02009;min), which cannot reveal the instantaneous physiological information. The literature often computes HR by examining videos with various lengths and subsequently calculating the average HR over that time period. In reality, the time interval between two connective heartbeats is not stable. By contrast to averaged HR, we refer to instantaneous or dynamic HR when talking about HR calculated for each IBI. This information can be used to reveal short lived phenomenon such as emotions. It can further be used to compute heart rate variability (HRV). According to Pavlidis et al. (<xref ref-type="bibr" rid="B38">2007</xref>) and Dawson et al. (<xref ref-type="bibr" rid="B13">2010</xref>), HRV is directly related with emotion and disease diagnosis, which is of great value in medical and affective computing domain.</p>
<p>To the best of our knowledge, so far there is no work focusing on dynamic HR. Also, for average HR estimation, there are no state-of-the-art approaches that are robust enough to be fully operated under real situations with grid and non-grid movements, illumination changes and noise caused by the camera. Even with well-established public dataset, one or two influential factors are still under controlled.</p>
</sec>
<sec id="S2-5-4">
<title>Commercial Applications and Software</title>
<p>Currently, for application purposes, PPG is mainly obtained from contact sensors. There are few commercial applications and software available estimating HR remotely from the color changes on faces such as Cardiio,<xref ref-type="fn" rid="fn1"><sup>1</sup></xref> Pulse Meter,<xref ref-type="fn" rid="fn2"><sup>2</sup></xref> and Vital Signs Camera.<xref ref-type="fn" rid="fn3"><sup>3</sup></xref> Cardiio and Pulse Meter are both phone applications. They present a circle or a rectangle on the screen for users to place their face areas and keep still for a certain time period. Vital sign camera is developed by Philips with both software and phone application. They apply frequency domain method to calculate HR. All these applications and software compute the average HR only and require users to stay stable. None of these products provide an estimation of their performance.</p>
</sec>
</sec>
</sec>
<sec id="S3">
<title>Comparative Analysis</title>
<p>It is of great importance to quantify and compare the performance of the main algorithms presented in the previous section. In this section, we study how the existing approaches perform in a close to realistic scenario with both gird and non-grid head movements.</p>
<p>Given the amount of workload, not all methods were tested exhaustively. For motion-based methods, only two methods were proposed and applied on face videos (Balakrishnan et al., <xref ref-type="bibr" rid="B4">2013</xref>; Irani et al., <xref ref-type="bibr" rid="B20">2014</xref>). Both of them required limited head movements. Our test dataset, MANNOB&#x02013;HCI, is not ideal for this group of methods. Therefore, we focus on the validation of intensity-based approaches which are most often applied. Despite of this restriction, some methods evaluated in this section can also be used as reference for motion-based methods. This is for instance the case for ROI selection which can be used in a motion-based framework.</p>
<p>To offer a panoramic view of the state-of-the-art, this article attempted to cover the analysis with methods from different categories within each processing stage as previously stated in Section &#x0201C;<xref ref-type="sec" rid="S2">Remote Methods for HR Detection</xref>.&#x0201D; The methods considered for comparative analysis are selected based on their popularity (times that they were adopted for other papers) and their category (supervised learning, dimensionality reduction, etc.).</p>
<p>The method validation is divided into two parts. In the first part, we test and compare the algorithms at each stage sequentially as shown in Figure <xref ref-type="fig" rid="F4">4</xref>. For one stage, we test several alternatives while keeping the algorithms applied at the other two stages fixed. Once a stage has been tested, the best method on this stage will be chosen for the following stage test. Before that, the algorithms selected for fixed stages are based on simplicity. At the pre-processing stage, the main target is to find out the efficient face segmentation. For signal extraction, various algorithms are tested to compare the most efficient method capable of separating the HR signal from noise and irrelevant information. At the post-processing stage, time domain and frequency domain methods are compared for HR computation. The second part focuses on the implementation of state-of-the-art methods presented in the work of Poh et al. (<xref ref-type="bibr" rid="B40">2011</xref>), Li et al. (<xref ref-type="bibr" rid="B29">2014</xref>), and Osman et al. (<xref ref-type="bibr" rid="B37">2015</xref>), which are used as baselines for validation.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Overview of region of interest selection (Stricker et al., <xref ref-type="bibr" rid="B50">2014</xref>).</p></caption>
<graphic xlink:href="fbioe-06-00033-g004.tif"/>
</fig>
<p>The state-of-art methods were tested on the MAHNOB-HCI (Soleymani et al., <xref ref-type="bibr" rid="B47">2012</xref>) database. In MAHNOB-HCI, 27 participants (15 females and 12 males) were recorded while watching movie clips to elicit emotions, which can influence their HR and stimulate facial expressions. This public dataset is multimodal including frontal face videos and HR information recorded using a gold standard technique: ECG. The frame rate is 61 fps and the ECG sampling frequency is 256&#x02009;Hz. Furthermore, this dataset is publicly accessible, allowing research to be easily reproduced. We filtered 465 samples (i.e., ECG and facial videos which each correspond to an emotion-elicited movie clip) from MAHNOB-HCI obtained from 24 participants, where 12 were males and 12 were females.</p>
<p>To reduce the possible risks of bias, we regulated the study from data source to performance validation. The face videos used for the study are from both genders and different skin tones. For one participant, 14&#x02013;20 videos are used to avoid bias from a certain scenario due to a specific stimulus. All participants we selected offered their consent before experimentation, where the recording duration surpasses 65&#x02009;s. Among the 465 samples, 20 samples did not contain corresponding ECG signals and 1 recording presented faulty filtering. These samples were removed leaving a total of 444 samples. The validation followed by Li et al. (<xref ref-type="bibr" rid="B29">2014</xref>) and Lam and Yoshinori (<xref ref-type="bibr" rid="B27">2015</xref>) started each video recording with a delay of 5&#x02009;s. Videos are subsequently cut into 30&#x02009;s segments, and subsequently synchronized with their corresponding ECG signal. All the methods and evaluations are implemented using MATLAB R2016a. To validate obtained results, the HR ground truth is obtained by first detecting the R peaks from ECG signals using the TEAP (Soleymani et al., <xref ref-type="bibr" rid="B48">2017</xref>), which uses the standard Tompkins&#x02019; method (Tompkins, <xref ref-type="bibr" rid="B55">1993</xref>). Absent and falsely detected peaks are then manually corrected. The mean HR is the averaged value computed from the instantaneous HR. Thus, a precise average HR is guaranteed as ground truth.</p>
<sec id="S3-1">
<title>Comparison at Each Stage</title>
<sec id="S3-1-1">
<title>Face Segmentation</title>
<p>As shown in Section &#x0201C;<xref ref-type="sec" rid="S2-2">Face Video Processing</xref>,&#x0201D; there is no consistent conclusion about which part of the face reveals most HR information. Thus, there is an interest in comparing HR detection accuracy for several facial regions. The face region contains useful features of HR information that may differ from every frame since appearance changes are spatially and temporally localized. Instead of using a constant and preselected ROI, we adopt OpenFace (Baltru&#x00161;aitis et al., <xref ref-type="bibr" rid="B5">2016</xref>) to detect face segmentations automatically and dynamically. OpenFace can extract 66 landmarks from the face, marking the location of eyes, nose, eyebrows, mouth, and the contour of visage. We compared the performance of forehead, cheeks, chin, whole face, and extracted skin area (accurate face contour area without eyes and mouth), which are frequently selected as ROIs in the state-of-the-art. For the rectangular forehead area, we used the distance between inner corners of the eyes as the rectangle width, while the distance from the uppermost face contour landmark to the uppermost eyes landmark constitutes the rectangle height. Similarly, the cheek areas had the same width as eyes and the height is the distance between upper lip border and the lower eye border. For the skin area, we removed the eye and mouth regions to avoid noise caused by blinks and other facial expressions. The regions we tested are shown in Figure <xref ref-type="fig" rid="F5">5</xref>. Since all the five ROI selections are determined by facial landmarks, the ROI areas are dynamic and may change due to the head motions from frame-to-frame. Given that in the MAHNOB-HCI database, the participants&#x02019; electroencephalogram (EEG) was recorded using a head cap, the forehead area was partially covered by the sensors that may influence the performance.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Face segmentation (1. Forehead; 2. Cheeks; 3. Chin; 4. Whole face; 5. Extracted skin).</p></caption>
<graphic xlink:href="fbioe-06-00033-g005.tif"/>
</fig>
<p>To avoid influential factors from other processing steps, we use the spatial-averaged green channel directly of each ROI and then apply temporal filtering to obtain the HR signal. The cutoff frequency is set from 0.7 to 2&#x02009;Hz which corresponds to HR between 42 and 120&#x02009;bpm. The PSD method (Poh et al., <xref ref-type="bibr" rid="B39">2010</xref>) is applied to calculate the averaged HR over 30&#x02009;s.</p>
</sec>
<sec id="S3-1-2">
<title>Face BVP Signal Extraction</title>
<p>For this part, we tested the main methods mentioned in Section &#x0201C;<xref ref-type="sec" rid="S2-3-1">Noise Reduction</xref>.&#x0201D; Temporal filtering with a detrending filter and a bandpass filter are applied on the raw featured signal before methods testing in this section. Three signal extraction methods were compared. PCA, ICA, and background noise estimation methods are evaluated on spatial-averaged RGB signals from extracted skin ROI. For ICA, the component with highest energy in the frequency domain is selected as the BVP signal, while the component with the highest variance is selected for PCA. The implementation of ICA and PCA follows Poh et al. (<xref ref-type="bibr" rid="B40">2011</xref>) and Rubinstein (<xref ref-type="bibr" rid="B44">2013</xref>), respectively. Background noise estimation is implemented following Li et al. (<xref ref-type="bibr" rid="B29">2014</xref>) and uses green channel for extracting HR signals after illumination rectification with normalized least mean square adaptive filter. HR is estimated using the PSD of the extracted face BVP signal (Poh et al., <xref ref-type="bibr" rid="B39">2010</xref>) as well.</p>
</sec>
<sec id="S3-1-3">
<title>HR Computing</title>
<p>The peak detection method and PSD method are evaluated for the calculation of HR. Selected ICA components extracted from skin ROI are used as the HR signal. For peak detection, we applied the algorithm from the open source Toolbox for Emotional feAture extraction from Physiological signals (TEAP) (Soleymani et al., <xref ref-type="bibr" rid="B48">2017</xref>). There are different ways of computing PSD. In this section, it is estimated <italic>via</italic> the periodogram method following Monkaresi et al. (<xref ref-type="bibr" rid="B34">2014</xref>).</p>
</sec>
</sec>
<sec id="S3-2">
<title>Comparison on Complete Methods</title>
<p>We reproduced three methods from the work of Poh et al. (<xref ref-type="bibr" rid="B40">2011</xref>), Li et al. (<xref ref-type="bibr" rid="B29">2014</xref>), and Osman et al. (<xref ref-type="bibr" rid="B37">2015</xref>). This was done to compare our analysis with the three main categories of HR estimation methods: ICA, background noise estimation, and machine learning, respectively. We implement these three approaches step-by-step and set parameters as mentioned in the papers. The implementation schematic is shown in Figure <xref ref-type="fig" rid="F6">6</xref>.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Schematic diagram for complete method comparison. <bold>(A)</bold> Schematic diagram from Poh et al. (<xref ref-type="bibr" rid="B40">2011</xref>). <bold>(B)</bold> Schematic diagram from Li et al. (<xref ref-type="bibr" rid="B29">2014</xref>). <bold>(C)</bold> Schematic diagram from Osman et al. (<xref ref-type="bibr" rid="B37">2015</xref>).</p></caption>
<graphic xlink:href="fbioe-06-00033-g006.tif"/>
</fig>
<p>Poh et al. and Osman et al. tested their method on self-collected datasets with BVP signals as ground truth, while Li et al. used both self-collected and the MAHNOB-HCI datasets.</p>
<p>The work of Osman et al. (<xref ref-type="bibr" rid="B37">2015</xref>) used finger BVP signals to label their extracted features. In our reproduction, ECG signals are used from MAHNOB-HCI. We followed the same method but labeled the feature positive if the R peak existed within the time tolerance of the detected peak from face videos. We randomly selected 12 subjects as training dataset and the other 12 subjects as testing dataset. For training, we have 5,000 positive features and 5,000 negative features as Osman et al. (<xref ref-type="bibr" rid="B37">2015</xref>). For testing, there are 6,765 positive features and 7,127 negative ones.</p>
</sec>
</sec>
<sec id="S4" sec-type="discussion">
<title>Results and Discussion</title>
<sec id="S4-1">
<title>Results at Each Stage</title>
<p>For face segmentation, we can see from Figure <xref ref-type="fig" rid="F7">7</xref>A that the extracted skin area performs better than the other facial areas, followed by the entire facial region. Though skin area performs a bit better than face area, there is no significant difference between them (<italic>t</italic>&#x02009;&#x0003D;&#x02009;1.72; <italic>p</italic>&#x02009;&#x0003D;&#x02009;0.07). There is no noticeable difference between the left cheek and the right cheek and the forehead does not perform better than other facial areas. According to our results, the more skin area used as BVP resource, the better performance we can achieve (extracted skin area performs better than cheek, forehead, and chin areas). The facial expressions, such as laugh and blinks, tends to add extra noise to the BVP signals but does not influence it considerably when taking the whole face into consideration. We also compared the whole face area with the area detected by Viola Jones face detector (Viola and Jones, <xref ref-type="bibr" rid="B61">2001</xref>). When there are no abrupt head moments and other interruptions, the featured raw signals obtained from these two methods are very similar. The Pearson correlation is statistically significant (<italic>r</italic>&#x02009;&#x0003D;&#x02009;0.99; <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.001). When the video includes more spontaneous movements, OpenFace is more robust with a higher success detection rate and the correlation between the two methods is degraded (<italic>r</italic>&#x02009;&#x0003D;&#x02009;0.83; <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.001). Viola Jones detector fails face detection on several frames and then uses the last successfully detected face information which adds noise in face BVP signal. Considering the computation complexity and detection efficiency, Viola Jones face detector is suitable for slight head movements and can be used as a prior method for more complex face detection algorithms.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Method performance at each stage. &#x0201C;best&#x0201D; indicate the best methods according to average performance. Other methods are tested against the best method (ns, non-significant; &#x0002A;, <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.05; &#x0002A;&#x0002A;, <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.01). <bold>(A)</bold> Face segmentation performance. <bold>(B)</bold> Performance at face blood volume pulse (BVP) extraction. <bold>(C)</bold> Performance at heart rate (HR) computation.</p></caption>
<graphic xlink:href="fbioe-06-00033-g007a.tif"/>
<graphic xlink:href="fbioe-06-00033-g007b.tif"/>
</fig>
<p>To see how head movements influence signal extraction, we tested the methods on one video with little head motions. For this situation, background noise estimation performs better since the illumination was the main source of noise. However, there is little difference between the root mean squared error of the green channel (12.12), PCA (12.51), ICA (11.80), and background noise estimation (11.69). Evaluating on all the samples, ICA performs better then background noise estimation (<italic>t</italic>&#x02009;&#x0003D;&#x02009;2.27; <italic>p</italic>&#x02009;&#x0003D;&#x02009;0.04) and PCA (<italic>t</italic>&#x02009;&#x0003D;&#x02009;4.84; <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.001) as shown in Figure <xref ref-type="fig" rid="F7">7</xref>B. Under the experimental setting of MAHNOB-HCI, head motions tend to have a higher influence on the HR detection, rather than illumination changes.</p>
<p>For HR computation, Figure <xref ref-type="fig" rid="F7">7</xref>C shows that the implemented peak detection method is more robust than the PSD method (<italic>t</italic>&#x02009;&#x0003D;&#x02009;2.52; <italic>p</italic>&#x02009;&#x0003D;&#x02009;0.03). Peak detection reduces the error by averaging all the IBI over the video duration. For PSD, once the noise takes the dominant frequency and lies in the human HR range, there is no solution to detect the right HR. The best result from each stage is shown in Table <xref ref-type="table" rid="T2">2</xref> with extracted skin area, ICA and peak detection method.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Obtained performance for the best method at each stage.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Stage</th>
<th valign="top" align="center">M (SD)/bpm</th>
<th valign="top" align="center">RMSE/bpm</th>
<th valign="top" align="center">&#x003C1;</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Face video processing (extracted skin area)</td>
<td align="center" valign="top">5.34 (14.98)</td>
<td align="center" valign="top">15.05</td>
<td align="center" valign="top">0.20</td>
</tr>
<tr>
<td align="left" valign="top">Blood volume pulse signal extraction (independent component analysis)</td>
<td align="center" valign="top">4.09 (13.37)</td>
<td align="center" valign="top">13.56</td>
<td align="center" valign="top">0.32</td>
</tr>
<tr>
<td align="left" valign="top">Heart rate computing (peak detection)</td>
<td align="center" valign="top">3.01 (12.14)</td>
<td align="center" valign="top">12.23</td>
<td align="center" valign="top">0.55</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Performance is measured by M (mean error), SD (standard derivation), RMSE (root mean squared error), and &#x003C1; (correlation coefficient)</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="S4-2">
<title>Results From Complete Methods</title>
<p>The results of three methods tested on MAHNOB-HCI are shown in Table <xref ref-type="table" rid="T3">3</xref>. Unsurprisingly the performance of Poh et al. (<xref ref-type="bibr" rid="B40">2011</xref>)&#x02019;s method is significantly dropped than its self-reported results using the proprietary dataset (mean bias is 0.64, RMSE is 4.63, and correlation coefficient is 0.95). This is probably due to the fact that the original dataset is rather stationary and avoids head motions. Li et al. (<xref ref-type="bibr" rid="B29">2014</xref>) and Lam and Yoshinori (<xref ref-type="bibr" rid="B27">2015</xref>) both tested the Poh et al. (<xref ref-type="bibr" rid="B40">2011</xref>) method on MAHNOB-HCI dataset. Our testing result is a bit worse than Li et al. (<xref ref-type="bibr" rid="B29">2014</xref>) whose mean bias is 2.04, RMSE is 13.6, and correlation coefficient is 0.36, but better than the result from Lam and Yoshinori (<xref ref-type="bibr" rid="B27">2015</xref>) (RMSE is 21.3). Following Li et al. (<xref ref-type="bibr" rid="B29">2014</xref>) method, we obtained better performance than Lam and Yoshinori (<xref ref-type="bibr" rid="B27">2015</xref>), but worse results than that presented by Li et al. (<xref ref-type="bibr" rid="B29">2014</xref>) themselves. Although evaluations are all based on the MAHNOB-HCI dataset, samples, and algorithm parameters are not exactly the same. As to the method from Osman et al. (<xref ref-type="bibr" rid="B37">2015</xref>), we cannot really compare the results since it was tested on self-collected dataset only. It shows better result than Poh et al. (<xref ref-type="bibr" rid="B40">2011</xref>), but not as good as Li&#x02019;s method. From our test, the method of Li et al. (<xref ref-type="bibr" rid="B29">2014</xref>) performs significantly better than Poh et al. (<xref ref-type="bibr" rid="B40">2011</xref>) (<italic>t</italic>&#x02009;&#x0003D;&#x02009;5.00; <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.001) and Osman et al. (<xref ref-type="bibr" rid="B37">2015</xref>) (<italic>t</italic>&#x02009;&#x0003D;&#x02009;4.51; <italic>p</italic>&#x02009;&#x0003C;&#x02009;0.001).</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Obtained performance for the complete methods.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Method</th>
<th valign="top" align="center">M (SD)/bpm</th>
<th valign="top" align="center">RMSE/bpm</th>
<th valign="top" align="center">&#x003C1;</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">Poh et al. (<xref ref-type="bibr" rid="B40">2011</xref>)</td>
<td align="center" valign="top">4.07 (13.04)</td>
<td align="center" valign="top">13.81</td>
<td align="center" valign="top">0.28</td>
</tr>
<tr>
<td align="left" valign="top">Li et al. (<xref ref-type="bibr" rid="B29">2014</xref>)</td>
<td align="center" valign="top">2.15 (10.04)</td>
<td align="center" valign="top">10.33</td>
<td align="center" valign="top">0.68</td>
</tr>
<tr>
<td align="left" valign="top">Osman et al. (<xref ref-type="bibr" rid="B37">2015</xref>)</td>
<td align="center" valign="top">3.37 (12.08)</td>
<td align="center" valign="top">12.79</td>
<td align="center" valign="top">0.47</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Performance is measured by M (mean error), SD (standard derivation), RMSE (root mean squared error), and &#x003C1; (correlation coefficient)</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>We can see from Tables <xref ref-type="table" rid="T2">2</xref> and <xref ref-type="table" rid="T3">3</xref> that the face segmentation influences results significantly. With the best performed method of each stage, it can achieve competitive results compared with complete state-of-the-art methods which are more complex. Thus, we obtained and proposed an efficient pipeline with extracted skin area, ICA and peak detection for detecting HR remotely from face videos.</p>
<p>This pipeline could be applied for other studies with front face recordings under environmental illumination. It is robust with facial expressions and limited head motions (translation and orientation), which is often the case for the majority of human&#x02013;computer interaction processes (online education, computer gaming, etc.). The pipeline is expected to perform better under more stable conditions with less head movements. Interruptions such as hands partially covering the face will influence the performance of this method. Furthermore, if the illumination changes significantly during the HR detection process, noise estimation methods could be applied to improve the overall performance.</p>
<p>The main limitation of this study comes specifically from the selected database and testing methods. Possible risks of bias can derive from the small number of subjects (24 participants) and the selected data source (444 videos and corresponding ECG signals from MAHNOB dataset). The specific experimental setting could favor some of the considered methods and could potentially be detrimental for others. For example, the forehead area is reported have a good performance (Lewandowska et al., <xref ref-type="bibr" rid="B28">2011</xref>; Stricker et al., <xref ref-type="bibr" rid="B50">2014</xref>), while our results showed that this was not significant, potentially due to the EEG head cap which partially covers the forehead area in MAHNOB videos.</p>
</sec>
</sec>
<sec id="S5">
<title>Conclusion</title>
<p>Remote HR measurements from face videos have improved during the last few years. Among the research in this domain, the designed models, parameter settings, chosen algorithms, and equipment are plenty, complex, and vary enormously. Some approaches achieve high accuracy under well-controlled situations but degrade with illumination changes and head motions. In this article, we performed (a) the collection and classification of state-of-the-art methods into three stages and (b) the comparison of their performance under HCI conditions.</p>
<p>The MAHNOB-HCI dataset is used for algorithm testing and analysis since it is a publicly accessible dataset. Our results showed that at the pre-processing stage, accurate face detection algorithms performed better than rough ROI detection. The extracted facial skin area used as the source HR signal obtained better result than any other facial ROI. For signal extraction, under most cases, ICA method obtained decent results. When the background is monotone, removing the noise estimated from the background increased the HR detection accuracy efficiently. As for post-processing, peak detection in time domain was more reliable than the frequency domain methods. In conclusion, we built an efficient pipeline for non-intrusive HR detection from face videos by combining the methods we found to be the best. This pipeline with skin area extraction, ICA and peak detection demonstrated a state-of-the-art accuracy.</p>
<p>Though considerable progress has been made in this domain, there are still many difficulties. The state-of-the-art approaches are not robust enough when applied under natural conditions and are still unable to detect HR in real-time. Technically speaking, machine learning methods may be promising for remote HR detection. Especially with finger oximeter as ground truth measurement, since the BVP signal detected from facial regions should have a similar shape with collected finger or ear BVP signals. So far there are only three papers using machine learning methods and both of them concentrate on the post-processing category. With the development of deep end-to-end learning, the robustness and accuracy of HR detection may increase efficiently even under naturalistic situations.</p>
<p>Instantaneous HR reveals more information about the subject&#x02019;s physiological and affective status than mean HR. According to Malik (<xref ref-type="bibr" rid="B31">1996</xref>) and Armony and Vuilleumier (<xref ref-type="bibr" rid="B3">2013</xref>), the heart beat variability can be used for investigating mental workload or to detect emotions such as anger and sadness. The work from Lakens (<xref ref-type="bibr" rid="B26">2013</xref>) has already shown the possibility of using smartphones to measure HR variation associated with relived experiences of anger and happiness. Thus psychological insights of HR changes and the link to affective status can be taken into consideration for further study. In short, more efforts could be devoted to reliable instantaneous HR detection in realistic scenarios.</p>
</sec>
<sec id="S6" sec-type="author-contributor">
<title>Author Contributions</title>
<p>CW designed and implemented the methods. CW and GC analyzed and interpreted the results. All authors contributed to the redaction of the manuscript.</p>
</sec>
<sec id="S7">
<title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<ack>
<p>The authors would like to show the gratitude to Louis Philippe Simoes Castelo Branco, Computer Vision and Multimedia Laboratory, Computer Science Department, University of Geneva, for reviewing the early version of the manuscript. His corrections and suggestions greatly improved the quality of the manuscript.</p>
</ack>
<fn-group>
<fn fn-type="financial-disclosure">
<p><bold>Funding.</bold> This work is supported by the Swiss National Science Foundation under the grant No. 200021E - 164326.</p></fn>
</fn-group>
<sec id="S9" sec-type="supplementary-material">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at <uri xlink:href="https://www.frontiersin.org/articles/10.3389/fbioe.2018.00033/full&#x00023;supplementary-material">https://www.frontiersin.org/articles/10.3389/fbioe.2018.00033/full&#x00023;supplementary-material</uri>.</p>
<supplementary-material xlink:href="Table_1.docx" id="SM1" mimetype="applicationn/docx" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abdi</surname> <given-names>H.</given-names></name> <name><surname>Williams</surname> <given-names>L. J.</given-names></name></person-group> (<year>2010</year>). <article-title>Principal component analysis</article-title>. <source>Wiley Interdiscip. Rev.</source> <volume>2</volume>, <fpage>433</fpage>&#x02013;<lpage>459</lpage>.<pub-id pub-id-type="doi">10.1002/wics.101</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Allen</surname> <given-names>J.</given-names></name></person-group> (<year>2007</year>). <article-title>Photoplethysmography and its application in clinical physiological measurement</article-title>. <source>Physiol. Meas.</source> <volume>28</volume>, <fpage>R1</fpage>.<pub-id pub-id-type="doi">10.1088/0967-3334/28/3/R01</pub-id><pub-id pub-id-type="pmid">17322588</pub-id></citation></ref>
<ref id="B3"><citation citation-type="book"><person-group person-group-type="editor"><name><surname>Armony</surname> <given-names>J.</given-names></name> <name><surname>Vuilleumier</surname> <given-names>P.</given-names></name></person-group> (eds) (<year>2013</year>). <source>The Cambridge Handbook of Human Affective Neuroscience</source>. <publisher-name>Cambridge University Press</publisher-name>.</citation></ref>
<ref id="B4"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Balakrishnan</surname> <given-names>G.</given-names></name> <name><surname>Durand</surname> <given-names>F.</given-names></name> <name><surname>Guttag</surname> <given-names>J.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;Detecting pulse from head motions in video,&#x0201D;</article-title> in <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>. <conf-loc>Portland, Oregon</conf-loc>.</citation></ref>
<ref id="B5"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Baltru&#x00161;aitis</surname> <given-names>T.</given-names></name> <name><surname>Robinson</surname> <given-names>P.</given-names></name> <name><surname>Morency</surname> <given-names>L.-P.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Openface: an open source facial behavior analysis toolkit,&#x0201D;</article-title> in <conf-name>Applications of Computer Vision (WACV), 2016 IEEE Winter Conference on (IEEE)</conf-name>.</citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cardoso</surname> <given-names>J.-F.</given-names></name></person-group> (<year>1999</year>). <article-title>High-order contrasts for independent component analysis</article-title>. <source>Neural Computat.</source> <volume>11</volume>, <fpage>157</fpage>&#x02013;<lpage>192</lpage>.<pub-id pub-id-type="doi">10.1162/089976699300016863</pub-id><pub-id pub-id-type="pmid">9950728</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cennini</surname> <given-names>G.</given-names></name> <name><surname>Arguel</surname> <given-names>J.</given-names></name> <name><surname>Ak&#x0015F;it</surname> <given-names>K.</given-names></name> <name><surname>van Leest</surname> <given-names>A.</given-names></name></person-group> (<year>2010</year>). <article-title>Heart rate monitoring via remote photoplethysmography with motion artifacts reduction</article-title>. <source>Opt. Exp.</source> <volume>18</volume>, <fpage>4867</fpage>&#x02013;<lpage>4875</lpage>.<pub-id pub-id-type="doi">10.1364/OE.18.004867</pub-id><pub-id pub-id-type="pmid">20389499</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chan</surname> <given-names>K. W.</given-names></name> <name><surname>Zhang</surname> <given-names>Y. T.</given-names></name></person-group> (<year>2002</year>). <article-title>Adaptive reduction of motion artifact from photoplethysmographic recordings using a variable step-size LMS filter</article-title>. <source>Sensors</source> <volume>2</volume>, <fpage>1343</fpage>&#x02013;<lpage>1346</lpage>.<pub-id pub-id-type="doi">10.1109/ICSENS.2002.1037314</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>D.-Y.</given-names></name> <name><surname>Wang</surname> <given-names>J. J.</given-names></name> <name><surname>Lin</surname> <given-names>K. Y.</given-names></name> <name><surname>Chang</surname> <given-names>H. H.</given-names></name> <name><surname>Wu</surname> <given-names>H. K.</given-names></name> <name><surname>Chen</surname> <given-names>Y. S.</given-names></name> <etal/></person-group> (<year>2015</year>). <article-title>Image sensor-based heart rate evaluation from face reflectance using hilbert&#x02013;huang transform</article-title>. <source>IEEE Sens. J.</source> <volume>15</volume>, <fpage>618</fpage>&#x02013;<lpage>627</lpage>.<pub-id pub-id-type="doi">10.1109/JSEN.2014.2347397</pub-id></citation></ref>
<ref id="B10"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>W.</given-names></name> <name><surname>Picard</surname> <given-names>R. W.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Eliminating physiological information from facial videos,&#x0201D;</article-title> in <conf-name>Automatic Face &#x00026; Gesture Recognition (FG 2017), 2017 12th IEEE International Conference on (IEEE)</conf-name>.</citation></ref>
<ref id="B11"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Da</surname> <given-names>H.</given-names></name> <name><surname>Winokur</surname> <given-names>E. S.</given-names></name> <name><surname>Sodini</surname> <given-names>C. G.</given-names></name></person-group> (<year>2011</year>). <article-title>&#x0201C;A continuous, wearable, and wireless heart monitor using head ballistocardiogram (BCG) and head electrocardiogram (ECG),&#x0201D;</article-title> in <conf-name>2011 Annual International Conference of the IEEE Engineering in Medicine and Biology Society</conf-name> (<conf-loc>Boston, Massachusetts</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>).</citation></ref>
<ref id="B12"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Datcu</surname> <given-names>D.</given-names></name> <name><surname>Cidota</surname> <given-names>M.</given-names></name> <name><surname>Lukosch</surname> <given-names>S.</given-names></name> <name><surname>Rothkrantz</surname> <given-names>L.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;Noncontact automatic heart rate analysis in visible spectrum by specific face regions,&#x0201D;</article-title> in <conf-name>Proceedings of the 14th International Conference on Computer Systems and Technologies</conf-name> (<conf-loc>New York</conf-loc>: <conf-sponsor>ACM</conf-sponsor>).</citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dawson</surname> <given-names>J. A.</given-names></name> <name><surname>Kamlin</surname> <given-names>C. O. F.</given-names></name> <name><surname>Wong</surname> <given-names>C.</given-names></name> <name><surname>Te Pas</surname> <given-names>A. B.</given-names></name> <name><surname>Vento</surname> <given-names>M.</given-names></name> <name><surname>Cole</surname> <given-names>T. J.</given-names></name> <etal/></person-group> (<year>2010</year>). <article-title>Changes in heart rate in the first minutes after birth</article-title>. <source>Arch. Dis. Childhood Fetal Neonatal Ed.</source> <volume>95</volume>, <fpage>F177</fpage>&#x02013;<lpage>F181</lpage>.<pub-id pub-id-type="doi">10.1136/adc.2009.169102</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>De Haan</surname> <given-names>G.</given-names></name> <name><surname>Vincent</surname> <given-names>J.</given-names></name></person-group> (<year>2013</year>). <article-title>Robust pulse rate from chrominance-based rPPG</article-title>. <source>IEEE Transac. Biomed. Eng.</source> <volume>60</volume>, <fpage>2878</fpage>&#x02013;<lpage>2886</lpage>.<pub-id pub-id-type="doi">10.1109/TBME.2013.2266196</pub-id><pub-id pub-id-type="pmid">23744659</pub-id></citation></ref>
<ref id="B15"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>De la Torre</surname> <given-names>F.</given-names></name> <name><surname>Chu</surname> <given-names>W.-S.</given-names></name> <name><surname>Xiong</surname> <given-names>X.</given-names></name> <name><surname>Vicente</surname> <given-names>F.</given-names></name> <name><surname>Ding</surname> <given-names>X.</given-names></name> <name><surname>Jeffrey Cohn</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;IntraFace,&#x0201D;</article-title> in <conf-name>Automatic Face and Gesture Recognition (FG), 2015 11th IEEE International Conference and Workshops on</conf-name>, Vol. <volume>1</volume> (<conf-loc>Pittsburgh</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>).</citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Garbey</surname> <given-names>M.</given-names></name> <name><surname>Sun</surname> <given-names>N.</given-names></name> <name><surname>Merla</surname> <given-names>A.</given-names></name> <name><surname>Pavlidis</surname> <given-names>I.</given-names></name></person-group> (<year>2007</year>). <article-title>Contact-free measurement of cardiac pulse based on the analysis of thermal imagery</article-title>. <source>IEEE Trans. Biomed. Eng.</source> <volume>54</volume>, <fpage>1418</fpage>&#x02013;<lpage>1426</lpage>.<pub-id pub-id-type="doi">10.1109/TBME.2007.891930</pub-id><pub-id pub-id-type="pmid">17694862</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hjelm&#x000E5;s</surname> <given-names>E.</given-names></name> <name><surname>Low</surname> <given-names>B. K.</given-names></name></person-group> (<year>2001</year>). <article-title>Face detection: a survey</article-title>. <source>Comput. Vis. Image Understand.</source> <volume>83</volume>, <fpage>236</fpage>&#x02013;<lpage>274</lpage>.<pub-id pub-id-type="doi">10.1006/cviu.2001.0921</pub-id></citation></ref>
<ref id="B18"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Huelsbusch</surname> <given-names>M.</given-names></name> <name><surname>Blazek</surname> <given-names>V.</given-names></name></person-group> (<year>2002</year>). <article-title>&#x0201C;Contactless mapping of rhythmical phenomena in tissue perfusion using PPGI,&#x0201D;</article-title> in <conf-name>Proc. SPIE 4683, Medical Imaging 2002: Physiology and Function from Multidimensional Images</conf-name>, Vol. <volume>110</volume>, <conf-loc>San Diego, CA</conf-loc>.</citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hyv&#x000E4;rinen</surname> <given-names>A.</given-names></name> <name><surname>Oja</surname> <given-names>E.</given-names></name></person-group> (<year>2000</year>). <article-title>Independent component analysis: algorithms and applications</article-title>. <source>Neural Networks</source> <volume>13</volume>, <fpage>411</fpage>&#x02013;<lpage>430</lpage>.<pub-id pub-id-type="doi">10.1016/S0893-6080(00)00026-5</pub-id></citation></ref>
<ref id="B20"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Irani</surname> <given-names>R.</given-names></name> <name><surname>Nasrollahi</surname> <given-names>K.</given-names></name> <name><surname>Moeslund</surname> <given-names>T. B.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;Improved pulse detection from head motions using DCT,&#x0201D;</article-title> in <conf-name>Computer Vision Theory and Applications (VISAPP), 2014 International Conference on</conf-name>, Vol. <volume>3</volume> (<conf-loc>Lisbon, Portugal</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>).</citation></ref>
<ref id="B21"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Jeanne</surname> <given-names>V.</given-names></name> <name><surname>Asselman</surname> <given-names>M.</given-names></name> <name><surname>den Brinker</surname> <given-names>B.</given-names></name> <name><surname>Bulut</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;Camera-based heart rate monitoring in highly dynamic light conditions,&#x0201D;</article-title> in <conf-name>2013 International Conference on Connected Vehicles and Expo (ICCVE)</conf-name> (<conf-loc>Las Vegas, NV</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>).</citation></ref>
<ref id="B22"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Jensen</surname> <given-names>J. N.</given-names></name> <name><surname>Hannemose</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <source>Camera-based Heart Rate Monitoring</source>. <publisher-loc>Lyngby, Denmark</publisher-loc>: <publisher-name>Department of Applied Mathematics and Computer Science, DTU Computer</publisher-name>, <fpage>17</fpage>.</citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kakumanu</surname> <given-names>P.</given-names></name> <name><surname>Makrogiannis</surname> <given-names>S.</given-names></name> <name><surname>Bourbakis</surname> <given-names>N.</given-names></name></person-group> (<year>2007</year>). <article-title>A survey of skin-color modeling and detection methods</article-title>. <source>Pattern Recognit.</source> <volume>40</volume>, <fpage>1106</fpage>&#x02013;<lpage>1122</lpage>.<pub-id pub-id-type="doi">10.1016/j.patcog.2006.06.010</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kumar</surname> <given-names>M.</given-names></name> <name><surname>Veeraraghavan</surname> <given-names>A.</given-names></name> <name><surname>Sabharwal</surname> <given-names>A.</given-names></name></person-group> (<year>2015</year>). <article-title>DistancePPG: robust non-contact vital signs monitoring using a camera</article-title>. <source>Biomed. Opt. Exp.</source> <volume>6</volume>, <fpage>1565</fpage>&#x02013;<lpage>1588</lpage>.<pub-id pub-id-type="doi">10.1364/BOE.6.001565</pub-id><pub-id pub-id-type="pmid">26137365</pub-id></citation></ref>
<ref id="B25"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Kwon</surname> <given-names>S.</given-names></name> <name><surname>Kim</surname> <given-names>H.</given-names></name> <name><surname>Suk Park</surname> <given-names>K.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Validation of heart rate extraction using video imaging on a built-in camera system of a smartphone,&#x0201D;</article-title> in <conf-name>2012 Annual International Conference of the IEEE Engineering in Medicine and Biology Society</conf-name> (<conf-loc>San Diego, CA</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>).</citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lakens</surname> <given-names>D.</given-names></name></person-group> (<year>2013</year>). <article-title>Using a smartphone to measure heart rate changes during relived happiness and anger</article-title>. <source>Trans. Affect. Comput.</source> <volume>4</volume>, <fpage>238</fpage>&#x02013;<lpage>241</lpage>.<pub-id pub-id-type="doi">10.1109/T-AFFC.2013.3</pub-id></citation></ref>
<ref id="B27"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Lam</surname> <given-names>A.</given-names></name> <name><surname>Yoshinori</surname> <given-names>K.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Robust heart rate measurement from video using select random patches,&#x0201D;</article-title> in <conf-name>Proceedings of the IEEE International Conference on Computer Vision</conf-name>. <conf-loc>Santiago, Chile</conf-loc>.</citation></ref>
<ref id="B28"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Lewandowska</surname> <given-names>M.</given-names></name> <name><surname>Rumi&#x00144;ski</surname> <given-names>J.</given-names></name> <name><surname>Kocejko</surname> <given-names>T.</given-names></name> <name><surname>Nowak</surname> <given-names>J.</given-names></name></person-group> (<year>2011</year>). <article-title>&#x0201C;Measuring pulse rate with a webcam&#x02014;a non-contact method for evaluating cardiac activity,&#x0201D;</article-title> in <conf-name>Computer Science and Information Systems (FedCSIS), 2011 Federated Conference on</conf-name> (<conf-loc>Szczecin, Poland</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>).</citation></ref>
<ref id="B29"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Chen</surname> <given-names>J.</given-names></name> <name><surname>Zhao</surname> <given-names>G.</given-names></name> <name><surname>Pietikainen</surname> <given-names>M.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;Remote heart rate measurement from face videos under realistic situations,&#x0201D;</article-title> in <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>. <conf-loc>Columbus, Ohio</conf-loc>.</citation></ref>
<ref id="B30"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Liao</surname> <given-names>X.</given-names></name> <name><surname>Carin</surname> <given-names>L.</given-names></name></person-group> (<year>2002</year>). <article-title>&#x0201C;A new algorithm for independent component analysis with or without constraints,&#x0201D;</article-title> in <conf-name>Sensor Array and Multichannel Signal Processing Workshop Proceedings, 2002</conf-name> (<conf-loc>Rosslyn, VA</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>).</citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Malik</surname> <given-names>M.</given-names></name></person-group> (<year>1996</year>). <article-title>Heart rate variability</article-title>. <source>Ann. Noninvasive Electrocardiol.</source> <volume>1</volume>, <fpage>151</fpage>&#x02013;<lpage>181</lpage>.<pub-id pub-id-type="doi">10.1111/j.1542-474X.1996.tb00275.x</pub-id></citation></ref>
<ref id="B32"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>McDuff</surname> <given-names>D. J.</given-names></name> <name><surname>Blackford</surname> <given-names>E. B.</given-names></name> <name><surname>Estepp</surname> <given-names>J. R.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;The impact of video compression on remote cardiac pulse measurement using imaging photoplethysmography,&#x0201D;</article-title> in <conf-name>Automatic Face &#x00026; Gesture Recognition (FG 2017), 2017 12th IEEE International Conference on (IEEE)</conf-name>.</citation></ref>
<ref id="B33"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Mestha</surname> <given-names>L. K.</given-names></name> <name><surname>Kyal</surname> <given-names>S.</given-names></name> <name><surname>Xu</surname> <given-names>B.</given-names></name> <name><surname>Lewis</surname> <given-names>L. E.</given-names></name> <name><surname>Kumar</surname> <given-names>V.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;Towards continuous monitoring of pulse rate in neonatal intensive care unit with a webcam,&#x0201D;</article-title> in <conf-name>2014 36th Annual International Conference of the IEEE Engineering in Medicine and Biology Society</conf-name> (<conf-loc>Chicago, IL</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>).</citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Monkaresi</surname> <given-names>H.</given-names></name> <name><surname>Calvo</surname> <given-names>R. A.</given-names></name> <name><surname>Yan</surname> <given-names>H.</given-names></name></person-group> (<year>2014</year>). <article-title>A machine learning approach to improve contactless heart rate monitoring using a webcam</article-title>. <source>IEEE J Biomed. Health Inform.</source> <volume>18</volume>, <fpage>1153</fpage>&#x02013;<lpage>1160</lpage>.<pub-id pub-id-type="doi">10.1109/JBHI.2013.2291900</pub-id><pub-id pub-id-type="pmid">25014930</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moreno</surname> <given-names>J.</given-names></name> <name><surname>Ramos-Castro</surname> <given-names>J.</given-names></name> <name><surname>Movellan</surname> <given-names>J.</given-names></name> <name><surname>Parrado</surname> <given-names>E.</given-names></name> <name><surname>Rodas</surname> <given-names>G.</given-names></name> <name><surname>Capdevila</surname> <given-names>L.</given-names></name></person-group> (<year>2015</year>). <article-title>Facial video-based photoplethysmography to detect HRV at rest</article-title>. <source>Int. J. Sports Med.</source> <volume>36</volume>, <fpage>474</fpage>&#x02013;<lpage>480</lpage>.<pub-id pub-id-type="doi">10.1055/s-0034-1398530</pub-id></citation></ref>
<ref id="B36"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Muender</surname> <given-names>T.</given-names></name> <name><surname>Miller</surname> <given-names>M. K.</given-names></name> <name><surname>Birk</surname> <given-names>M. V.</given-names></name> <name><surname>Mandryk</surname> <given-names>R. L.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Extracting heart rate from videos of online participants,&#x0201D;</article-title> in <conf-name>Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (CHI&#x02019;2016)</conf-name>. <conf-loc>San Jose, CA</conf-loc>.</citation></ref>
<ref id="B37"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Osman</surname> <given-names>A.</given-names></name> <name><surname>Turcot</surname> <given-names>J.</given-names></name> <name><surname>El Kaliouby</surname> <given-names>R.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Supervised learning approach to remote heart rate estimation from facial videos,&#x0201D;</article-title> in <conf-name>Automatic Face and Gesture Recognition (FG), 2015 11th IEEE International Conference and Workshops on</conf-name>, Vol. <volume>1</volume> (<conf-loc>Washington</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>).</citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pavlidis</surname> <given-names>I.</given-names></name> <name><surname>Dowdall</surname> <given-names>J.</given-names></name> <name><surname>Sun</surname> <given-names>N.</given-names></name> <name><surname>Puri</surname> <given-names>C.</given-names></name> <name><surname>Fei</surname> <given-names>J.</given-names></name> <name><surname>Garbey</surname> <given-names>M.</given-names></name></person-group> (<year>2007</year>). <article-title>Interacting with human physiology</article-title>. <source>Comput. Vis. Img. Understand.</source> <volume>108</volume>, <fpage>150</fpage>&#x02013;<lpage>170</lpage>.<pub-id pub-id-type="doi">10.1016/j.cviu.2006.11.018</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Poh</surname> <given-names>M.-Z.</given-names></name> <name><surname>McDuff</surname> <given-names>D. J.</given-names></name> <name><surname>Picard</surname> <given-names>R. W.</given-names></name></person-group> (<year>2010</year>). <article-title>Non-contact, automated cardiac pulse measurements using video imaging and blind source separation</article-title>. <source>Opt. Exp.</source> <volume>18</volume>, <fpage>10762</fpage>&#x02013;<lpage>10774</lpage>.<pub-id pub-id-type="doi">10.1364/OE.18.010762</pub-id><pub-id pub-id-type="pmid">20588929</pub-id></citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Poh</surname> <given-names>M.-Z.</given-names></name> <name><surname>McDuff</surname> <given-names>D. J.</given-names></name> <name><surname>Picard</surname> <given-names>R. W.</given-names></name></person-group> (<year>2011</year>). <article-title>Advancements in noncontact, multiparameter physiological measurements using a webcam</article-title>. <source>IEEE Trans. Biomed. Eng.</source> <volume>58</volume>, <fpage>7</fpage>&#x02013;<lpage>11</lpage>.<pub-id pub-id-type="doi">10.1109/TBME.2010.2086456</pub-id><pub-id pub-id-type="pmid">20952328</pub-id></citation></ref>
<ref id="B41"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Pursche</surname> <given-names>T.</given-names></name> <name><surname>Krajewski</surname> <given-names>J.</given-names></name> <name><surname>Moeller</surname> <given-names>R.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Video-based heart rate measurement from human faces,&#x0201D;</article-title> in <conf-name>2012 IEEE International Conference on Consumer Electronics (ICCE)</conf-name> (<conf-loc>Berlin</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>).</citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rother</surname> <given-names>C.</given-names></name> <name><surname>Kolmogorov</surname> <given-names>V.</given-names></name> <name><surname>Blake</surname> <given-names>A.</given-names></name></person-group> (<year>2004</year>). <article-title>Grabcut: interactive foreground extraction using iterated graph cuts</article-title>. <source>ACM Trans. Graph.</source> <volume>23</volume>, <fpage>309</fpage>&#x02013;<lpage>314</lpage>.<pub-id pub-id-type="doi">10.1145/1015706.1015720</pub-id></citation></ref>
<ref id="B43"><citation citation-type="thesis"><person-group person-group-type="author"><name><surname>Ruben</surname> <given-names>N. E.</given-names></name></person-group> (<year>2015</year>). <source>Remote Heart Rate Estimation Using Consumer-Grade Cameras</source> [Dissertation]. <publisher-loc>Logan, UT</publisher-loc>: <publisher-name>Utah State University</publisher-name>.</citation></ref>
<ref id="B44"><citation citation-type="thesis"><person-group person-group-type="author"><name><surname>Rubinstein</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <source>Analysis and Visualization of Temporal Variations in Video</source> [Dissertation]. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>Massachusetts Institute of Technology</publisher-name>.</citation></ref>
<ref id="B45"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Sahindrakar</surname> <given-names>P.</given-names></name> <name><surname>de Haan</surname> <given-names>G.</given-names></name> <name><surname>Kirenko</surname> <given-names>I.</given-names></name></person-group> (<year>2011</year>). <source>Improving Motion Robustness of Contact-Less Monitoring of Heart Rate Using Video Analysis</source>. <publisher-loc>Eindhoven, The Netherlands</publisher-loc>: <publisher-name>Technische Universiteit Eindhoven, Department of Mathematics and Computer Science</publisher-name>.</citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Saragih</surname> <given-names>J. M.</given-names></name> <name><surname>Lucey</surname> <given-names>S.</given-names></name> <name><surname>Cohn</surname> <given-names>J. F.</given-names></name></person-group> (<year>2011</year>). <article-title>Deformable model fitting by regularized landmark mean-shift</article-title>. <source>Int. J. Comput. Vis.</source> <volume>91</volume>, <fpage>200</fpage>&#x02013;<lpage>215</lpage>.<pub-id pub-id-type="doi">10.1007/s11263-010-0380-4</pub-id></citation></ref>
<ref id="B47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Soleymani</surname> <given-names>M.</given-names></name> <name><surname>Lichtenauer</surname> <given-names>J.</given-names></name> <name><surname>Pun</surname> <given-names>T.</given-names></name> <name><surname>Pantic</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>A multimodal database for affect recognition and implicit tagging</article-title>. <source>IEEE Trans. Affect. Comput.</source> <volume>3</volume>, <fpage>42</fpage>&#x02013;<lpage>55</lpage>.<pub-id pub-id-type="doi">10.1109/T-AFFC.2011.25</pub-id></citation></ref>
<ref id="B48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Soleymani</surname> <given-names>M.</given-names></name> <name><surname>Villaro-Dixon</surname> <given-names>F.</given-names></name> <name><surname>Pun</surname> <given-names>T.</given-names></name> <name><surname>Chanel</surname> <given-names>G.</given-names></name></person-group> (<year>2017</year>). <article-title>Toolbox for emotional feAture extraction from physiological signals (TEAP)</article-title>. <source>Front. ICT</source> <volume>4</volume>:<fpage>1</fpage>.<pub-id pub-id-type="doi">10.3389/fict.2017.00001</pub-id></citation></ref>
<ref id="B49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Starr</surname> <given-names>I.</given-names></name> <name><surname>Rawson</surname> <given-names>A. J.</given-names></name> <name><surname>Schroeder</surname> <given-names>H. A.</given-names></name> <name><surname>Joseph</surname> <given-names>N. R.</given-names></name></person-group> (<year>1939</year>). <article-title>Studies on the estimation of cardiac output in man, and of abnormalities in cardiac functions, from the heart&#x02019;s recoil and the blood&#x02019;s impacts; the ballistocardiogram</article-title>. <source>Am. J. Physiol.</source> <volume>127</volume>, <fpage>1</fpage>&#x02013;<lpage>28</lpage>.</citation></ref>
<ref id="B50"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Stricker</surname> <given-names>R.</given-names></name> <name><surname>M&#x000FC;ller</surname> <given-names>S.</given-names></name> <name><surname>Gross</surname> <given-names>H. M.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;Non-contact video-based pulse rate measurement on a mobile service robot&#x0201D;</article-title>, in <conf-name>Robot and Human Interactive Communication, 2014 RO-MAN: The 23rd IEEE International Symposium on</conf-name> (<conf-loc>Edinburgh, Scotland</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>), p. <fpage>1056</fpage>&#x02013;<lpage>1062</lpage>.</citation></ref>
<ref id="B51"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>Y.</given-names></name> <name><surname>Hu</surname> <given-names>S.</given-names></name> <name><surname>Azorin-Peris</surname> <given-names>V.</given-names></name> <name><surname>Kalawsky</surname> <given-names>R.</given-names></name> <name><surname>Greenwald</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <article-title>Noncontact imaging photoplethysmography to effectively access pulse rate variability</article-title>. <source>J. Biomed. Opt.</source> <volume>18</volume>, <fpage>061205</fpage>&#x02013;<lpage>061205</lpage>.<pub-id pub-id-type="doi">10.1117/1.JBO.18.6.061205</pub-id><pub-id pub-id-type="pmid">23111602</pub-id></citation></ref>
<ref id="B52"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>Y.</given-names></name> <name><surname>Papin</surname> <given-names>C.</given-names></name> <name><surname>Azorin-Peris</surname> <given-names>V.</given-names></name> <name><surname>Kalawsky</surname> <given-names>R.</given-names></name> <name><surname>Greenwald</surname> <given-names>S.</given-names></name> <name><surname>Hu</surname> <given-names>S.</given-names></name></person-group> (<year>2012</year>). <article-title>Use of ambient light in remote photoplethysmographic systems: comparison between a high-performance camera and a low-cost webcam</article-title>. <source>J. Biomed. Opt.</source> <volume>17</volume>, <fpage>0370051</fpage>&#x02013;<lpage>03700510</lpage>.<pub-id pub-id-type="doi">10.1117/1.JBO.17.3.037005</pub-id><pub-id pub-id-type="pmid">22502577</pub-id></citation></ref>
<ref id="B53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tarassenko</surname> <given-names>L.</given-names></name> <name><surname>Villarroel</surname> <given-names>M.</given-names></name> <name><surname>Guazzi</surname> <given-names>A.</given-names></name> <name><surname>Jorge</surname> <given-names>J.</given-names></name> <name><surname>Clifton</surname> <given-names>D. A.</given-names></name> <name><surname>Pugh</surname> <given-names>C.</given-names></name></person-group> (<year>2014</year>). <article-title>Non-contact video-based vital sign monitoring using ambient light and auto-regressive models</article-title>. <source>Physiol. Meas.</source> <volume>35</volume>, <fpage>807</fpage>.<pub-id pub-id-type="doi">10.1049/htl.2014.0077</pub-id><pub-id pub-id-type="pmid">24681430</pub-id></citation></ref>
<ref id="B54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tarvainen</surname> <given-names>M. P.</given-names></name> <name><surname>Ranta-Aho</surname> <given-names>P. O.</given-names></name> <name><surname>Karjalainen</surname> <given-names>P. A.</given-names></name></person-group> (<year>2002</year>). <article-title>An advanced detrending method with application to HRV analysis</article-title>. <source>IEEE Trans. Biomed. Eng.</source> <volume>49</volume>, <fpage>172</fpage>&#x02013;<lpage>175</lpage>.<pub-id pub-id-type="doi">10.1109/10.979357</pub-id><pub-id pub-id-type="pmid">12066885</pub-id></citation></ref>
<ref id="B55"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Tompkins</surname> <given-names>W. J.</given-names></name></person-group> (<year>1993</year>). <source>Biomedical Digital Signal Processing: C-language Examples and Laboratory Experiments for the IBM PC. Hauptbd</source>. <publisher-name>Editorial Prentice Hall</publisher-name>.</citation></ref>
<ref id="B56"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Tran</surname> <given-names>D. N.</given-names></name> <name><surname>Lee</surname> <given-names>H.</given-names></name> <name><surname>Kim</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;A robust real time system for remote heart rate measurement via camera,&#x0201D;</article-title> in <conf-name>2015 IEEE International Conference on Multimedia and Expo (ICME)</conf-name> (<conf-loc>Torino, Italy</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>).</citation></ref>
<ref id="B57"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Tulyakov</surname> <given-names>S.</given-names></name> <name><surname>Alameda-Pineda</surname> <given-names>X.</given-names></name> <name><surname>Ricci</surname> <given-names>E.</given-names></name> <name><surname>Yin</surname> <given-names>L.</given-names></name> <name><surname>Cohn</surname> <given-names>J. F.</given-names></name> <name><surname>Sebe</surname> <given-names>N.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Self-adaptive matrix completion for heart rate estimation from face videos under realistic conditions,&#x0201D;</article-title> in <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name> <conf-loc>(Las Vegas, NV: IEEE)</conf-loc>.</citation></ref>
<ref id="B58"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Verkruysse</surname> <given-names>W.</given-names></name> <name><surname>Svaasand</surname> <given-names>L. O.</given-names></name> <name><surname>Stuart Nelson</surname> <given-names>J.</given-names></name></person-group> (<year>2008</year>). <article-title>Remote plethysmographic imaging using ambient light</article-title>. <source>Opt. Exp.</source> <volume>16</volume>, <fpage>21434</fpage>&#x02013;<lpage>21445</lpage>.<pub-id pub-id-type="doi">10.1364/OE.16.021434</pub-id><pub-id pub-id-type="pmid">19104573</pub-id></citation></ref>
<ref id="B59"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vezhnevets</surname> <given-names>V.</given-names></name> <name><surname>Sazonov</surname> <given-names>V.</given-names></name> <name><surname>Andreeva</surname> <given-names>A.</given-names></name></person-group> (<year>2003</year>). <article-title>&#x0201C;A survey on pixel-based skin color detection techniques&#x0201D;</article-title>, in <source>Proc. Graphicon</source>, Vol. <volume>3</volume>, <fpage>85</fpage>&#x02013;<lpage>92</lpage>.</citation></ref>
<ref id="B60"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Villarroel</surname> <given-names>M.</given-names></name> <name><surname>Jorge</surname> <given-names>J.</given-names></name> <name><surname>Pugh</surname> <given-names>C.</given-names></name> <name><surname>Tarassenko</surname> <given-names>L.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Non-contact vital sign monitoring in the clinic,&#x0201D;</article-title> in <conf-name>Automatic Face &#x00026; Gesture Recognition (FG 2017), 2017 12th IEEE International Conference on (IEEE)</conf-name>.</citation></ref>
<ref id="B61"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Viola</surname> <given-names>P.</given-names></name> <name><surname>Jones</surname> <given-names>M.</given-names></name></person-group> (<year>2001</year>). <article-title>&#x0201C;Rapid object detection using a boosted cascade of simple features,&#x0201D;</article-title> in <conf-name>Computer Vision and Pattern Recognition, 2001. CVPR 2001. Proceedings of the 2001 IEEE Computer Society Conference on</conf-name>, Vol. <volume>1</volume> (<conf-loc>Kauai, HI</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>).</citation></ref>
<ref id="B62"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>den Brinker</surname> <given-names>A. C.</given-names></name> <name><surname>Stuijk</surname> <given-names>S.</given-names></name> <name><surname>de Haan</surname> <given-names>G.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Color-distortion filtering for remote photoplethysmography,&#x0201D;</article-title> in <conf-name>Automatic Face &#x00026; Gesture Recognition (FG 2017), 2017 12th IEEE International Conference on (IEEE)</conf-name>.</citation></ref>
<ref id="B63"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Wei</surname> <given-names>L.</given-names></name> <name><surname>Tian</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Ebrahimi</surname> <given-names>T.</given-names></name> <name><surname>Huang</surname> <given-names>T.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Automatic webcam-based human heart rate measurements using laplacian eigenmap,&#x0201D;</article-title> in <conf-name>Asian Conference on Computer Vision</conf-name> (<conf-loc>Berlin, Heidelberg</conf-loc>: <conf-sponsor>Springer</conf-sponsor>).</citation></ref>
<ref id="B64"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Werner</surname> <given-names>P.</given-names></name> <name><surname>Al-Hamadi</surname> <given-names>A.</given-names></name> <name><surname>Walter</surname> <given-names>S.</given-names></name> <name><surname>Gruss</surname> <given-names>S.</given-names></name> <name><surname>Traue</surname> <given-names>H. C.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;Automatic heart rate estimation from painful faces,&#x0201D;</article-title> in <conf-name>2014 IEEE International Conference on Image Processing (ICIP)</conf-name> (<conf-loc>Paris, France</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>).</citation></ref>
<ref id="B65"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wold</surname> <given-names>S.</given-names></name> <name><surname>Esbensen</surname> <given-names>K.</given-names></name> <name><surname>Geladi</surname> <given-names>P.</given-names></name></person-group> (<year>1987</year>). <article-title>Principal component analysis</article-title>. <source>Chemom. Intell. Lab. Syst.</source> <volume>2.1-3</volume>, <fpage>37</fpage>&#x02013;<lpage>52</lpage>.<pub-id pub-id-type="doi">10.1016/0169-7439(87)80084-9</pub-id></citation></ref>
<ref id="B66"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>H. Y.</given-names></name> <name><surname>Rubinstein</surname> <given-names>M.</given-names></name> <name><surname>Shih</surname> <given-names>E.</given-names></name> <name><surname>Guttag</surname> <given-names>J. V.</given-names></name> <name><surname>Durand</surname> <given-names>F.</given-names></name> <name><surname>Freeman</surname> <given-names>W.</given-names></name></person-group> (<year>2012</year>). <conf-name>ACM Transactions on Graphics (TOG) &#x02013; Proceedings of ACM SIGGRAPH 2012</conf-name>, Vol. <volume>31</volume>. <conf-loc>New York, NY</conf-loc>: <conf-sponsor>ACM</conf-sponsor>.</citation></ref>
<ref id="B67"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>L.</given-names></name> <name><surname>Kunde Rohde</surname> <given-names>G.</given-names></name></person-group> (<year>2014</year>). <article-title>Robust efficient estimation of heart rate pulse from video</article-title>. <source>Biomedical Opt. Exp.</source> <volume>5</volume>, <fpage>1124</fpage>&#x02013;<lpage>1135</lpage>.<pub-id pub-id-type="doi">10.1364/BOE.5.001124</pub-id><pub-id pub-id-type="pmid">24761294</pub-id></citation></ref>
<ref id="B68"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>F.</given-names></name></person-group> (<year>2006</year>). <article-title>&#x072EC;&#x07ACB;&#x05206;&#x091CF;&#x05206;&#x06790;&#x07684;&#x0539F;&#x07406;&#x04E0E;&#x05E94;&#x07528;</article-title>, <publisher-name>Tsinghua University Press.</publisher-name></citation></ref>
<ref id="B69"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>Y.-P.</given-names></name> <name><surname>Kwan</surname> <given-names>B. H.</given-names></name> <name><surname>Lim</surname> <given-names>C. L.</given-names></name> <name><surname>Wong</surname> <given-names>S. L.</given-names></name> <name><surname>Raveendran</surname> <given-names>P.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x0201C;Video-based heart rate measurement using short-time Fourier transform,&#x0201D;</article-title> in <conf-name>Intelligent Signal Processing and Communications Systems (ISPACS), 2013 International Symposium on</conf-name> (<conf-loc>Okinawa, Japan</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>).</citation></ref>
<ref id="B70"><citation citation-type="confpoc"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>Y.-P.</given-names></name> <name><surname>Raveendran</surname> <given-names>P.</given-names></name> <name><surname>Lim</surname> <given-names>C.-L.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;Heart rate estimation from facial images using filter bank,&#x0201D;</article-title> in <conf-name>Communications, Control and Signal Processing (ISCCSP), 2014 6th International Symposium on</conf-name> (<conf-loc>Athens, Greece</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>).</citation></ref>
<ref id="B71"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>Y.-P.</given-names></name> <name><surname>Raveendran</surname> <given-names>P.</given-names></name> <name><surname>Lim</surname> <given-names>C.-L.</given-names></name></person-group> (<year>2015</year>). <article-title>Dynamic heart rate measurements from video sequences</article-title>. <source>Biomed. Opt. Exp.</source> <volume>6</volume>, <fpage>2466</fpage>&#x02013;<lpage>2480</lpage>.<pub-id pub-id-type="doi">10.1364/BOE.6.002466</pub-id><pub-id pub-id-type="pmid">26203374</pub-id></citation></ref>
<ref id="B72"><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Zaunseder</surname> <given-names>S.</given-names></name> <name><surname>Heinke</surname> <given-names>A.</given-names></name> <name><surname>Trumpp</surname> <given-names>A.</given-names></name> <name><surname>Malberg</surname> <given-names>H.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;&#x0201C;Heart beat detection and analysis from videos.&#x0201D; Electronics and Nanotechnology (ELNANO),&#x0201D;</article-title> in <conf-name>2014 IEEE 34th International Conference on</conf-name> (<conf-loc>Kyiv, Ukraine</conf-loc>: <conf-sponsor>IEEE</conf-sponsor>).</citation></ref>
<ref id="B73"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Zha</surname> <given-names>H.</given-names></name></person-group> (<year>2004</year>). <article-title>Principal manifolds and nonlinear dimensionality reduction via tangent space alignment</article-title>. <source>SIAM J Sci. Comput.</source> <volume>26</volume>, <fpage>313</fpage>&#x02013;<lpage>338</lpage>.<pub-id pub-id-type="doi">10.1137/S1064827502419154</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="fn1"><p><sup>1</sup>Cardiio. Cardiio. <uri xlink:href="https://www.cardiio.com/">https://www.cardiio.com/</uri> (Accessed: April 09, 2018).</p></fn>
<fn id="fn2"><p><sup>2</sup>Rapsodo. Pulse Meter. <uri xlink:href="http://www.appspy.com/app/667129/pulse-meter">http://www.appspy.com/app/667129/pulse-meter</uri> (Accessed: April 09, 2018).</p></fn>
<fn id="fn3"><p><sup>3</sup>Philips. Philips vitals signs camera. <uri xlink:href="http://www.vitalsignscamera.com">http://www.vitalsignscamera.com</uri> (Accessed: April 09, 2018).</p></fn>
</fn-group>
</back>
</article>
